Image processing method, system, electronic device, storage medium, and program product

By combining full-image segmentation with quadtree algorithms and user interaction, the problem of over- or under-cutting in automated image cutout is solved, achieving efficient and accurate cutout results.

WO2026016679A1PCT designated stage Publication Date: 2026-01-22HANGZHOU ALIBABA INT INTERNET IND CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/100280
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-16
Filing Date
2025-06-10
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing automated background removal technologies are prone to over- or under-removal when dealing with multiple subjects and complex scenes, failing to meet users' actual background removal needs. Furthermore, offline background removal tools with high professional requirements are inefficient.

Method used

Image sampling is performed using full-image segmentation technology combined with quadtree algorithm. The sampling point information is used to segment the entire image, and the segmented selectable regions are displayed on the client interface. Users can interactively select the target region for image cutout.

Benefits of technology

It improves the accuracy and efficiency of image cutout, avoiding problems of over-cutting or under-cutting. Users can complete image cutout through simple interaction without professional knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025100280_22012026_PF_FP_ABST
    Figure CN2025100280_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing method, a system, an electronic device, a storage medium, and a program product. A method comprises: in combination with image content of a target image, performing point sampling on the target image to obtain sampling point information; performing full image segmentation on the target image on the basis of the sampling point information; displaying, on a client interface, the target image which has undergone full image segmentation, the target image which has undergone full image segmentation comprising multiple segmented selectable regions; and in response to a region selection operation triggered by a user for the multiple selectable regions, performing image matting of the target image on the basis of at least one selected region, to obtain a matted image. By using the present solution, the segmented target image can be adaptively changed on the basis of the image content of the target image, which can improve the full image region division effect of the target image, and balance the time consumption of functions; moreover, interaction between the target image which has undergone the full image segmentation and the user can be used to clarify an actual intention of the user on the basis of the region selected by the user, so that excessive matting or missing matting is avoided, and better image matting service can be provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing methods, systems, electronic devices, storage media, and software products

[0001] This disclosure claims priority to Chinese Patent Application No. 202410957465.X, filed with the China Patent Office on July 16, 2024, entitled “Image Processing Method, System, Electronic Device, Storage Medium and Program Product”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of image matting technology, and in particular to an image processing method, system, electronic device, storage medium, and program product. Background Technology

[0003] In many scenarios, it's necessary to perform image cutout to extract regions or objects of interest from images and create separate images. Currently, image cutout technologies fall into two categories: one is using offline cutout tools, which requires a high level of user expertise and manual cutout, resulting in slow speed and low efficiency; the other is using online automated cutout functions, however, automated cutout is prone to over-cutting or under-cutting in complex images with multiple subjects, leading to results that fail to meet users' actual needs. Therefore, there is an urgent need for an automated image cutout technology solution that can meet users' practical cutout requirements. Summary of the Invention

[0004] In view of the above-mentioned problems mentioned in the background art, this disclosure provides an image processing method, system, electronic device, storage medium and program product that solves the above problems or at least partially solves the above problems, so as to clarify the user's actual image cutout intention through user interaction and improve the quality of automated image cutout.

[0005] In a first aspect, embodiments of this disclosure provide an image processing method, the method comprising:

[0006] By combining the image content of the target image, point sampling is performed on the target image to obtain sampling point information;

[0007] Based on the sampling point information, the target image is segmented in its entirety;

[0008] The client interface displays a segmentation map corresponding to the target image, the segmentation map including multiple selectable regions segmented out;

[0009] In response to a selection operation triggered for the plurality of selectable regions, the target image is cut out based on at least one selected selectable region to obtain a cut-out image.

[0010] Secondly, embodiments of this disclosure provide another image processing method. This method includes:

[0011] Based on the color information of the target image, obtain the segmentation prompt information of the target image;

[0012] Using the segmentation prompt information and the target image, a preset full-image segmentation model is triggered to perform full-image segmentation on the target image;

[0013] On the client interface, a segmentation map corresponding to the target image is displayed, the segmentation map including multiple selectable regions segmented out;

[0014] In response to a selection operation triggered for the plurality of selectable regions, the target image is processed by masking based on at least one selected selectable region to obtain a masked image;

[0015] The cut-out image is displayed on the client interface.

[0016] Thirdly, embodiments of this disclosure provide yet another image processing method. This method includes:

[0017] Obtain the product image to be cut out;

[0018] Based on the image content of the product image, point sampling is performed on the product image to obtain sampling point information;

[0019] Based on the sampling point information, the product image is segmented into a whole image;

[0020] On the client interface, a segmentation diagram corresponding to the product image is displayed, the segmentation diagram including multiple segmented object regions;

[0021] In response to a selection operation triggered for the multiple object regions, based on at least one selected object region, the main body of the target object is separated from the product image to obtain a display image of the main body of the target object;

[0022] Displaying a diagram of the main body of the target object.

[0023] Fourthly, embodiments of this disclosure provide an image processing system. The system includes:

[0024] The client is used to acquire the target image to be cut out;

[0025] The server is used to implement the steps in the image processing methods provided in the embodiments of this disclosure.

[0026] Fifthly, embodiments of this disclosure provide an electronic device. The electronic device includes a memory and a processor, wherein the memory is used to store a program; the processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the various method embodiments provided in this disclosure;

[0027] Sixthly, embodiments of this disclosure provide a computer-readable storage medium. This computer-readable storage medium stores a computer program or instructions; when the computer program or instructions are executed by a processor, the steps in the various method embodiments provided in this disclosure can be implemented.

[0028] In a seventh aspect, embodiments of this disclosure provide a computer program product. This computer program product includes a computer program or instructions that, when executed by a processor, cause the processor to perform the steps described in the various method embodiments of this disclosure.

[0029] The technical solution provided in this disclosure is based on the image content of the target image (such as the color information of the target image), performing point sampling on the target image to obtain sampling point information, and then performing full-image segmentation on the target image based on this sampling point information. This full-image segmentation method allows the segmentation to adapt to the image content information of the target image, improving the full-image region division effect of the target image and balancing the time consumption of the function. For example, the sampling density can be increased for color-rich and complex regions in the target image to achieve more refined segmentation, while the sampling density can be reduced for regions with similar colors to reduce redundant segmentation calculations. Furthermore, a segmentation map corresponding to the target image will be displayed on the client interface. This segmentation map includes multiple selectable regions, and the user can select regions from these regions. Based on at least one selected region, the target image can be cut out to obtain a corresponding cut-out image. The target image can be a product image, and the obtained cut-out image can be a display image of the target object separated from the product image. This solution provides an interactive background removal method. By interacting with the user, it can clarify the user's actual intention based on the selected area, avoiding excessive or incomplete background removal and providing users with better background removal services. Moreover, users can complete the selection in a relatively simple way (such as clicking) without the need for professional knowledge. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 is a schematic flowchart of an image processing method provided in an embodiment of this disclosure;

[0032] Figure 2a is a schematic diagram illustrating the principle of image matting provided in an embodiment of this disclosure;

[0033] Figure 2b is an example diagram of displaying a cutout image provided in an embodiment of this disclosure;

[0034] Figure 3 is a schematic diagram of the results of a full-image segmentation model provided in an embodiment of this disclosure;

[0035] Figures 4 and 5 are schematic diagrams illustrating the principles of image matting provided in other embodiments of this disclosure;

[0036] Figures 6 and 7 are schematic flowcharts of image processing methods provided in two other embodiments of this disclosure;

[0037] Figures 8 to 10 are schematic diagrams of the image processing apparatus provided in the embodiments of this disclosure;

[0038] Figure 11 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0039] Image cutout primarily involves extracting the target object from an image, and it has wide applications in image editing and creative compositing. For example, for merchants on e-commerce platforms, whether it's displaying main or secondary images of products or creating product creative images for advertising, they all need to extract the main product image from the product image through image cutout. Currently, automated image cutout tools are mainly used to achieve this. However, when the image to be cut out contains multiple and complex objects, existing automated image cutout tools inevitably result in over-cutting or under-cutting, or manual cutout may be necessary.

[0040] For example, all existing photo editing tools offer a brush-based background removal function, but this function requires manual background removal, which demands a high level of expertise from the user.

[0041] For example, various existing image editing tools include one-click background removal functionality; or, with the release and rapid development of technologies such as GPT (Generative Pre-Training), SAM (Segment Anything Model), and image-to-image processing, some platform-based merchant assistant products have also launched automated background removal functions. These merchant assistant products, in addition to automated background removal, also include other image functions such as creating white background images and scene images. Automated background removal is a core and fundamental capability, and it's involved in almost all other image function workflows. The aforementioned one-click or automated background removal functions utilize image understanding, detection, and segmentation modules to allow users to directly obtain the algorithm's output image background removal results without any interaction cost. However, based on real-world user experience, while automated background removal solutions can solve most input image background removal problems, there are still some difficult background removal scenarios that cannot be resolved. For example, when the original image contains multiple objects, the lack of interactive input makes it difficult to identify the user's actual intent, making it impossible to determine whether the user wants to extract all objects in the image or just one. In addition, when the original image has a complex scene, the automatic image cutout process is relatively subjective in whether to extract accompanying items, packaging boxes, icons, logos, etc. All of these factors may lead to the final cutout result being either over-cut or under-cut, making the cutout result incomprehensible and inconsistent with the user's actual intent.

[0042] To address the problems of the aforementioned image matting schemes, this disclosure provides an image processing technique. This technique primarily involves segmenting the entire image, numbering each segmented region, and providing these regions to the front-end (client) for user selection. The selected regions are then merged and post-processed to achieve the matting function. Specifically, the full image segmentation process combines a quadtree algorithm with a full image segmentation model to achieve more reasonable and precise segmentation of the entire image. The quadtree algorithm is mainly used for image sampling, employing the obtained sampling points as cue points to guide the detection of object content within the image. The quadtree algorithm is detailed below and will not be elaborated upon here.

[0043] To enable those skilled in the art to better understand the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0044] In some processes described in the specification, claims, and accompanying drawings of this disclosure, multiple operations appearing in a specific order are included. These operations may be executed out of order or in parallel. Operation numbers such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the terms "first," "second," etc., used herein are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types. The term "or / and" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A or / and B indicates that A can exist alone, A and B can exist simultaneously, or B can exist alone. The character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship. It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system including said element. Furthermore, the following embodiments are merely some embodiments of this disclosure, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0045] The technical solutions provided in the embodiments of this disclosure will be described below.

[0046] The execution subject of the various method embodiments provided in this disclosure can be an electronic device with logical operation functions, and the electronic device can be a server. The server can be a single server, a service cluster composed of multiple servers, a cloud server, or a virtual server, etc., and this disclosure does not specifically limit it in this regard.

[0047] Figure 1 shows a schematic flowchart of an image processing method provided in an embodiment of this disclosure. As shown in Figure 1, the image processing method includes the following steps:

[0048] 101. Based on the image content of the target image, perform point sampling on the target image to obtain sampling point information;

[0049] 102. Based on the sampling point information, perform full-image segmentation on the target image;

[0050] 103. Display the segmentation map corresponding to the target image on the client interface, the segmentation map including multiple selectable regions segmented out;

[0051] 104. In response to a selection operation triggered for the plurality of selectable regions, the target image is cut out based on at least one selected selectable region to obtain a cut-out image.

[0052] In the above 101, the target image can be, but is not limited to, a product image, and can be input by the user. In specific implementation, the target image can be selected by the user from a gallery. The gallery can be locally stored by the execution entity as described in this embodiment (such as a photo album), or it can be stored on other devices on the network side. This embodiment does not limit this.

[0053] For example, as shown in Figure 2a, the merchant opens the Merchant Business Assistant page provided by the e-commerce platform through client 1 and clicks the "Smart Background Removal" control on the Merchant Business Assistant page. At this time, the Smart Background Removal page 10 will be displayed on the client interface. The user can operate the corresponding image selection control on the Smart Background Removal page 10 to input the target image to be removed. For example, the user can operate the "Select Image from Network Side" control or the "Select Local Image" control to input the target image to be removed. The input target image (such as product image P) can be displayed in area Z and can also be obtained by server 2 for subsequent processing.

[0054] Based on the above, that is, before performing step 101, the method provided in this embodiment further includes:

[0055] 100. Obtain the target image to be cut out.

[0056] In one specific implementable solution, step 100, "obtaining the target image to be cut out," includes:

[0057] 1001. Display the background removal service page (such as the intelligent background removal page 10) on the client;

[0058] 1002. In response to the user's input operation through the image cutout service page, obtain the target image input by the user.

[0059] The client described in this embodiment may be, but is not limited to, a smartphone, tablet, laptop, desktop computer, etc.

[0060] In addition to acquiring the target image, considering that the target object may have multiple subjects and the scene is complex, in order to enable the automated image cutout to better meet the actual needs of users, this embodiment will first perform full image segmentation on the target image to predict and segment out all possible objects in the target image for user interaction, thereby clarifying the user's intention through interaction and providing the user with better image cutout services.

[0061] Currently, full-image segmentation is primarily achieved using corresponding full-image segmentation models, such as the SAM model. The SAM model can receive inputs such as text, coordinate points, and bounding boxes as prompts, and then, based on the auxiliary information provided by the prompts (fuzzy hints), it can identify and segment every possible object element in the image. Existing SAM models, when implementing automatic full-image segmentation, use sampling points obtained through uniform sampling as prompts. Uniform sampling involves dividing an image into M*N networks with equal spacing in the horizontal and vertical directions, and using the center point of each network as a sampling point (as shown in Figure 4, 256 points). Because uniform sampling does not incorporate the image's own features, it often suffers from insufficient sampling of complex images, leading to problems such as inaccurate extraction of fine structures and unreasonable image region division during final segmentation. Furthermore, uniform sampling introduces many unnecessary sampling points, resulting in a large amount of redundant computation during subsequent segmentation. To improve the rationality of image region division and reduce redundant computation during full image segmentation, the solution provided in this embodiment is to optimize sampling by combining the image content itself (such as color information (which is also pixel value information)) so that the sampled point set obtained by sampling can be used as the prompt of a full image segmentation model such as the SAM model.

[0062] Therefore, in a specific implementable solution, the image content of the target image includes color information; and step 101, "combining the image content of the target image, performing point sampling on the target image to obtain sampling point information," can specifically mean: performing point sampling on the target image based on the color information of the target image to obtain sampling point information. In one implementable technical solution, a quadtree algorithm can be used, but is not limited to, to perform point sampling on the target image based on the color information of the target image. Therefore, a specific implementation of step 101 may include:

[0063] 10111. Using the quadtree algorithm, the target image is recursively divided based on its color information, so that the center point of each sub-region is determined as the sampling point.

[0064] The quadtree algorithm described above recursively divides the image into regions based on its color information until a certain termination condition is met. In practice, the specific implementation process of using the quadtree algorithm to divide and sample the target image is as follows:

[0065] 1) Treat the entire target image as a single rectangular region, and use it as the root node N0 of the quadtree.

[0066] 2) Divide the rectangular region of root node N0 into four equal sub-regions along the horizontal and vertical directions, forming four sub-nodes N1, N2, N3, and N4. If the region represented by any of the sub-nodes N1, N2, N3, and N4 does not meet the set division termination condition, it can be further divided into four equal sub-regions along the horizontal and vertical directions. This process is repeated to perform image region hierarchical division.

[0067] When performing image region hierarchical segmentation, the segmentation termination conditions may include:

[0068] Homogeneity condition: When the color richness of the current region is lower than the preset first threshold (such as when the pixel values ​​of all pixels in the region are basically the same), the division will not continue.

[0069] Tree depth condition: When the depth of the partitioned tree is greater than the preset second threshold, the partitioning will stop.

[0070] Region area threshold condition: When the area of ​​a sub-region is less than the preset third threshold, the sub-region will no longer be divided.

[0071] 3) The center point of each sub-region is determined as the sampling point.

[0072] This embodiment utilizes a quadtree algorithm to partition the target image into regions for sampling. The partitioned regions can adaptively change according to the complexity of the color information in the input image, while the set partition termination condition avoids overly fine partitioning. After partitioning, the center points of each sub-region can be used as sampling points. Since regions with similar colors may belong to the same object, the sampling point information obtained based on the quadtree algorithm is beneficial for guiding subsequent full-image segmentation of the target image. The sampling point information includes the coordinates of each sampling point (157 points shown in Figure 4). Each sampling point is used to guide the full-image segmentation model to understand and analyze different parts of the image, detecting and exploring all possible foreground objects in the image.

[0073] Compared to uniform sampling, this embodiment uses a quadtree algorithm to sample the target image, allowing the sampling point density to adaptively change according to the image information (such as color information) of the target image. For example, the sampling density can be reduced in areas with similar colors, thereby reducing subsequent redundant calculations and the service time of the image matting function. Furthermore, it can achieve more refined subdivision of complex areas with rich colors in the image to increase the sampling point density.

[0074] Accordingly, step 102 above, "performing full-image segmentation of the target image based on the sampling point information," may include the following steps:

[0075] 1021. Input the sampling point information and the target image into the full image segmentation model, and execute the full image segmentation model to output multiple mask images; the mask images are used to represent the regions where the corresponding objects are located in the target image;

[0076] 1022. Perform full-image segmentation on the target image based on the multiple mask images to divide the target image into the multiple selectable regions.

[0077] In the above 1021, the preset full-image segmentation model can be, but is not limited to, the SAM model. Inputting the sampling point information (as segmentation prompt information) into the full-image segmentation model can guide the model to complete the corresponding segmentation task.

[0078] For example, referring to the schematic diagram of the full-image segmentation model shown in Figure 3, the full-image segmentation model includes a first encoder, a second encoder, and a mask encoder. The first encoder is an image encoder, and the second encoder is a prompt encoder. After the target image and segmentation prompt information are input into the full-image segmentation model, the first encoder encodes the target image to generate a first feature vector (i.e., an image embedding vector, used to describe the image features), and the second encoder encodes the segmentation prompt information to generate a second feature vector (i.e., a prompt embedding vector, used to describe the sampling point location features). The mask encoder analyzes and predicts all possible objects in the target image based on the first and second feature vectors, thereby outputting multiple corresponding mask images. The mask image is a binary image, its size matching the target image size, used to indicate the region where the corresponding object is located in the target image.

[0079] Considering that full-image segmentation models often encounter duplicate detections during the process of detecting objects based on segmentation prompts and target images, resulting in some predicted masks being redundant, this embodiment of the full-image segmentation model performs post-processing on the predicted mask images to remove duplicates. The post-processing method used for deduplication is the NMS (Non-Maximum Suppression) algorithm, whose main function is to eliminate redundant masks and retain the masks with better potential indicators. For a detailed description of using the NMS algorithm for deduplication, please refer to existing related materials. Finally, the full-image segmentation model outputs multiple mask images obtained after deduplication.

[0080] In the above 1022, when performing full-image segmentation of the target image based on multiple mask images, the target image will be divided into multiple selectable regions. Different selectable regions can be marked with different colors (as shown in image P' in Figure 2a), and different selectable regions represent different objects or parts in the target image.

[0081] In practice, during the segmentation process, the target image can be further divided based on the connectivity of the non-masked regions. This serves two purposes: firstly, it facilitates subsequent display and interactive selection; secondly, since disconnected regions in the non-masked areas may not belong to the background, it is better to perform finer region segmentation based on the connectivity of the non-masked areas to avoid omissions.

[0082] That is, a specific implementation of the above-mentioned 1022 "performing full-image segmentation of the target image based on the multiple mask images to divide the target image into the multiple selectable regions" may include:

[0083] 10221. Based on the plurality of mask images, determine the non-mask region of the target image;

[0084] 10222. Based on the multiple mask images and the connectivity information of the non-mask regions, the target image is divided into multiple selectable regions.

[0085] Furthermore, for the multiple selectable regions, each region can be numbered according to the area of ​​its mask image. In practice, the regions can first be sorted in descending order of mask image area, and then numbered starting from 0 or 1 based on the sorting result. Numbering regions based on the descending order results allows for smaller region numbers for regions with larger mask image areas. Of course, other sorting methods can also be used, such as ascending order of mask image area, followed by numbering starting from 0 or 1. This embodiment does not impose specific limitations on this. Descending order arranges regions from largest to smallest area, while ascending order arranges regions from smallest to largest area.

[0086] The area of ​​the aforementioned mask refers to the area representing the object region within the mask. When numbering, if at least two overlapping selectable regions exist, the overlapping portion can be marked according to the region number of the two selectable regions with the largest region number. Then, the region numbers of each selectable region can be sent to the client so that the client can associate the region numbers of each selectable region with the corresponding selectable regions when displaying the segmented target image, making it easy for the user to identify which regions are selectable.

[0087] Therefore, the method provided in this embodiment may further include the following steps:

[0088] 105. Based on the area of ​​the mask map corresponding to each of the multiple optional regions, determine the region number of each of the multiple optional regions, so as to associate and display the multiple optional regions and their respective region numbers on the client interface.

[0089] Where at least two of the multiple optional regions partially overlap, the region number with the largest region number among the at least two optional regions is determined as the region number corresponding to the overlapping region of the at least two optional regions.

[0090] For example, the area number can be displayed in one corner of the corresponding selectable area, etc., without specific limitations.

[0091] Furthermore, by performing step 103 above, the segmentation map corresponding to the target image can be displayed on the client interface, such as in the selection page 11 shown in Figure 2a. When displaying the segmentation map P” corresponding to the target image on the selection page, the corresponding region number can also be displayed in association with each selectable region segmented by the segmentation map (not fully shown in Figure 2a).

[0092] Users can select from multiple selectable regions in the segmented image corresponding to the displayed target image using interactive methods provided by the client (such as mouse, voice, touch, etc.).

[0093] For example, referring to Figure 2a, if a user selects the optional region numbered ①, the selected region will be highlighted (e.g., its color will change to blue), indicating that the optional region is selected. Afterwards, when the user clicks on the optional region numbered ① again, the selection of that optional region will be deselected. Correspondingly, the optional region will return to its original color, indicating that it is unselected. Similarly, the user can select or deselect other optional regions. After selecting at least one optional region according to their needs, the user can click the "Confirm Selection" control. In response to the user's operation on the "Confirm Selection" control, the server will trigger the execution of step 104 above, "Cutting out the target image based on at least one selected optional region." Alternatively, the user can click the "Cancel" control to cancel all selected regions with a single click.

[0094] In addition to selecting selectable areas by clicking, users can also select areas through other interactive methods. For example, if the selection page 11 provides text or voice interaction methods, users can also enter the corresponding area number information to complete the selection. This embodiment does not specifically limit the specific implementation method of the selection.

[0095] Based on this example, in a specific implementable solution, the method described in this embodiment may further include the following steps:

[0096] In response to a click operation, at least one selectable region is selected; or

[0097] In response to the area number input operation, the selectable area corresponding to the input area number is the selected selectable area.

[0098] Based on the above, this embodiment uses full-image segmentation and displays the selectable regions divided by the full-image segmentation, allowing users to select regions through corresponding interactive methods (such as clicking, entering region numbers, etc.). Thus, the user's actual intention can be clearly understood based on the selected regions, which can effectively solve problems such as over-slicing or under-slicing caused by unclear cutout targets. It can also provide users with more flexible and simple automated cutout services, optimizing the user's interactive cutout service experience.

[0099] In step 104 above, the mask images corresponding to at least one selected region can be merged, and the merged mask image can be processed by erosion, dilation, edge optimization, etc., for use in matting the target image. That is, in a specific implementable technical solution, step 104 above, "matting the target image based on at least one selected selectable region to obtain a matted image," can include:

[0100] 1031. Merge the mask images corresponding to at least one selected region to obtain a merged mask image;

[0101] 1032. Process the merged mask image to obtain the target mask image; the processing includes erosion expansion and edge optimization;

[0102] 1033. Based on the target mask image, the target image is cut out to obtain a cut-out image.

[0103] The purpose of performing erosion and dilation on the obtained merged mask image is to construct a trimap (also known as a three-part image).

[0104] For example, in the merged mask image, the white area represents the target area to be extracted, and the pixel value in the white area is 255. An erosion operation is performed on the merged mask image to obtain an erosion map. In the erosion map, the pixel value of areas with a pixel value greater than 127 is defined as 255, and this area is definitely a foreground area. A dilation operation is performed on the merged mask image to obtain a dilation map. In the dilation map, the pixel value of areas with a pixel value greater than 127 that do not overlap with the aforementioned definitely foreground areas is defined as 127 (gray, indicating this area is undetermined), and the pixel value of areas with a pixel value less than or equal to 127 is defined as 0 (black, indicating this area is a background area). Based on the obtained erosion and dilation maps, a Trimap image can be generated, as shown in Figure 5(b). In the Trimap image, three values ​​are used to identify the following three types of areas: definitely foreground (white), definitely background (black), and uncertain (gray).

[0105] Furthermore, the Trimap image and the target image are input into a corresponding matting model, such as a matting model. Executing the matting model will output the corresponding matting result (including the matted image). Specifically, as shown in Figure 5(b), after inputting the Trimap image and the target image into the corresponding matting model, the matting model will perform edge optimization processing on the Trimap image to obtain a target mask image. Then, based on the target mask image, the target region is separated from the target image, and the matting result is output (the matted image shown in Figure 5(b)). In addition, the target mask image can also be output. That is, the matting model provided in this embodiment can be used to implement edge optimization and matting. The algorithm used for edge optimization can be, but is not limited to, edge sharpening algorithms based on specific grayscale distributions, optimization algorithms based on line and surface features, etc., and is not specifically limited here. Among them, the edge optimization performed by the matting model on the Trimap image is reflected in predicting values ​​between 0 and 255 for edge regions, rather than direct binarization.

[0106] Figure 5(a) shows the cutout result of directly cutting out the target image based on the obtained merged mask. The cutout result is the cutout image shown in Figure 5(a). It can be clearly seen that the edges of the aforementioned merged mask and the corresponding cutout image are relatively flawed, such as having jagged edges.

[0107] As mentioned above, various post-processing techniques such as erosion and expansion, and edge optimization can effectively improve the quality of image matting.

[0108] The cut-out images obtained above can be used in subsequent processes such as generating white background images, scene images, and image editing.

[0109] Furthermore, the method provided in this embodiment may also include the following steps:

[0110] 106. Display the cut-out image on the client interface.

[0111] In practice, the cut-out image and the original target image can be displayed on the same client interface (as shown in Figure 2b); alternatively, the cut-out image can be displayed separately on a separate page. The purpose of displaying the cut-out image is to facilitate user review and confirmation of whether to proceed with subsequent processing. For example, the cut-out image can be uploaded as a white background image to the corresponding publishing page, downloaded to the local machine, or used for scene editing to generate a corresponding scene image and publish it, etc.

[0112] In summary, the technical solution provided in this embodiment is based on the image content of the target image (such as the color information of the target image) and uses a preset full-image segmentation model to perform full-image segmentation on the target image. This full-image segmentation method allows the segmentation to adapt to the image content information of the target image, improving the full-image region division effect of the target image and balancing the time consumption of the function. For example, the sampling density can be increased for rich and complex color regions in the target image to achieve more refined segmentation, while the sampling density can be reduced for regions with similar colors to reduce redundant segmentation calculations. Furthermore, the target object after full-image segmentation will be displayed on the client interface. The target image after full-image segmentation is divided into multiple selectable regions. The user can select regions from these regions, and the target image can be cut out based on at least one region selected by the user to obtain the corresponding cutout interface. This solution provides an interactive cutout solution. Through interaction with the user, the actual intention of the user can be clarified based on the user's selection, avoiding over-cutting or under-cutting, and providing the user with a better cutout service; moreover, the user can complete the selection in a relatively simple way (such as point-and-click), without the need for professional knowledge.

[0113] This disclosure also provides several other embodiments of image processing methods. Specifically:

[0114] Figure 6 illustrates an image processing method provided by another embodiment of this disclosure. As shown in Figure 6, the image processing method includes the following steps:

[0115] 201. Obtain segmentation prompt information for the target image based on its color information;

[0116] 202. Using the segmentation prompt information and the target image, trigger a preset full-image segmentation model to perform full-image segmentation on the target image;

[0117] 203. On the client interface, a segmentation map corresponding to the target image is displayed, the segmentation map including multiple selectable regions segmented out;

[0118] 204. In response to a selection operation triggered for the plurality of selectable regions, perform image cutout processing on the target image based on at least one selected selectable region to obtain a cutout image;

[0119] 205. Display the cutout image on the client interface.

[0120] In the above 201, the segmentation prompt information is also the sampling point information described in other embodiments of this disclosure.

[0121] In one feasible technical solution, the above-mentioned 201 "obtaining segmentation prompt information of the target image based on the color information of the target image" may include:

[0122] 2011. Using the quadtree algorithm, the target image is recursively divided based on its color information, so that the center point of each sub-region is determined as the sampling point;

[0123] The segmentation prompt information refers to the identified sampling point information.

[0124] For a detailed description of the implementation of step 2011 and steps 202 to 204, please refer to the relevant content in other embodiments.

[0125] In step 205 above, the cutout image and the original target image can be displayed on the same client interface (as shown in Figure 2b), or the cutout image can be displayed separately. The purpose of displaying the cutout image is to facilitate user review and confirmation of whether to proceed with subsequent processing, such as uploading the cutout image as a white background image to the corresponding publishing page, downloading it locally, or editing the scene to generate a corresponding scene image and publishing it, etc.

[0126] In addition to the steps described above, the method provided in this embodiment may also include other steps. For details of these other steps and their specific implementation, please refer to the relevant content in other embodiments, which will not be repeated here.

[0127] Figure 7 shows a schematic flowchart of an image processing method provided in another embodiment of this disclosure. As shown in Figure 7, the image processing method includes:

[0128] 301. Obtain the product image to be cut out;

[0129] 302. Based on the image content of the product image, perform point sampling on the product image to obtain sampling point information;

[0130] 303. Based on the sampling point information, perform full image segmentation on the product image;

[0131] 304. On the client interface, a segmentation diagram corresponding to the product image is displayed, the segmentation diagram including multiple segmented object regions;

[0132] 305. In response to a selection operation triggered for the plurality of object regions, based on at least one selected object region, separate the main body of the target object from the product image to obtain a display image of the main body of the target object;

[0133] 306. Display a diagram of the main body of the target object.

[0134] For a detailed description of the implementation of steps 301 to 306 above, please refer to the relevant content in other embodiments. Furthermore, in addition to the steps described above, the method provided in this disclosure may also include other steps. For details of these other steps and their specific implementation, please refer to the relevant content in other embodiments, and will not be repeated here.

[0135] Furthermore, this disclosure also provides embodiments of image processing systems and image processing apparatuses corresponding to the above-described method embodiments. Specifically, they are as follows:

[0136] As shown in Figure 2a, the image processing system provided in this embodiment includes: a client 1 and a server 2.

[0137] In one example, the functions of client 1 and server 2 are as follows:

[0138] Client 1 is used to acquire the target image to be cut out and send the target image to the server; receive and display the segmentation map of the target image fed back by the server, the segmentation map including multiple selectable regions; respond to the selection operation triggered for the multiple selectable regions, cut out the target image according to at least one selected selectable region, and obtain and display the cut-out image;

[0139] Server 2 is used to perform point sampling on the target image based on the image content of the target image to obtain sampling point information; and to perform full-image segmentation on the target image based on the sampling point information to obtain a segmentation map; or, server 2 is used to trigger a preset full-image segmentation model to perform full-image segmentation on the target image to obtain a segmentation map using the segmentation prompt information and the target image.

[0140] The aforementioned "cutting out the target image based on at least one selected optional region to obtain a cutout image" can also be performed by the server. The client sends at least one selected optional region to the server. The server performs cutout processing on the target image based on the selected at least one region to obtain a cutout image, and then sends the cutout image to client 2.

[0141] In another example, the client and server functionalities are as follows:

[0142] Client 1 is used to obtain a product image to be cut out and send the product image to the server; receive and display the segmentation image fed back by the server, the segmentation image including multiple segmented object regions; respond to the selection operation triggered for the multiple object regions, and separate the target object body from the product image according to at least one selected object region to obtain a display image of the target object body; and display the display image of the target object body.

[0143] Server 2 is used to perform point sampling on the product image based on its image content to obtain sampling point information. Based on the sampling point information, the product image is segmented to obtain a segmented image corresponding to the product image.

[0144] Similarly, the aforementioned "extracting the target object from the product image" can be performed by the server. The client sends at least one selected object region to the server, and the server, based on the selected at least one object region, extracts the target object from the product image to obtain a display image of the target object; the display image is then sent back to the client.

[0145] For a detailed description of the functions of each terminal in the system provided in this embodiment, please refer to the relevant content in other embodiments.

[0146] Figure 8 shows a schematic diagram of an image processing apparatus according to an embodiment of the present disclosure. As shown in Figure 8, the image processing apparatus includes: a sampling and segmentation module 41, a display module 42, and a matting module 43.

[0147] The sampling and segmentation module 41 is used to combine the image content of the target image, perform point sampling on the target image to obtain sampling point information, and perform full image segmentation on the target image based on the sampling point information;

[0148] The display module 42 is used to display the segmentation map corresponding to the target image on the client interface; wherein the segmentation map includes multiple selectable regions segmented out;

[0149] The image cutout module 43 is used to respond to the selection operation triggered by the user for the plurality of selectable regions, and to cut out the target image based on at least one selected selectable region to obtain a cutout image.

[0150] Optionally, the sampling segmentation module 41 described above, when used to segment the target image based on the sampling point information, can specifically be used to: take the sampling point information and the target image as input, execute a full-image segmentation model to output multiple mask images; the mask images are used to represent the regions where the corresponding objects are located in the target image; and perform full-image segmentation on the target image according to the multiple mask images to divide the target image into the multiple selectable regions.

[0151] Optionally, the image content includes color information. Furthermore, the sampling segmentation module 41, when used to combine the image content of the target image to perform point sampling on the target image to obtain sampling point information, can specifically be used to: recursively divide the target image based on the color information of the target image using a quadtree algorithm to obtain multiple sub-regions; and determine the center point of each sub-region within the multiple sub-regions as a sampling point.

[0152] Optionally, the sampling and segmentation module 41 described above, when used to perform full-image segmentation of the target image based on the multiple mask images to divide the target image into the multiple selectable regions, may specifically be used to: determine the non-masked regions of the target image based on the multiple mask images; and divide the target image into multiple selectable regions based on the multiple mask images and the connectivity information of the non-masked regions.

[0153] Optionally, the device further includes: a determining module, configured to determine the region number of the optional region based on the area of ​​the mask image corresponding to the optional region, so as to display the optional region and the region number of the optional region together on the client interface; when multiple optional regions have partial overlap, the largest region number among the region numbers is determined as the region number corresponding to the overlapping region.

[0154] Optionally, the above-mentioned cutout module 43, when used to determine at least one selected optional region in response to a selection operation triggered by the user for the plurality of optional regions, may be specifically used to: determine at least one optional region in the selected state as at least one selected optional region in response to a click operation; or, in response to a region number input operation, determine at least one selected optional region based on the input region number information.

[0155] Optionally, the above-mentioned matting module 43, when used to matte the target image based on at least one selected optional region to obtain a matted image, may specifically be used to: merge the mask images corresponding to the at least one selected optional region to obtain a merged mask image; process the merged mask image to obtain a target mask image; wherein, the processing includes erosion dilation and edge optimization; and matte the target image based on the target mask image to obtain a matted image.

[0156] It should be noted that the image processing apparatus provided in the above embodiments can implement the technical solutions described in the corresponding method embodiments. The specific implementation principles of each module or unit can be found in the relevant content of the corresponding method embodiments, and will not be elaborated here.

[0157] Figure 9 shows a schematic diagram of an image processing apparatus according to another embodiment of the present disclosure. As shown in Figure 9, the image processing apparatus includes: an acquisition module 51, a segmentation module 52, a display module 53, a background removal module 54, and a display module 55.

[0158] The acquisition module 51 is used to acquire segmentation prompt information of the target image based on the color information of the target image;

[0159] The segmentation module 52 is used to trigger a preset full-image segmentation model to perform full-image segmentation on the target image using the segmentation prompt information and the target image;

[0160] The display module 53 is used to display the segmentation map corresponding to the target image on the client interface; wherein the segmentation map includes multiple selectable regions segmented out.

[0161] The cutout module 54 is used to respond to the user's selection operation triggered for the plurality of selectable regions, and to perform cutout processing on the target image based on at least one selected selectable region to obtain a cutout image;

[0162] Display module 55 is used to display the cutout image on the client interface.

[0163] It should be noted that the image processing apparatus provided in the above embodiments can implement the technical solutions described in the corresponding method embodiments. The specific implementation principles of each module or unit can be found in the relevant content of the corresponding method embodiments, and will not be elaborated here.

[0164] Figure 10 shows a schematic diagram of an image processing apparatus according to another embodiment of the present disclosure. As shown in Figure 10, the image processing apparatus includes: an acquisition module 61, a segmentation and display module 62, a separation module 63, and a display module 64.

[0165] The segmentation module 61 is used to obtain the product image to be cut out;

[0166] The sampling segmentation display module 62 is used to perform point sampling on the product image based on the image content of the product image to obtain sampling point information; perform full image segmentation on the product image based on the sampling point information; and display the segmentation map corresponding to the product image on the client interface; wherein, the segmentation map includes multiple segmented object regions (which are optional object regions);

[0167] The separation module 63 is used to respond to the user's selection operation triggered for the multiple object areas, and to separate the main body of the target object from the product image based on at least one selected object area to obtain a display image of the main body of the target object.

[0168] Display module 64 is used to display a display image of the target object on the client interface so that the user can review and confirm whether it is displayed on the corresponding product page.

[0169] It should be noted that the image processing apparatus provided in the above embodiments can implement the technical solutions described in the corresponding method embodiments. The specific implementation principles of each module or unit can be found in the relevant content of the corresponding method embodiments, and will not be elaborated here.

[0170] Figure 11 shows a schematic diagram of the structure of an electronic device provided according to an embodiment of the present disclosure. As shown in Figure 11, the electronic device includes a memory 71 and a processor 72. The memory 71 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Specifically,

[0171] The aforementioned memory 71 is used to store programs;

[0172] The processor 72, coupled to the memory 71, is used to execute the program stored in the memory for steps or functions in the methods provided in the embodiments of this disclosure.

[0173] Furthermore, as shown in Figure 11, the electronic device also includes other components such as a communication component 73, a power supply component 74, and an audio component 75. Figure 11 only schematically shows some of the components and does not imply that the electronic device includes only the components shown in Figure 11. The electronic device may be a server.

[0174] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a computer, can implement the method steps provided in the above embodiments.

[0175] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, enables the processor to implement the method steps or functions provided in the above embodiments.

[0176] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.

Claims

1. An image processing method, wherein, The method comprises the following steps: Point sampling is performed on the target image based on the image content of the target image to obtain sample point information; Full-image segmentation is performed on the target image based on the sample point information; A segmentation map corresponding to the target image is displayed on a client interface, and the segmentation map comprises a plurality of selectable regions segmented out; In response to a selection operation triggered on the plurality of selectable regions, the target image is cut out based on at least one selected region to obtain a cut-out image.

2. The method of claim 1, wherein, The full-image segmentation of the target image based on the sample point information comprises: The sample point information and the target image are input into a full-image segmentation model to output a plurality of mask maps; the mask maps are used to represent the regions of the corresponding objects in the target image; The target image is segmented based on the plurality of mask maps to divide the target image into the plurality of selectable regions.

3. The method of claim 2, wherein, The image content comprises color information; and the point sampling of the target image based on the image content of the target image comprises: The target image is recursively divided based on the color information of the target image by using a quadtree algorithm to obtain a plurality of sub-regions; The center point of each sub-region in the plurality of sub-regions is determined as a sample point.

4. The method of claim 2 or 3, wherein, The full-image segmentation of the target image based on the plurality of mask maps to divide the target image into the plurality of selectable regions comprises: The non-mask region of the target image is determined based on the plurality of mask maps; The target image is divided into the plurality of selectable regions based on the plurality of mask maps and the connectivity information of the non-mask region.

5. The method of any one of claims 2 to 4, wherein, Further comprising: The area number of the selectable region is determined based on the area of the mask map corresponding to the selectable region, so that the selectable region and the area number of the selectable region are associated and displayed on the client interface; When the plurality of selectable regions exist in a partially overlapped manner, the largest area number in the area numbers is determined as the area number corresponding to the overlapped region.

6. The method of any one of claims 1 to 5, wherein, Further comprising: In response to a point selection operation, at least one selectable region in a selected state is at least one selected region; Or In response to an area number input operation, the selectable region corresponding to the input area number is the selected region.

7. The method of any one of claims 2 to 5, wherein, The cut-out of the target image based on at least one selected region to obtain a cut-out image comprises: The mask maps corresponding to the at least one selected region are merged to obtain a merged mask map; The merged mask map is processed to obtain a target mask map; wherein the processing comprises erosion and expansion, edge optimization; The target image is cut out based on the target mask map to obtain a cut-out image.

8. An image processing method, wherein, The method comprises the following steps: Segmentation prompt information of the target image is obtained based on the color information of the target image; A preset full-image segmentation model is triggered based on the segmentation prompt information and the target image to perform full-image segmentation on the target image; A segmentation map corresponding to the target image is displayed on a client interface, and the segmentation map comprises a plurality of selectable regions segmented out; In response to a region selection operation triggered for the multiple selectable regions, the target image is subjected to a matting process according to the at least one selected region, to obtain a matting image. The matting image is displayed on the client interface.

9. An image processing method, wherein, The method comprises: The image content of the product image is subjected to point sampling, to obtain sampling point information; The product image is subjected to full-image segmentation based on the sampling point information; The segmented image corresponding to the product image is displayed on the client interface, and the segmented image comprises multiple object regions; In response to a region selection operation triggered for the multiple object regions, a target object subject is separated from the product image according to the at least one selected object region, to obtain a display image of the target object subject.

10. An image processing system, wherein, The method comprises: A client acquires a target image to be subjected to matting; A server implements the method of any one of claims 1-9.

11. An electronic device, comprising: The method comprises: A memory and a processor; wherein The memory is configured to store a program; The processor is coupled to the memory and is configured to execute the program stored in the memory, to implement the steps of the image processing method of any one of claims 1-7, or to implement the steps of the image processing method of claim 8, or to implement the steps of the image processing method of claim 9.

12. A computer readable storage medium, wherein, The computer readable storage medium stores a computer program or instructions; the computer program or instructions are executed by the processor to implement the steps of the image processing method of any one of claims 1-7, or to implement the steps of the image processing method of claim 8, or to implement the steps of the image processing method of claim 9.

13. A computer program product, wherein, The computer program product comprises a computer program or instructions, and when the computer program or instructions are executed by the processor, the processor executes the steps of the image processing method of any one of claims 1-7, or the steps of the image processing method of claim 8, or the steps of the image processing method of claim 9.

Citation Information

Patent Citations

  • Image matting method

    CN101989353A

  • Interactive matting system, method and device based on superpixel segmentation

    CN110969629A

  • Image processing method and system, electronic equipment, storage medium and program product

    CN118918125A

  • Updating Image Segmentation Following User Input

    US20110216976A1

  • Method and apparatus for segmenting images

    US7031517B1