Image Cleaning on Mobile Devices
By using machine learning models to identify and generate masks on user equipment and perform image cleaning, the problem of high resource consumption in the prior art is solved, and efficient image cleaning effect on mobile devices is achieved.
Patent Information
- Application Number
- CN202111267747.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-04
- Filing Date
- 2021-10-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-10-28
AI Technical Summary
The prior art requires a large number of computer processor and memory resources when performing image cleaning on mobile devices, and it is difficult to efficiently implement on resource-constrained user devices.
By obtaining a first image depicting the foreground item on the user device, subsampling is performed to generate a first-level image of multi-scale transformation, and identifying the foreground and background parts based on the machine learning model, generating and enlarging the mask, and finally applying the mask on the user device to remove the image background, forming a cleaned image.
It realizes efficient image cleaning on user equipment with resource-constrained resources, reduces the burden on the back-end server system, and improves the efficiency of image search and publishing systems.
Smart Images

Figure CN114445286B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image cleaning on a mobile device. Background Art
[0002] Computer image processing may include automatically extracting meaningful information from an image (such as a digital image) using digital image processing techniques. Some examples of such image processing involve the automatic recognition of foreground objects and background objects in a photo or picture. The segmentation of foreground objects and background objects can be used as part of a process for cleaning an image before publishing or initiating a search (such as an image search). Summary of the Invention
[0003] Methods, systems, and articles of manufacture including computer program products for image cleaning are provided. In some embodiments, a method is provided. The method may include: obtaining, at a user device, a first image depicting at least one foreground item; subsampling the first image at the user device to a first-level image of a multiscale transform; performing, at the user device and based on a machine learning model, identification of a foreground portion and a background portion of the first-level image; generating, at the user device and based on the identification of the foreground portion and the background portion, a first mask; magnifying the first mask at the user device to a resolution corresponding to the first image depicting the foreground item; applying the magnified first mask to the first image at the user device to form a second image, the second image depicting the foreground item with at least a portion of the background portion in the second image removed; and / or providing, at the user device, the second image depicting the foreground item to a publishing system.
[0004] In some variations, one or more of the features disclosed herein that include the following features may optionally be included in any feasible combination. The magnification may further include: applying a domain transformation to magnify and smooth the first mask. A first pixel boundary surrounding at least a portion of the first-stage image may be specified to initialize the machine learning model to classify that portion as pure background. Based on the machine learning model and the specified first boundary as the pure background, the pixels of the first-stage image may be classified as a likely foreground portion or a likely background portion, wherein the first mask is generated based on the classified pixels. The machine learning model may include a Gaussian mixture model, GrabCut, and / or a neural network. The first pixel boundary may include a set amount of pixels surrounding at least a portion of the first-stage image. The set amount may be 5. The set amount may provide an initial state of the machine learning model to enable classification of one or more pixels of the first-stage image as pure background, likely background, pure foreground, or likely foreground. The subsampling may further include smoothing the first image to form the first-stage image. The application may include: causing a display of the second image; receiving an indication of a scribble on the second image; modifying the first mask based on the scribble or an enlarged scribble by at least adding or removing one or more pixels of the first mask; and applying the modified first mask to the second image. The first mask may be smoothed before application. The user device may obtain the first image by: capturing an image from an image sensor of the user device, downloading the first image, downloading the first image from a third-party website, and / or capturing the first image as a screenshot from a display of the user device. The foreground portion of the first-stage image and the background portion of the first-stage image may correspond to subsampled and smoothed versions of the background portion and the foreground portion of the first image. The second image may be caused to be displayed as a list of the publishing system.
[0005] Embodiments of the present subject matter may include methods consistent with the description provided herein and articles of manufacture including tangible, machine-readable media operably causing one or more machines (such as computers, etc.) to perform operations that implement one or more of the described features. Similarly, a computer system is also described as including one or more processors and one or more memories coupled to the one or more processors. The memories may include non-transitory computer-readable or machine-readable storage media that may include, encode, store one or more programs, etc., that cause the one or more processors to perform one or more of the operations described herein. Computer-implemented methods consistent with one or more embodiments of the present subject matter may be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computing systems may be connected via a direct connection, etc. between one or more of the multiple computing systems and may exchange data and / or commands or other instructions, etc. via one or more connections including, for example, a connection to a network (such as the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.).
[0006] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the specification, the drawings, and the claims. Although certain features of the presently disclosed subject matter are described for illustrative purposes in connection with the virtualization of configuration data, it will be readily understood that such features are not intended to be limiting. The claims following this disclosure are intended to define the scope of the protected subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The accompanying drawings, which are included and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed embodiments. In the drawings,
[0008] Figure 1 is a diagram of an example of a system suitable for image cleaning according to some example embodiments;
[0009] Figure 2 is a diagram showing an example of a user equipment according to some example embodiments;
[0010] Figure 3 depicts a block diagram of an image cleaning system configured to clean an image according to some example embodiments;
[0011] Figure 4 depicts an example of a process for image cleaning according to some example embodiments;
[0012] Figure 5Depicts another example of a process for image cleaning according to some example embodiments;
[0013] Figure 6 Depicts an example of a multi - scale method for subsampling according to some example embodiments;
[0014] Figure 7 Depicts an example of a process for image trimming according to some embodiments; and
[0015] Figure 8 Depicts another example of a process for image trimming according to some embodiments.
[0016] In fact, like reference numerals are used to refer to the same or similar items in the drawings. Detailed Description
[0017] Image cleaning can include removing the background portion of an image so that the image mainly includes one or more foreground items of interest. Accurately separating the foreground portion of the image from the background can improve the quality of the image and / or improve the ability of a system (such as an image search system, a publishing system, etc.) to automatically identify the foreground items in the image. By way of another example, if a user is seeking to publish an item on a publishing system at a mobile device, removing the background portion of the image can enable other users on other mobile devices to access the actual item being published without interference from the background. Similarly, image search can focus the image search on the foreground items of interest in the image rather than the background (which is not of interest). The removed background can be replaced with a plain background, such as a white or other color background, or a bokeh effect can be applied to further emphasize the foreground items. The cleaned image including the foreground items of interest (without the background) can be used, as mentioned, for enhanced image search, enhanced publishing, and / or enhanced insertion into other media.
[0018] While the image cleaning for removing background portions as mentioned may be desirable, this image cleaning may require a relatively large amount of computer processor and memory resources. Implementing the image cleaning (e.g., implementing the image cleaning as mentioned) on the user's mobile device rather than on the backend server system of the publishing system can distribute the burden from the backend server system to the mobile device, such as a smart phone, a tablet computer, or other processor-based devices (collectively referred to herein as user devices) associated with a user (e.g., an end user). Since the user device may be more limited in processing and resources than the backend server system, the image cleaning process may need to be adapted to operate using fewer memory and / or computer processor resources. To this end, one or more processes are disclosed herein that facilitate image cleaning on user devices where memory and / or processor resources may be more limited relative to the physical servers (e.g., image search systems, publishing systems, etc.) typically used in backend systems.
[0019] In some example embodiments, a user device may obtain a first image depicting at least one foreground item. The user device may then subsample the first image to a first-level image of a multiscale transform. The user device may automatically perform identification of the foreground portion and the background portion of the first-level image based on a machine learning model. Next, the user device may generate a first mask based on the identified portions. This first mask may then be upscaled to a resolution corresponding to the first image depicting the foreground item. The user device may then apply the upscaled first mask to the first image to form a second image. This second image depicts the foreground item with at least a portion (if not all) of the background portion removed using the first mask. The user device may then provide the second image depicting the foreground item to another system (e.g., a publishing system) to enable enhanced image search, enhanced publishing, and / or enhanced insertion into other media.
[0020] Figure 1 FIG. is an example of a system 100 suitable for image cleaning according to some example embodiments. The system 100 may include a server machine 110, a database 115, and one or more user devices 140A - 140B, all of which may be communicatively coupled to each other via a network 190.
[0021] The publishing system 105 may include one or more server machines (e.g., server machine 110) and may also include one or more databases (e.g., database 115). The publishing system may be a network-based publishing system configured to publish one or more items to other devices (e.g., other user devices). Although a publishing system is depicted, the publishing system may also be included in or include an image search system, a social media system, an e-commerce system, and / or other types of systems. In Figure 1In the example, the database 115 may store images of searchable items published to other devices. In some embodiments, the publishing system may list over a billion items, some of which may be stored in the database 115 as a list (although other quantities of items may also be implemented). For example, according to the embodiments disclosed herein, a user device (associated with a user) may capture an image, clean the captured image, and then send an image to the publishing system that mainly includes an item of interest in the foreground (e.g., to publish an image depicting the item to others, initiate a search for other similar or related items, etc.).
[0022] The server machine 110 may form all or part of the publishing system 105. For example, one or more cloud-based server systems may be used at 110 to provide one or more services to the user devices 140A - 140B. Both the server machine 110 and the user devices 140A - 140B may be implemented in a processor-based device described below. The server machine 110 may include a server, a web server, a database management system, or other machine capable of receiving and processing information (such as image data of a captured image depicting one or more items). The server machine 110 may be part of the illustrated publishing system, but the server machine may be part of an image search system, a social media system, a website, a database, and / or other types of systems. Figure 2 The depicted server 110, database 115, and / or user devices 140A - 140B may be implemented in a general-purpose computer modified (e.g., configured or programmed) by software to be a special-purpose computer to perform one or more of the functions described herein for the server machine 110, database 115, and / or user devices 140A - 140B. For example, a computer system can implement any one or more of the methods described herein. As used herein, a "database" is a data storage resource and may store data structured as text files, tables, spreadsheets, relational databases (such as object-relational databases), triple stores, in-memory databases, hierarchical data stores, or any suitable combination thereof. Additionally,
[0023] Figure 1 Any two or more of the illustrated devices may be combined into a single device, and the functions described herein for any single device may be subdivided among multiple devices. Figure 1
[0024] User devices 140A-140B may include or be included in the following processor-based devices, such as, for example, smart phones, desktop computers, vehicle computers, tablet computers, navigation devices, portable media devices, Internet of Things devices (such as sensors, actuators, etc.), wearable devices (such as smart watches, smart glasses), etc. For example, the user device may be accessed by a user to effectuate the capture of an image. For example, the user device may obtain a first image depicting a foreground item and a background. The user device may also subsample the first image to a first-level image of a multiscale transform. In addition, the user device may perform identification of a foreground portion of the first-level image and a background portion of the first-level image based on a machine learning (ML) model. The user device may generate a first mask based on the foreground portion and the background portion. Next, the user device may upscale the first mask to a resolution corresponding to the first image depicting the foreground item. The user device may then apply the upscaled first mask to the first image to form a second image depicting the foreground item and cause at least a portion of the background portion in the second image to be removed. The user device may then provide the second image depicting the foreground item to a publishing system. In some examples mentioned herein, a user may refer to a human user, a machine user (such as a computer configured to interact with the user device via a software program), or a combination thereof (such as a machine-assisted human or a human-supervised machine). However, the user is not considered part of system 100 but may access the user device or otherwise be associated therewith as mentioned.
[0025] Network 190 may be any network that enables communication between or among devices 140A-140B, 110, and 115. Thus, network 190 may be a wired network, a wireless network (such as a mobile or cellular network), or any suitable combination thereof. Network 190 may include one or more portions that constitute a private network, a public network (such as the Internet), or any suitable combination thereof. Thus, network 190 may include one or more portions that include a local area network (LAN), a wide area network (WAN), the Internet, a mobile phone network (such as a cellular network), a wired network (such as a plain old telephone system (POTS) network), a wireless data network (such as a WiFi network), or any suitable combination thereof.
[0026] Figure 2FIG. 0 is a diagram illustrating an example of a user equipment 200 that can be implemented to provide user equipments 140A - 140B according to some example embodiments. As mentioned, according to some example embodiments, the user equipment can be used to perform image cleaning. The user equipment can provide the cleaned images to other devices (such as the server machine 110 at the publishing system 105). The user equipment can be configured to perform one or more of the aspects discussed herein for the user equipment (such as user equipments 140A - 140B). Additionally, the user equipment can include an image cleaning system 300 as further described below. Figure 3 The image cleaning system 300 described further below.
[0027] The user equipment 200 can include one or more processors (such as processor 202). Processor 202 can be any of a variety of different commercially available processors (such as a central processing unit, a graphics processing unit, an image processing unit, a digital processing unit, a neural processing unit, an ARM - based processor, etc.). The user equipment can also include a memory 204, such as random access memory (RAM), flash memory, or other types of memory. Memory 204 can be accessed by processor 202. Memory 204 can be adapted to store an operating system (OS) 206 and application programs 208. Processor 202 can be directly or via a suitable intermediate hardware connection to a display 210 and one or more input / output (I / O) devices 212 (such as a keyboard, a touchpad sensor, a microphone, an image capture device 213, etc.). The image capture device 213 can form part of the Figure 3 image cleaning system 300 described below.
[0028] Processor 202 can be coupled to a transceiver 214 that interfaces with at least one antenna 216. Depending on the nature of the user equipment 200, transceiver 214 can be configured to wirelessly transmit and wirelessly receive cellular signals, WiFi signals, or other types of signals via antenna 216. Additionally, in some configurations, a GPS receiver 218 can also utilize antenna 216 (or another antenna) to receive GPS signals.
[0029] The user equipment 200 can include additional components or operate with fewer components than those described above. Additionally, the user equipment 200 can be implemented as a Figure 2Some or all of the cameras, Internet of Things devices (such as sensors or other smart devices) among the components described above. The user device 200 may be configured to perform any one or more aspects described herein. For example, the memory 204 of the user device 200 may include: instructions including one or more units for performing the methods discussed herein. The units may configure the processor 202 of the user device 200 or at least one processor of the user device 200 having multiple processors to perform one or more of the operations outlined below for each unit. In some embodiments, both the user device 200 and the server machine 110 may store at least a portion of the units discussed above and cooperate to perform the methods of the present disclosure, which will be explained in more detail below.
[0030] Figure 3 A block diagram of an image cleaning system 300 configured to clean images is depicted in accordance with some example embodiments. If not all of the image cleaning system is included in the user device 200, some of the image cleaning system may be included therein. Alternatively or additionally, if not all of the image cleaning system is included in the backend server system or service (such as a cloud-based service accessible by the user device 200), some of the image cleaning system may be included therein. The image cleaning system 300 may include or be coupled to an image capture device (such as the image capture device 213). For example, the image capture device may include an image sensor (such as a camera) coupled to or included in the user device, but the image capture device or the image cleaning system 300 may capture images in other ways (such as screen capture of a displayed image, obtained from a storage device containing images, downloaded from another device, etc.). The image cleaning system 300 may temporarily or permanently include a processor (such as an image processor communicating with the image capture device 213), a memory (such as the memory 204), and a communication component (such as the transceiver 214, the antenna 216, and the GPS receiver 218) in the form of one or more units for performing the methods described herein.
[0031] The image cleaning system 300 may include a receiver unit 302, an identification unit 304, a mask generation unit 306, a zooming unit 308, an image application unit 309, a painting unit 312, and a communication unit 310, all of which may be configured to communicate with each other (e.g., via a bus, shared memory, or switch). Although described as part of the mobile device 200, the image cleaning system 300 may be distributed across multiple devices such that some of the units associated with the image cleaning system 300 may reside in the user device 200 and some other units may reside in the server machine 110 or the publishing system 105. Whether the units are on a single device or distributed across multiple systems or devices, they may cooperate to perform any one or more of the methods described herein. Any one or more of the units described herein may be implemented using hardware (e.g., one or more processors of a machine) or a combination of hardware and software. For example, any one of the units described herein may configure a processor (a processor among one or more processors of a machine) to perform the operations described herein for that unit. Additionally, any two or more of these units may be combined into a single unit, and the functions described herein for a single unit may be subdivided into multiple units. Further, as described above and according to various example embodiments, units described herein as being implemented in a single device may be distributed across multiple devices. For example, as mentioned above for Figure 1 the server machine 110 may cooperate with the user device 140A or 140B via the network 190 to perform the methods described herein. In some embodiments, one or more of the units or portions of the units discussed above may be stored on both the server machine 110 and the mobile device 200.
[0032] The receiver unit 302 may obtain an image depicting at least one foreground item. For example, the receiver unit may receive an image (e.g., pixel data) depicting at least one foreground item of interest and a background. The image may be obtained from an image capture device 213 (e.g., an image sensor, such as a camera coupled to or included in the user device), a screen capture of a displayed image, obtained from a storage device containing the image, downloaded from another device, downloaded from a third-party website, etc. The image capture device 213 may include a hardware component including an image sensor (e.g., a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS)). The image sensor may detect image data to be captured, and the image capture device 213 may transfer the image data to the receiver unit 302. The receiver unit 302 may include a part of the image capture device 213, or the image capture device 213 may include a part of the receiver unit 302. The receiver unit 302 may communicate with the image capture device 213 and other units 304 - 309 via the communication unit 310.
[0033] The image cleaning system 300 may include the zooming unit 308 as mentioned. For example, the zooming unit may downsample the first image to a first-level image. Regarding Figure 6 Examples of the first-level image are further described at 608. In this example, the first-level image is downsampled to 1 / 16 of its original size. Alternatively or additionally, the zooming unit may also be used to upsample an image. For example, a downsampled mask or image may be upsampled to a larger pixel size. In Figure 6 the example of, for example, the downsampled image may be upsampled back to its original resolution shown at 602. In some example embodiments, domain transformation may be used for zooming in and / or out. Although various types of transformations may be used for zooming in and / or out of an image, in some embodiments, an edge-aware domain transformation may be used, such as "Domain Transform for Edge-Aware Image and Video Processing", E. Gastal and M. Oliveira, ACM SIGGRAPH 2011.
[0034] The image cleaning system 300 may include the recognition unit 304 as mentioned to recognize the foreground portion of the image from the background portion of the image. The recognition unit 304 may include a machine learning (ML) model algorithm or model for at least recognizing the foreground portion (and / or foreground portions) of the image, such as a neural network, a convolutional neural network (CNN), a generative adversarial network (GAN), the GrabCut algorithm, or other techniques. The GrabCut algorithm typically assumes that the user provides an initial bounding box surrounding the object. In contrast, the image cleaning system 300 may provide full automation according to some example embodiments by: initializing the GrabCut algorithm with a predefined (e.g., 5-pixel) wide boundary along the edges of the image as pure background, such that the remaining regions of the image are initialized as possible foreground.
[0035] The image cleaning system 300 may include a mask generator unit 306 as mentioned. The mask generator unit may generate a first mask, and the first mask may be applied to an image to mask the image so that the foreground portion of the image is retained. The image cleaning system 300 may include an application unit 309 as mentioned. The application unit may apply the mask to the image to form another image that mainly depicts the foreground portion. This another image may be provided to another system (such as the publishing system 105) via the communication unit 310. The image cleaning system 300 may also include a painting unit 312 as mentioned to handle painting for modifying the mask. For example, the GrabCut algorithm may allow a user device to paint on a part of the image or the mask to indicate that the painted part is misclassified as foreground (or background). In some instances, the GrabCut algorithm (or other cleaning algorithms) may then be run to modify the image based on the painting. When this is the case, the painted part may be removed as part of the background (or added in the case of misclassified background).
[0036] The communication unit 310 implements Figure 3 communications between the depicted units and / or implements communications with other devices and / or systems. For example, the communication unit 350 may include communication mechanisms such as an antenna, a transmitter, one or more buses, and other suitable communication mechanisms capable of implementing communications between the unit, user devices, the publishing system, and / or other devices or systems.
[0037] Figure 4 An example of a process 400 for image cleaning according to some example embodiments is depicted. Figure 4 The description also refers to Figure 1 、 Figure 5 and Figure 6 .
[0038] At 405, a first image depicting a foreground item may be obtained by a user device. For example, the first image (which may include one or more items of interest in the foreground) may be captured by a user device (such as user device 140A). The first image may be captured by a user device using an image sensor (such as a camera coupled to or included in the user device). Alternatively, the image may be captured via screen capture of a displayed image, obtained from a storage device containing the image, downloaded from another device, downloaded from a third - party website, or obtained in other ways. Referring to Figure 5 , an example of a first image 502 depicting a foreground item 503 (which is a camera in this example) and a background is depicted. In this example, the first image may have a resolution of 1600 pixels by 1066 pixels, but images of other resolutions may also be used. Each pixel may represent data, such as pixel data.
[0039] To further illustrate by way of example, a user device may capture a first image including a foreground portion of the image and a background portion of the image. The foreground portion may include foreground items of interest. To be able to enhance posting or image search, the user device may remove some of the background portion of the image if it does not remove all of the background portion of the image, so that the foreground items are mainly depicted in the image. This focusing on the foreground portion of the image (which depicts items of interest) may enable other users on other devices to access the actual item being posted without interference from the background. This focusing may also enhance image search by focusing on the items of interest in the image rather than the items depicted in the background (which are not of interest).
[0040] At 412, the first image may be subsampled to a first-level image of a multiscale transform. For example, the user device may subsample the first image 503 from a resolution of 1600 pixels by 1066 pixels to a resolution of 500 pixels by 333 pixels. Figure 6 An example of a multiscale method for performing subsampling is depicted. In Figure 6 the example, the first image is represented by image 602 and is subsampled to a first level of approximately 1 / 16, so that image 610 has a resolution of 1 / 16 of image 602. In the multiscale method, each level undergoes smoothing (or blurring) and subsampling. For example, smoothing may be performed using a smoothing function (such as Gaussian averaging). When Gaussian averaging or blurring is used, each pixel may include a local average that corresponds to Figure 6 adjacent pixels on another level (such as a lower level or a higher level) of the multiscale pyramid depicted. In some embodiments, the multiscale method includes first subsampling the image and applying GrabCut only to the lowest resolution. After this initial step, a domain transform (DT) filter may be used to upsample the image subsequently while only optimizing the image near the boundary region. This may provide processing that is 3 times faster than traditional GrabCut. Although the previous example describes applying GrabCut at the lowest (1 / 16) resolution of the multiscale method, GrabCut may also be applied at other resolution levels.
[0041] At 414, the foreground portion of the first-level image and the background portion of the first-level image are identified based on a machine learning model. For example, the user device may include a machine learning model to at least identify the foreground portion of the image. The machine learning model may include neural networks, convolutional neural networks (CNNs), generative adversarial networks (GANs), GrabCut algorithms, and / or other techniques. The machine learning model may include aspects of a multiscale GrabCut algorithm modified according to some example embodiments to Figure 6The GrabCut algorithm is performed on one level of the depicted multi-scale pyramid (e.g., the subsampled image 608 at the lowest resolution). Additionally, the multi-scale GrabCut algorithm can be automatically initialized with a predetermined bounding box (e.g., a 5-pixel-wide border surrounding the image) according to some example embodiments, although the width of the bounding box can also have other predetermined values.
[0042] In Figure 6 the example of, each level of the multi-scale is graphically depicted. As shown at 604 - 610, the image 602 is subsampled at resolutions of 1 / 2, 1 / 4, 1 / 8, and 1 / 16. Different from traditional multi-scale GrabCut algorithms, the GrabCut algorithm can be applied only at a given resolution level (e.g., the lowest resolution of the subsampled image 610) according to some example embodiments. In the case of GrabCut, image segmentation can be iteratively performed using the image 608 based on a graph that classifies background pixels and foreground pixels and then a cut through the graph to identify the foreground portion. Although a user accessing the user device can select an initial bounding box surrounding the foreground item depicted in the image on the display to initialize GrabCut, according to some example embodiments, the predetermined bounding box can be automatically determined by the user device. For example, the predetermined bounding box can be automatically determined as a 5-pixel-wide border surrounding the first image, although the width of the bounding box can also have other predetermined values. In the GrabCut algorithm, the bounding box pixels are initialized by GrabCut as "pure background". Then, GrabCut continues to determine whether the pixels of the first image are likely background or likely foreground. In this way, the foreground portion and the background portion of the image can be classified as likely foreground pixels or likely foreground pixels.
[0043] In some implementations, the GrabCut iteration is such that the initial classification of foreground or background can be corrected by the user painting (e.g., marking) a portion of the classified first image as misclassified, in which case GrabCut can be run again to provide another estimate of which pixels are likely foreground pixels or likely foreground pixels.
[0044] At 416, a first mask is generated based on the foreground portion and the background portion identified for the subsampled image. Figure 5 An example of a first image 502 that is subsampled by the multi-scale 505 pyramid algorithm to form a subsampled image 510 is depicted. Then the subsampled image 510 is provided to an ML model (e.g., GrabCut 515). Based on the subsampled image 510, GrabCut is able to provide a classification of which pixels of the subsampled image are likely foreground pixels and which pixels are likely background pixels. The estimate is then used to generate a first mask 520.
[0045] At 418, the first mask can be magnified to a resolution corresponding to the first image depicting the foreground item. Referring again to Figure 5 , the first mask 520 can be magnified 525 to form a magnified first mask 530 having the same or a similar resolution as the original image 502. For example, the first mask 520 can have a resolution of 500 by 333 pixels and then be magnified to a resolution of 1600 pixels by 1066 pixels to match the resolution of the original first image 502. The magnification can be performed in a variety of ways to change the size of the image to the resolution of the first image. For example, interpolation (such as nearest neighbor interpolation) can be used to change the size of the image. In some example embodiments, an edge-aware domain transform can be used for magnification. "Domain Transform for Edge-Aware Image and Video Processing", E. Gastal and M. Oliveira, ACM SIGGRAPH 2011 provides an example of an edge-aware domain transform, but other types of transforms can also be used to upsample or downsample an image. For example, a domain transform (DT) filter can be used to upsample (or downsample) an image while only optimizing the image near the boundary regions.
[0046] Next, the magnified first mask can be applied to the first image at 420 to form a second image depicting the foreground item, removing at least some of the background portion of the second image if not all of it. Referring again to Figure 5 , the magnified first mask 530 is applied to the first image 502 to form a second image 530. This second image 530 masks off (e.g., removes) some of the background portion if not all of the background portion from the first image 502 so that the foreground item 503 remains in the second image 530.
[0047] At 422, the second image depicting the foreground item can be provided to a system (such as a publishing system). Referring again to Figure 5 , the second image 530 can be provided to the publishing system 105 or other types of systems or devices. For example, the second image can be displayed as a list in the publishing system so that multiple user devices can view the list. Removing the background to leave the foreground item (as shown in the second image 530) can improve the ability of the publishing system to identify similar items or publish the image (e.g., so that other user devices can access the foreground item). Alternatively or additionally, a user device can provide the second image depicting the foreground item to an image search system for searching for items similar or related to the foreground item.
[0048] As mentioned, process 400 can be used to clean an image. In some instances, automatic image cleaning may require what is referred to as "trimming." Figure 7 Illustrated in Figure 7 is an example process 700 for image trimming according to some embodiments. In some embodiments, process 700 can be used to modify a first mask 710A automatically generated by process 400, but process 700 can also be used without using process 400. Figure 7 The description of also refers to Figure 5 .
[0049] At 705, a first mask 710B can be received. For example, the first mask 710B can correspond to a first mask automatically generated by process 400 without user intervention. The first mask 710B can then be presented on a display of a user device to enable painting on the first mask 710B at 715. For example, the first mask 710B can be presented on a display of a user device to enable trimming by painting to indicate which portion of the mask is to be removed from (or added to) the first mask. The painted mask is depicted at 710C. In this example, the painting indicates which portion needs to be excluded from the first mask 710B and thus from the corresponding foreground image 710A. In some embodiments, alternatively, the painting can be applied to the first image 710A displayed on the user device. At 720, the painting can be enlarged (e.g., expanded). For example, the painting at 710C can be expanded to be as large as the expanded painting shown at 710D.
[0050] At 730, the first image 710A is processed to generate a filtered scribble mask. For example, a domain transform (DT) filter, such as "Domain Transform for Edge-Aware Image and Video Processing" by E. Gastal and M. Oliveira, ACM SIGGRAPH 2011, can be used to filter the image 710A while only refining the image near the boundary regions. The filtered image is then binarized and thresholded to form a binarized and thresholded image that can be used as a mask. Then, the expanded scribble mask 710D is applied to the binarized and thresholded mask to form a filtered scribble mask 710E. The filtered scribble mask 710E can also be processed at 750 using contour smoothing. The filtered scribble mask 710E (which can be smoothed at 750) can be combined at 760 with the original filter mask 705 / 710B to form a new updated filter mask 710F. As can be seen, the new updated filter mask 710F can now exclude the part 766A associated with the scribble from the foreground. When the new updated filter mask 710F is applied to the original image 710A to form an updated image, as shown at the new updated image 710G, the updated image 710G can exclude the part 766B. If it is desired to further clean the image 710G (in this case, the new updated image 710G will be input at 730 and the scribbling will occur on the updated mask 710F), then the process at 700 can be repeated. In some embodiments, the DT filter applied at 730 is applied to the subsampled image 710A and scribble mask 710D at a scale of 1 / 16 (e.g., from 1600 pixels by 1200 pixels to 400 pixels by 300 pixels) of, for example, the pyramid multiscale method as mentioned above. Figure 6 The subsampled image 710A and scribble mask 710D are at a scale of 1 / 16 (e.g., from 1600 pixels by 1200 pixels to 400 pixels by 300 pixels) of, for example, the pyramid multiscale method as mentioned above.
[0051] Figure 8 FIG. depicts an example of a process 800 for image cleaning according to some example embodiments. Figure 8 The description also refers to Figure 7 .
[0052] At 810, a first image for image cleaning and / or its corresponding first image can be received. For example, the first mask 710B and / or the first image 710 can be received by an image cleaning system of a user device. At 812, the first mask and / or the corresponding first image can be displayed on the user device. This display enables a user associated with the user device to draw marks (or provide other forms of drawing) on the display. The drawing indicates which part of the mask or image needs to be cleaned. For example, if the mask classifies or identifies a part that should be recognized as the background as the foreground, the drawing (such as the drawing shown at 710C) can indicate the part (such as pixels) that needs to be cleaned by removing the misclassified part (but in some cases, misclassified parts can also be added). In response to the drawing, the image cleaning system can receive an indication of the drawing at 814, such as which pixels are indicated as being drawn. At 816, the drawing can be enlarged. For example, the enlargement can expand the size of the drawing by a predetermined number of pixels (such as expanding the size of the pixels by a predetermined percentage of the area of the drawing, such as increasing by 1%, 2%, 3%, 4%, 5% or other percentages).
[0053] A filtered drawing mask is generated at 818. An image (such as image 710A) can be binarized and thresholded to form a binarized and thresholded image that can be used as a mask. The enlarged drawing mask 710D is applied to the binarized and thresholded mask to form a filtered drawing mask 710E. The image 710A and the drawing mask 710D can be subsampled to enable binarization, thresholding, and the application of the drawing mask. After that, DT can be applied to upsample the filtered drawing mask 710E to have the same or similar resolution as the original image 710A. At 822, the filtered drawing mask 710E can be combined with the original filter mask 710B to form a new updated filter mask. At 824, the new updated filter mask 710F can then be applied to the original image 710A to generate an updated image 710G, which in this example excludes the part corresponding to the drawing.
[0054] In some example embodiments, generation of a mask (e.g., mask 530) may be performed using a neural network (e.g., a convolutional neural network including pooling). Examples of this type of neural network may be found in "A Simple Pooling-Based Design for Real-Time Salient Object Detection", Jiang-Jiang Liu et al., CVPR 2019. For example, supervised learning may be used to train a neural network (e.g., using training images labeled to distinguish foreground pixels from background pixels) to automatically segment foreground and background portions of an image. Alternatively, unsupervised learning, such as by using a generative adversarial network, may be used to generate a mask, such as mask 530. In some example embodiments, a saliency map may be used to further modify the binary mask.
[0055] In some example embodiments, the GrabCut algorithm automatically initializes with a predefined (e.g., 5 pixel wide) border along the edge of the image as pure background, while the rest of the image is initialized as possible foreground. GrabCut is then run to identify pixels as foreground pixels or background pixels. For example, a predefined border (e.g., a rectangle) is set around the image. Boundary pixels are classified as pure background, while the interior of the border is unknown. The GrabCut algorithm then performs an initial hard labeling of each pixel, such as pure background based on the border or any user-provided indication (e.g., a paint indicating foreground pixels or background pixels). Next, a Gaussian mixture model (GMM) models the foreground and background. The GMM learns and creates a pixel distribution to label unknown pixels as possible foreground or possible background based on the relationship between the unknown pixel and other hard-labeled pixels (e.g., classification or clustering based on color statistics). GrabCut then generates a graph based on this pixel distribution, with the nodes in the graph being the pixels of the image. The graph includes two nodes / additional nodes called source nodes and sink nodes. Each foreground pixel is connected to the source node and each background pixel is connected to the sink node. The weight of the edge connecting a pixel to a source node is defined by the probability that the pixel is a foreground pixel, while the weight of the edge connecting a pixel to a sink node is defined by the probability that the pixel is a background pixel. The weights (between pixels) are defined by edge information or pixel similarity. If there is a large difference in pixel color, the edge between pixels will have a correspondingly low weight. The mincut algorithm is then used to segment the graph. The minicut algorithm will Figure 1Split into two to segment the source node and the sink node based on a cost function (e.g., a minimum cost function). For example, the cost function can correspond to the sum of all weights of the edges being segmented. After segmentation, all pixels connected to the source node are classified as foreground, while pixels connected to the sink node are classified as background. The GrabCut process can be iterated until the classification of pixels converges.
[0056] In some example embodiments, the multi-scale method is configured to subsample the image and perform GrabCut only at the lowest resolution. In some example embodiments, upsampling is performed using a domain transform (DT) filter on the image near the optimized boundary region.
[0057] One or more aspects or features of the subject matter described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs, field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These different aspects or features can include implementations employing one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. The relationship of client and server arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0058] These computer programs (which may also be referred to as programs, software, software applications, applications, components, or code) include machine instructions for a programmable processor and can be implemented in a high-level procedural and / or object-oriented programming language and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as in a non-transitory solid-state memory or a magnetic hard disk drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transitory manner, such as in a processor cache or other random access memory associated with one or more physical processor cores.
[0059] To provide for interaction with a user, one or more aspects or features of the subject matter described in this document can be implemented on a computer having a display device (such as a cathode ray tube (CRT), liquid crystal display (LCD), or light emitting diode (LED) monitor) for displaying information to the user and a keyboard and an indicating device (such as a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, speech recognition hardware and software, optical scanners, optical pointers, digital image capture devices, and associated interpretation software, etc.
[0060] The subject matter described in this document can be embodied in a system, apparatus, method, and / or article of manufacture according to desired configurations. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described in this document. On the contrary, they are only some examples consistent with aspects related to the subject matter described. Although some variations have been described in detail above, other modifications or additions are possible. Specifically, additional features and / or variations can be provided in addition to those features and / or variations set forth herein. For example, the above embodiments can be directed to various combinations and sub-combinations of the disclosed features and / or combinations and sub-combinations of several additional features described above. Moreover, the logical flows depicted in the figures and / or described herein do not necessarily need to be in the particular order shown or sequential order to achieve the desired result. For example, the logical flow can include different and / or additional operations than those shown without departing from the scope of the disclosure. One or more operations of the logical flow can be repeated and / or omitted without departing from the scope of the disclosure. Other embodiments are within the scope of the appended claims.
Claims
1. A method for image cleaning, comprising: obtaining, at a user device, a first image depicting at least a foreground item; subsampling the first image at the user device to a first-level image of a multiscale transform; performing, at the user device and based on a machine learning model, identification of a foreground portion and a background portion of the first-level image; generating, at the user device and based on the identification of the foreground portion and the background portion, a first mask; causing display of the first mask; receiving a painting instruction on the first mask to obtain a painting mask indicating which part of the first mask to remove or add from the first mask; enlarging the painting mask; applying a domain transform filter to the first image to refine the image around a boundary region; performing binarization and thresholding on the filtered image to form a binarized and thresholded image; applying the enlarged painting mask to the binarized and thresholded image to form a filtered painting mask; combining the first mask with the filtered painting mask to form an updated mask; applying the updated mask to the first image at the user device to form a second image, the second image depicting the foreground item with at least a part of the background portion in the second image removed; and providing, at the user device, the second image depicting the foreground item to a publishing system.
2. The method according to claim 1, wherein, the performing includes: specifying a first pixel boundary surrounding at least a part of the first-level image to initialize the machine learning model to classify this part as pure background; and classifying pixels of the first-level image as a possible foreground part or a possible background part based on the machine learning model and the first boundary specified as the pure background, and wherein the first mask is generated based on the classified pixels.
3. The method according to claim 1, wherein, the machine learning model includes GrabCut.
4. The method according to claim 2, wherein, the first pixel boundary includes a set amount of pixels surrounding at least a part of the first-level image.
5. The method according to claim 4, wherein, the set amount is 5.
6. The method according to claim 4, wherein, the set amount provides an initial state of the machine learning model to enable classification of one or more pixels of the first-level image as pure background, possible background, pure foreground, or possible foreground.
7. The method according to claim 1, wherein, the subsampling further includes smoothing the first image to form the first-level image.
8. The method according to claim 1, wherein, the first mask is smoothed before application.
9. The method according to claim 1, wherein, the user device obtains the first image by: capturing an image from an image sensor of the user device, downloading the first image, downloading the first image from a third-party website, and / or capturing the first image as a screenshot from a display of the user device.
10. The method according to claim 1, wherein, The foreground part of the first - level image and the background part of the first - level image both correspond to the subsampled and smoothed versions of the background part and the foreground part of the first image.
11. The method according to claim 1, further comprising: causing the second image depicting the foreground item to be displayed at the user device as a list of items of the publishing system.
12. The method according to claim 1, wherein, the machine - learning model includes a Gaussian mixture model and / or a neural network.
13. An image - cleaning system, comprising: at least one processor; and at least one memory including instructions that, when executed by the at least one processor, cause the system to provide operations including the following: obtaining, at a user device, a first image depicting at least a foreground item; subsampling the first image at the user device to a first - level image of a multiscale transform; performing, at the user device and based on a machine - learning model, identification of a foreground part of the first - level image and a background part of the first - level image; generating, at the user device and based on the identification of the foreground part and the background part, a first mask; causing the display of the first mask; receiving a painting indication on the first mask to obtain a painted mask indicating which part of the first mask is to be removed or added from the first mask; enlarging the painted mask; applying a domain - transform filter to the first image to refine the image around the boundary region; performing binarization and thresholding on the filtered image to form a binarized and thresholded image; applying the enlarged painted mask to the binarized and thresholded image to form a filtered painted mask; combining the first mask with the filtered painted mask to form an updated mask; applying the updated mask to the first image at the user device to form a second image that depicts the foreground item with at least a part of the background part in the second image removed; and providing, at the user device, the second image depicting the foreground item to a publishing system.
14. The system according to claim 13, wherein, the performing includes: specifying a first pixel boundary surrounding at least a part of the first - level image to initialize the machine - learning model to classify that part as pure background; and classifying the pixels of the first - level image as a likely foreground part or a likely background part based on the machine - learning model and the first boundary specified as the pure background, and wherein the first mask is generated based on the classified pixels.
15. The system according to claim 13, wherein, the machine - learning model includes GrabCut.
16. The system according to claim 14, wherein, the first pixel boundary includes a set amount of pixels surrounding at least a part of the first - level image.
17. The system according to claim 16, wherein, the set amount provides an initial state of the machine - learning model to enable classification of one or more pixels of the first - level image as pure background, likely background, pure foreground, or likely foreground.
18. A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor, cause operations including the following: Obtain a first image depicting at least a foreground item at a user device; Subsample the first image at the user device to a first-level image of a multiscale transform; Perform identification of a foreground portion and a background portion of the first-level image at the user device and based on a machine learning model; Generate a first mask at the user device and based on the identification of the foreground portion and the background portion; Cause display of the first mask; Receive a painting indication on the first mask to obtain a painted mask indicating which part of the first mask to remove or add from the first mask; Enlarge the painted mask; Apply a domain transform filter to the first image to refine the image around a boundary region; Binarize and threshold the filtered image to form a binarized and thresholded image; Apply the enlarged painted mask to the binarized and thresholded image to form a filtered painted mask; Combine the first mask with the filtered painted mask to form an updated mask; Apply the updated mask to the first image at the user device to form a second image that depicts the foreground item with at least a portion of the background portion in the second image removed; And Provide the second image depicting the foreground item to a publishing system at the user device.