Integrated Interactive Image Segmentation

Through integrated segmentation system and convolutional RNN technology, combining multiple segmentation methods to generate the optimal segmentation mask, the problems of inaccurate image segmentation and high computational cost in the prior art are solved, and efficient and accurate image editing is achieved.

CN113554661BActive Publication Date: 2025-08-01ADOBE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110070052.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-03
Filing Date
2021-01-19
Publication Date
2025-08-01
Estimated Expiration
2041-01-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively combine multiple segmentation methods in the image segmentation process, resulting in inaccurate segmentation masks and lack of scalability, especially when retraining the neural network when replacing the segmentation tool, which has high computing cost and low efficiency.

Method used

The integrated segmentation system is adopted, and the convolutional RNN combined with the probability distribution map is used to generate the optimal segmentation mask through iterative integration of multiple segmentation methods, including deep learning technology, color range detection, etc., to achieve optimal segmentation of the image.

Benefits of technology

Improve the accuracy and efficiency of image segmentation, reduce computing costs, and achieve seamless integration and scalability of multiple segmentation methods, so that users can intuitively and interactively edit images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113554661B_ABST
    Figure CN113554661B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to integrated interactive image segmentation. Optimal segmentation methods and systems based on multiple segmented images are provided. In particular, multiple segmentation methods can be combined by considering previous segmentations. For example, an optimal segmentation can be generated by iteratively integrating a previous segmentation (e.g., using an image segmentation method) with a current segmentation (e.g., using the same or a different image segmentation method). To allow for optimal segmentation of an image based on multiple segmentations, one or more neural networks can be used. For example, when transitioning from one segmentation method to the next, a convolutional RNN can be used to maintain information about one or more previous segmentations. The convolutional RNN can combine the (multiple) previous segmentations with the current segmentation without any information about the image segmentation method used to generate the segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to image segmentation, and more particularly to integrated interactive image segmentation. Background Art

[0002] For example, images are often segmented to allow for editing of the images. Image segmentation is generally the process of generating image segments. Such segments can be visualized in an editing environment using a mask that indicates the segmented portions of the image to which the editing will be applied, as opposed to the portions of the image that will not be affected by the editing. The segments can be created along boundaries within the image such that the segments represent objects and / or features within the image. For example, in a portrait of an individual, the image can be segmented into segments of the individual's face or segments of the background, or if more detail is desired, the image can be segmented into segments of the individual's eyes, segments of the individual's teeth, segments of the individual's hair, etc. Thus, using a mask for such a portrait can indicate that the editing will be applied only to the emphasized portion(s) of the image, as opposed to the non-emphasized portion(s). Summary of the Invention

[0003] Embodiments of the present disclosure relate to an integrated segmentation system that allows for optimal segmentation based on multiple segmented images. According to embodiments of the present disclosure, the integrated segmentation system allows for the combination of multiple segmentation methods (e.g., different segmentation techniques implemented using various segmentation tools) by considering previous segmentations. For example, an optimal segmentation can be generated by iteratively integrating a previous segmentation (e.g., using an image segmentation method) with a current segmentation (e.g., using the same or a different image segmentation method). To create such an integrated segmentation system, one or more neural networks can be used. For example, when transitioning from one segmentation method to the next, the integrated segmentation system can implement a convolutional RNN to maintain information about one or more previous segmentations. The convolutional RNN can combine the previous segmentation(s) with the current segmentation without any information about the (multiple) image segmentation methods used to generate the segmentation.

[0004] A convolutional RNN can integrate a segmentation method that uses information about a probability distribution map related to image segmentation. For example, a convolutional RNN can be used to determine a feature map for an image. Then, this feature map can be combined with a current probability distribution map (e.g., from a current segmentation method) and a previous probability distribution map (e.g., from a previous segmentation method). In particular, the feature map can be combined with the current probability distribution map to generate a first feature (e.g., the combination of the feature map and the current probability distribution map). Additionally, the feature map can be combined with the previous probability distribution map to generate a second feature (e.g., the combination of the feature map and the previous probability distribution map). Then, the first feature and the second feature can be concatenated to generate an updated probability distribution map. This updated probability distribution map incorporates information about the current segmentation method and the previous segmentation method. This process can be repeated until the updated segmentation mask is the optimal segmentation mask. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1A Depicts an example configuration of an operating environment according to various embodiments of the present disclosure, in which some implementations of the present disclosure may be employed.

[0006] Figure 1B Depicts an example configuration of an operating environment according to various embodiments of the present disclosure, in which some implementations of the present disclosure may be employed.

[0007] Figure 2 Depicts aspects of an illustrative integrated segmentation system according to various embodiments of the present disclosure.

[0008] Figure 3 Illustrates a processing flow showing an embodiment for performing integration of multiple segmentation methods according to an embodiment of the present disclosure.

[0009] Figure 4 Illustrates a processing flow showing an embodiment for integrating multiple segmentation methods according to an embodiment of the present disclosure.

[0010] Figure 5 Illustrates a processing flow showing an embodiment for performing integration of multiple segmentation methods to generate an optimal segmentation mask according to an embodiment of the present disclosure.

[0011] Figure 6 Illustrates a processing flow showing an embodiment for integrating multiple segmentation methods including a removal segmentation method according to an embodiment of the present disclosure.

[0012] Figure 7 Illustrates an example environment according to an embodiment of the present disclosure, which can be used to integrate information about a (previous) segmentation when performing a current segmentation.

[0013] Figure 8 FIG. illustrates an example environment in accordance with an embodiment of the present disclosure, which can be used to integrate information about (a) previous segmentation(s) when performing a current segmentation that includes a removal action.

[0014] Figure 9 FIG. illustrates an example environment in accordance with an embodiment of the present disclosure, which can be used for joint embedding supervision of an integrated segmentation system that allows for optimal segmentation of an image based on multiple segmentations.

[0015] Figure 10 is a block diagram of an example computing device in which embodiments of the present disclosure may be employed. DETAILED DESCRIPTION

[0016] To meet statutory requirements, the subject matter of the present disclosure is described in detail herein. However, the specification itself is not intended to limit the scope of this patent. Rather, the inventors have contemplated that the claimed subject matter might also be implemented in other ways, in combination with other current or future technologies, to include different steps or combinations of steps similar to those described herein. Additionally, although the terms "step" and / or "block" may be used herein to denote different elements of methods employed, these terms should not be construed as implying any particular order among the steps disclosed herein unless and except when the order of individual steps is explicitly described.

[0017] Generally, users expect to edit images intuitively. For example, users may expect an edit that does not require expertise in difficult tools within an editing system, or an edit that is not highly repetitive and time-consuming. One way to edit an image is to use segmentation. Image segmentation can specify segments within the image so that the segments can be edited. A segmentation mask can be used to visualize such segments, which indicates the segmented part of the image to which the edit will be applied, rather than the part of the image that will not be affected by the edit. However, there may be difficulties in segmenting an image such that the generated segmentation mask is about the exact part of the image that the user expects to segment. For example, in an image of a dog in a field, a first segmentation tool (e.g., using a first segmentation method) may not include the dog's ears in the generated segmentation mask. In an attempt to modify the generated segmentation mask, a second segmentation tool may be selected. However, in most conventional systems, based on the selection of the second segmentation tool (e.g., using a second segmentation method), the initially generated segmentation mask is typically deleted and an entirely new segmentation mask is generated (the dog's ears may still not be included in the generated segmentation mask). Thus, it is difficult to combine various segmentation tools when attempting to segment an image. In an attempt to overcome the difficulty of combining various segmentation tools, segmentation masks generated using different segmentation tools have been combined by averaging the two segmentation masks. However, this may result in the combined segmentation mask still containing inaccuracies (e.g., the dog's ears are not included in the combined segmentation mask).

[0018] As technology has advanced, various methods for image segmentation have been developed to attempt to segment images more easily and accurately. For example, various (e.g., neural network-based) deep learning techniques for quickly and intelligently segmenting images have been implemented. However, even when using these deep learning-based techniques to generate a segmentation for an image, it is difficult to combine multiple segmentation tools when segmenting the image. As previously described, generally, if a new segmentation tool is applied to an image, any segmentation done using a previous segmentation tool will be lost. One way that has been attempted to combine segmentation tools using various deep learning techniques requires a complete re-training of the neural network related to the segmentation tools each time a new segmentation tool is added to combine these tools. Re-training the neural network to account for previous segmentations generated by other tools (e.g., using different methods) is computationally expensive and limits the scalability of the system. In this way, existing technologies are generally deficient in allowing for the combination of various segmentation tools in a scalable and computationally efficient manner that takes into account previous segmentations of the image.

[0019] Accordingly, embodiments of the present disclosure relate to an integrated segmentation system that allows for optimal segmentation based on multiple segmented images. In particular, the integrated segmentation system allows for the combination (e.g., performed using various segmentation tools) of multiple segmentation methods by considering previous segmentations. Such segmentation methods may include segmentation based on deep learning techniques, color range or saliency detection, thresholding, clustering methods, compression-based methods, histogram-based methods, edge detection, double clustering methods, region growing methods, partial differential equation-based methods, variational methods, graph partitioning methods, watershed transform, model-based segmentation, multi-scale segmentation, and semi-automatic segmentation.

[0020] Combining multiple segmentation methods allows the user to easily obtain an optimal segmentation of the image. For example, an optimal segmentation can be generated by iteratively integrating a previous segmentation (e.g., using an image segmentation method) with a current segmentation (e.g., using the same or a different image segmentation method). As an example, as the user interacts with the image, the (multiple) previous segmentation methods that have been applied to the image can be integrated into an updated image segmentation. In this way, using the integrated segmentation system to segment an image allows for the incorporation of multiple interactive image segmentation methods into an overall integrated image segmentation process. Advantageously, when generating a segmentation mask for optimal image segmentation based on an image, combining various image segmentation methods in this way leverages the strengths of each of the various image segmentation methods.

[0021] More specifically, embodiments of the present disclosure relate to an integrated segmentation system that is user-friendly and compatible with any interactive image segmentation method. In particular, as described herein, the integrated segmentation system can integrate various interactive image segmentation methods into a unified library. This unified library allows various image segmentation methods to build on each other in response to interactions (e.g., clicks, swipes, bounding boxes) indicating a desired segmentation for the image. For example, the user can interact with the image using a first image segmentation method to indicate that a dog should be included in the segmentation, and the first image segmentation method can generate a first segmentation mask that does not include the dog's ears. Using a second image segmentation method, the user can indicate that the dog's ears should be included in the segmentation. The integrated segmentation system of the present disclosure allows the first segmentation mask to be considered when segmenting the image using the second image segmentation method to generate an optimal segmentation mask (e.g., including the dog and the dog's ears). Advantageously, the unified library of the integrated segmentation system allows the user to interact with the image using any type of interactive input that best indicates the user's desired segmentation for the image. In this way, the integrated segmentation system allows the user to interact with the image in an intuitive and straightforward manner to obtain an optimal segmentation mask.

[0022] To integrate multiple segmentation methods into an optimal segmentation of an image, information about previous segmentations can be integrated into the current segmentation. In one embodiment, one or more neural networks can be used to implement the integrated segmentation system. A neural network generally refers to a computational method using a large cluster of connected neurons. For example, a neural network can consist of fully connected layers. Neural networks are self-learning and trained rather than being explicitly programmed, such that the generated output of the neural network reflects the desired result. In various embodiments, the integrated segmentation system can include one or more neural networks based on a convolutional recurrent neural network (RNN) architecture. For example, when transitioning from one segmentation method to the next, the integrated segmentation system can implement a convolutional RNN to maintain information about one or more previous segmentations.

[0023] When performing the current segmentation, the integrated segmentation system can use a convolutional RNN to integrate information about the previous segmentation(s). The convolutional RNN can combine the previous segmentation(s) with the current segmentation without any information about the image segmentation method(s) used to generate the segmentation. In this way, since any image segmentation method can be used to generate the segmentation, the integrated segmentation system is highly scalable. For example, even if a convolutional RNN was not previously used to integrate image segmentation methods, a new image segmentation method can be added to the integrated segmentation system.

[0024] The integration of combining the previous segmentation(s) with the current segmentation can be performed based on a convolutional RNN that receives and combines information about the previous segmentation of the image and information about the current segmentation of the image. Such information can at least include a probability distribution map. The probability distribution map can generally be the output of a segmentation method (e.g., a segmentation mask, a heatmap, etc.). A segmentation mask can be generated based on the probability distribution map. In particular, the convolutional RNN can receive information about the previous probability distribution map. Using this information, the convolutional RNN can generate a hidden state based on the information about the previous probability distribution map. Subsequently, when a subsequent segmentation method is used to segment the image, the hidden state can be used to update the current probability distribution map (e.g., determined using the subsequent segmentation method) to generate an updated probability distribution map. In this way, the hidden state of the convolutional RNN can be used to incorporate information about the previous segmentation into the current segmentation. Then, the resulting updated probability distribution map can be used to generate an updated segmentation mask (e.g., combining the previous segmentation and the current segmentation). This process can be repeated until the updated segmentation mask is the optimal segmentation mask.

[0025] More specifically, the convolutional RNN can use information about the probability distribution map to integrate a segmentation method to segment an image. For example, the convolutional RNN can be used to determine a feature map for the image. Then, this feature map can be combined with a current probability distribution map (e.g., from a current segmentation method) and a previous probability distribution map (e.g., from a previous segmentation method). In particular, the feature map can be combined with the current probability distribution map to generate a first feature (e.g., the combination of the feature map and the current probability distribution map). Additionally, the feature map can be combined with the previous probability distribution map to generate a second feature (e.g., the combination of the feature map and the previous probability distribution map). Then, the first feature and the second feature can be concatenated to generate an updated probability distribution map. This updated probability distribution map incorporates information about the current segmentation method and the previous segmentation method.

[0026] In some embodiments, the convolutional RNN can use information about segmenting an image to integrate a removal image method. The removal segmentation method can occur when an interactive object selection indicates that an object, a separate feature, or a portion of the image should be removed from the desired segmentation mask. Generally, when performing a removal image segmentation method, there needs to be a pre-existing segmentation mask from which the object, feature, or portion to be excluded from the desired segmentation mask can be selected. Generally, when performing a removal image segmentation method on an image that is a first segmentation, there is not enough information to indicate what object, feature, or portion to exclude from the image selection. For example, when only a dog is desired, the first segmentation method can generate a first segmentation mask of a cat and a dog. When a second segmentation method (e.g., a removal image segmentation method) is used to indicate that the cat should not be included in the segmentation mask, generally the removal segmentation tool will not have any information about the first segmentation mask of the cat and the dog. For example, the integrated segmentation system does not have any knowledge of the type of image segmentation method used. Thus, the integrated segmentation system can incorporate the removal information into the convolutional RNN so that the system keeps track of what objects, features, or portions of the image should be excluded from the desired segmentation mask.

[0027] In particular, a convolutional RNN can be used to determine a feature map for an image. Such a feature map can incorporate removal information. For example, if an interactive object selection indicates the removal of a cat, information about the cat can be removed from the determined feature map. Then, this feature map (e.g., having information about the cat being removed) can be combined with a current probability distribution map (e.g., from a current segmentation method) and a previous probability distribution map (e.g., from a previous segmentation method). In particular, the feature map can be combined with the current probability distribution map to generate a first feature (e.g., the combination of the feature map and the current probability distribution map). Additionally, the feature map can be combined with the previous probability distribution map to generate a second feature (e.g., the combination of the feature map and the previous probability distribution map). Then, the first feature and the second feature can be concatenated to generate an updated probability distribution map. This updated probability distribution map incorporates information about the current segmentation method and the previous segmentation method.

[0028] The convolutional RNN can be trained to integrate various interactive image segmentation methods. In one embodiment, two image segmentation methods can be used to train the convolutional RNN (e.g., PhraseCut and Deep Interactive Object Selection (“DIOS”). For example, when a spoken command is received, a language-based segmentation method (e.g., PhraseCut) can be used. When a click is received, a click-based segmentation method (e.g., DIOS) can be used. First, to train the convolutional RNN, a first segmentation method can be performed on an image. To perform the first segmentation method, an interactive object selection indicating which segmentation should be used can be received (e.g., a spoken command for PhraseCut and a click for DIOS). Then, the first segmentation method can be run to generate a probability distribution map. This probability distribution map can be fed, along with the image, into the convolutional RNN. The convolutional RNN can store a hidden state about the first segmentation method and output an updated probability mask. The loss between the updated probability mask and the ground-truth mask can be used to update the convolutional RNN. For example, a pixel loss can be used.

[0029] In an embodiment, the convolutional RNN does not have knowledge of the method for generating the probability distribution map. Training the convolutional RNN without this knowledge ensures that the trained convolutional RNN will be scalable to any type of image segmentation method. For example, although the convolutional RNN can be trained using two image segmentation methods, when performing an optimal segmentation of an image, the trained convolutional RNN can be used to integrate any number of image segmentation methods.

[0030] In an embodiment, the classifier neural network can generate a segmentation mask from the updated probability distribution map. Such a classifier neural network can receive features (e.g., in the form of an updated probability distribution map) and generate a final output (e.g., in the form of an optimal segmentation mask). For example, the classifier neural network can include a decoder portion that can extract features into a feature space that is not interpretable by humans and transform the features back into an image state. In some embodiments, the classifier neural network can be trained on how to combine features from a feature map combined with the current probability distribution map and from a feature map combined with the previous probability distribution map. For example, because the convolutional RNN does not have any knowledge of the image segmentation method used, the classifier neural network can use this information to intelligently combine a first feature regarding a first image segmentation method with a second feature regarding a second image segmentation method. As an example, if the first image segmentation method is more accurate with respect to the interior of an object but less accurate for the edges, and the second image segmentation method is less accurate with respect to the interior of the object but more accurate for the edges, then the classifier neural network can combine the first and second features accordingly (e.g., favoring portions regarding the more reliable / accurate method).

[0031] Go to Figure 1A , Figure 1A FIG. depicts an example configuration of an operating environment in accordance with various embodiments of the present disclosure, in which some implementations of the present disclosure may be employed. It should be understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, and groupings of functions, etc.) may be used in addition to or in place of the arrangements and elements shown, and some elements may be omitted altogether for clarity. Moreover, many of the elements described herein are functional entities that may be implemented as discrete or distributed components, or in conjunction with other components, and in any suitable combination and location. The various functions described herein as being performed by one or more entities may be implemented by hardware, firmware, and / or software. For example, some functions may be performed by a processor executing instructions stored in a memory, as further described with reference to Figure 10 is further described.

[0032] It should be understood that Figure 1A The operating environment 100 shown in is an example of a suitable operating environment. Among other components not shown, the operating environment 100 includes a plurality of user devices, such as user devices 102a and 102b through 102n, a network 104, and (one or more) servers 108. Each of the components shown in Figure 1A can be implemented via any type of computing device, e.g., such as in conjunction with Figure 10One or more computing devices in the described computing device 1000. These components may communicate with each other via network 104, which may be wired, wireless, or both. Network 104 may include multiple networks or a network of networks, but is shown in a simplified form so as not to obscure various aspects of the present disclosure. By way of example, network 104 may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet), and / or one or more private networks. In the case where network 104 includes a wireless telecommunications network, components such as base stations, communication towers, and even access points (and other components) may provide wireless connectivity. Networked environments are common in offices, enterprise-wide computer networks, intranets, and the Internet. Network 104 may be any network capable of communicating between machines, databases, and (mobile or other) devices. Thus, network 104 may be a wired network, a wireless network (e.g., a mobile network or a cellular network), a storage area network (SAN), or any suitable combination thereof. In one example embodiment, network 104 includes one or more portions of a private network, a public network (such as the Internet), or a combination thereof. Thus, network 104 is not described in detail.

[0033] It should be understood that any number of user devices, servers, and other components may be employed within the operating environment 100 within the scope of the present disclosure. Each component may include a single device or multiple devices that cooperate in a distributed environment.

[0034] User devices 102a through 102n may be any type of computing device capable of being operated by a user. For example, in some implementations, user devices 102a through 102n are of the type of computing device Figure 10 described. By way of example and not limitation, a user device may be implemented as a personal computer (PC), laptop computer, mobile device, smart phone, tablet computer, smart watch, wearable computer, personal digital assistant (PDA), MP3 player, global positioning system (GPS) or device, video player, handheld communication device, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, device, consumer electronic device, workstation, any combination of these depicted devices, or any other suitable device.

[0035] A user device may include one or more processors and one or more computer-readable media. The computer-readable media may include computer-readable instructions executable by one or more processors. The instructions may be from one or more applications (such as in Figure 1Aimplemented by the application 110 shown in. For simplicity, the application 110 is referred to as a single application, but its functionality can in practice be implemented by one or more applications. As indicated above, other user devices may include one or more applications similar to the application 110.

[0036] (Multiple) applications can generally be any application capable of supporting information exchange between the user device and the (multiple) servers 108 when performing image editing, such as splitting an image to generate a segmentation mask for editing the image. In some implementations, the (multiple) applications include web applications that can run in a web browser and can be at least partially hosted on the server side of the environment 100. Additionally, or conversely, the (multiple) applications can include dedicated applications, such as applications with image editing and / or processing functionality. For example, such an application can be configured to display an image and / or allow a user to input or identify an image for editing. In some cases, the application is integrated into the operating system (e.g., as a service). Thus, "application" is expected to be interpreted broadly herein. Example applications include etc.

[0037] According to embodiments herein, the application 110 can support segmenting an image, presenting the segmentation as a segmentation mask, and using the optimal segmentation mask to edit the image. In particular, a user can select or input an image or picture for segmentation. The image and / or picture can be selected or input in any way. The application can support access to one or more images stored on the user device 102a (e.g., in a photo library), and / or import images from remote devices 102b - 102n and / or applications (such as from the server 108). For example, a user can take a photo using a camera on the device (e.g., the user device 102a). As another example, a user can select a desired image from a repository (e.g., stored in a network-accessible data repository or locally stored at the user device 102a). Based on the input image, some of these techniques are further discussed with reference to Figure 2 the integrated segmentation system 204, and the segmentation mask can be provided to the user via the user device 102a.

[0038] In particular, the user can use Application 110 to perform interactive object selection of an image. Such interactive object selection can be based on interactive actions (e.g., click, scribble, bounding box, and / or language). Based on the received interactive object selection, the image can be segmented using the selected segmentation method. After undergoing segmentation, the segmentation mask can be displayed to the user. The user can further interact with the image and the displayed segmentation mask using additional interactive object selections. Such interactive object selections can indicate further refinement that the user desires to the displayed segmentation mask. From these additional interactive object selections, an updated segmentation mask (e.g., an optimized segmentation mask) can be displayed to the user. The updated segmentation mask can be generated using the integrated segmentation techniques further discussed below with reference to Figure 2 the integrated segmentation system 204 of

[0039] The user device can communicate with the server 108 (e.g., a software as a service (SAAS) server) via the network 104, and the server 108 provides a cloud-based and / or network-based integrated segmentation system 106. The integrated segmentation system can communicate with the user device and the corresponding user interface to support segmentation of an image via, for example, Application 110 and / or presentation of the image to the user.

[0040] As described herein, the server 108 can support segmenting an image, generating an optimized segmentation mask, and presenting such segmentation as a segmentation mask via the integrated segmentation system 106. The server 108 includes one or more processors and one or more computer-readable media. The computer-readable media includes computer-readable instructions executable by one or more processors. As described in further detail below, the instructions can selectively implement one or more components of the integrated segmentation system 106. The server 108 is capable of integrating various interactive image segmentation methods to generate an optimized segmentation mask. Such interactive image segmentation methods can be stored in a unified library. The unified library allows various image segmentation methods to build on each other in response to interactions (e.g., click, slide, bounding box) indicating a desired segmentation of the image.

[0041] For cloud-based implementations, the instructions on the server 108 can implement one or more components of the integrated segmentation system 106. Application 110 can be used by the user to interface with functionality implemented on the server(s) 108, such as the integrated segmentation system 106. In some cases, Application 110 includes a web browser. In other cases, as further discussed with reference to Figure 1B the server 108 may not be required.

[0042] Accordingly, it should be understood that the integrated segmentation system 106 can be provided via multiple devices that jointly provide the functionality described herein in a distributed environment. Additionally, other components not shown may also be included within the distributed environment. Additionally or alternatively, the integrated segmentation system 106 can be at least partially integrated into a user device (such as user device 102a).

[0043] Reference Figure 1B is made to various aspects of an illustrative integrated segmentation system in accordance with various embodiments of the present disclosure. Figure 1B A user device 114 is depicted in accordance with an example embodiment and is configured to allow for optimal segmentation of an image by integrating multiple segmentation methods. The user device 114 can be the same as or similar to user devices 102a - 102n and can be configured to support the integrated segmentation system 116 (as a stand-alone or networked device). For example, the user device 114 can store and execute software / instructions to support interaction between the user and the integrated segmentation system 116 via the user interface 118 of the user device.

[0044] The user device can be utilized by a user to support segmenting an image, presenting the segmentation as a segmentation mask, and editing the image using the segmentation mask (e.g., an optimized segmentation mask). In particular, the user can use the user interface 118 to select or input an image or picture for segmentation. The image and / or picture can be selected or input in any manner. The user interface can support the user in accessing one or more images stored on the user device (e.g., in a photo library), and / or importing images from remote devices and / or applications. Based on the input image, the input image can be segmented using various techniques (some of which are further discussed below with reference to Figure 2 the integrated segmentation system 204), and an optimized segmentation mask can be provided to the user via the user interface. After combining multiple image segmentation methods to generate an optimized segmentation mask, the segmentation mask can be used to edit the image.

[0045] Referring Figure 2 to, various aspects of an illustrative image segmentation environment 200 in accordance with various embodiments of the present disclosure are shown. The integrated segmentation system 204 includes an interaction analysis engine 206, a unified library engine 208, and an integration engine 210. The foregoing engines of the integrated segmentation system 204 can be, for example, in Figure 1A the operating environment 100 and / or Figure 1Bimplemented in the operating environment 112. In particular, these engines can be integrated into any suitable combination of user devices 102a and 102b through 102n and (a) server(s) 108 and / or user device 114. Although the interaction analysis engine, the unified library engine, and the integration engine are depicted as separate engines, it should be understood that a single engine can perform the functionality of one or more of the engines. Additionally, in an implementation, the functionality of the engines can be performed using additional engines. Further, it should be understood that the functionality of the unified library engine can be provided by a system separate from the integration segmentation system (e.g., an image segmentation system).

[0046] Such an integration segmentation system can work in conjunction with the data repository 202. The data repository 202 can store computer instructions (e.g., software program instructions, routines, or services), data, and / or models used in the embodiments described herein. In some implementations, the data repository 202 can store information or data received via the various engines and / or components of the integration segmentation system 204 and provide access to such information or data to the various engines and / or components as needed. Although depicted as a single component, the data repository 202 can be implemented as one or more data repositories. Additionally, the information in the data repository 202 can be distributed in any suitable manner across one or more data repositories for storage (which can be hosted externally).

[0047] In an embodiment, the data repository 202 can be used to store a neural network system that can be used to perform optimal segmentation of an image by integrating multiple segmentation methods. In particular, such optimal segmentation can be based on deep learning techniques (discussed further below with reference to the integration engine 210). Such a neural network system can consist of one or more neural networks.

[0048] In an embodiment, the data stored in the data repository 202 can include an image that a user can select for segmentation using, for example, the integration segmentation system. The image can include a visual representation of a person, object, or scene. Examples of images can include digital versions of pictures, paintings, drawings, and / or photographs. Such an image can be input into the data repository 202 from a remote device (such as from a server or a user device). The data stored in the data repository 202 can also include a segmentation mask generated for the image. Such a segmentation mask can be stored as multiple segments and / or masks. The segments can be created along boundaries within the image, and / or the segments can be used to specify objects and / or features within the image. The data stored in the data repository 202 can also include an edited image. The saved edits can include manipulation of the image by applying the edits to the corresponding portions of the image using the selected segmentation mask.

[0049] The integrated segmentation system 204 can generally be used to segment images. Specifically, the integrated segmentation system can be configured to perform optimal segmentation of an image by integrating multiple segmentation methods. As used herein, image segmentation is the process of partitioning an image into a segmentation mask based on one or more interactions with the image. Such an interaction can be an interactive object selection (e.g., indicating a portion of the image to be segmented into the segmentation mask). Such a segmentation mask can be used in image editing to selectively apply edits only to the portion of the image indicated by the interaction (e.g., interactive object selection).

[0050] The interaction analysis engine 206 can receive and analyze interactions with the image. These interactions can be interactive object selections. Such interactive object selections can be based on interactive actions performed by the user. For example, the interactive actions can include clicking, scribbling, drawing a bounding box, and / or speaking a language command. The interactive actions can indicate an object, region, and / or a portion of the image to be included or excluded from the segmentation. The user can perform an interactive object selection on the image using a graphical user interface (GUI). As an illustrative example, the user can click on a dog in the image to indicate that the dog should be included in the segmentation. As another example, the user can speak a command to indicate that the dog should be included in the image segmentation.

[0051] In some embodiments, the interaction analysis engine 206 can further determine a specific image segmentation method that should be used to segment the image. In some embodiments, the user can select a specific image segmentation method. For example, the user can explicitly select the method (e.g., via graphical user interface interaction) by selecting the image segmentation method. In an embodiment, when the user explicitly selects the method, the interaction analysis engine 206 can receive the specific image segmentation method that should be used to segment the image. For example, the user can select a segmentation tool for indicating a specific image segmentation method. For example, the user can explicitly select the method (e.g., via graphical user interface interaction) by selecting the image segmentation method. In other embodiments, the user can implicitly select the method (e.g., the method can be selected based on the interactive object selection).

[0052] From the received interactive object selection, the integrated segmentation system can run a specific image segmentation method. In some embodiments, the interaction analysis engine 206 can analyze the received interactive object selection to determine a specific image segmentation method that should be used to segment the image. For example, the method can be selected based on the interactive object selection. For example, if the interactive object selection is a spoken command, the selected image segmentation method can be a language-based segmentation method (e.g., PhraseCut). In some instances, a trained neural network can be used to determine a specific image segmentation method that should be used to segment the image.

[0053] Once a specific image segmentation method for segmenting an image (e.g., using the interactive analysis engine 206) is selected, the unified library engine 208 can run the image segmentation method. The unified library engine 208 can include any number of image segmentation methods. Such image segmentation methods can be implemented using techniques including: deep learning techniques, color range or saliency detection, thresholding, clustering methods, compression-based methods, histogram-based methods, edge detection, double clustering methods, region growing methods, partial differential equation-based methods, variational methods, graph partitioning methods, watershed transformation, model-based segmentation, multi-scale segmentation, and semi-automatic segmentation. Specifically, deep learning techniques can include instance-level semantic segmentation, human parsing with automatic boundary awareness, object detection using cascaded CNNs, general segmentation algorithms, etc.

[0054] As shown, the unified library engine 208 can include a segmentation component 212. The foregoing components of the unified library engine 208 can be implemented, for example, in Figure 1A the operating environment 100 and / or Figure 1B the operating environment 112. In particular, these components can be integrated into any suitable combination of the user devices 102a and 102b to 102n and the (multiple) servers 106 and / or the user device 114. It should be understood that although the segmentation component is depicted as a single component, one or more additional components can be used to perform the functionality of the component in implementation.

[0055] Generally, the segmentation component 212 can be configured to perform an image segmentation method. In particular, the segmentation component can perform a specific image segmentation method selected for segmenting an image (e.g., using the interactive analysis engine 206). When performing the image segmentation method, the segmentation component 212 can generate a probability distribution map based on the received image and the interactive object selection regarding the image. The probability distribution map can generally be the output of the segmentation method (e.g., a segmentation mask, a heat map, etc.). For example, when segmenting an image, the probability distribution map can be the information generated by the segmentation method. A segmentation mask can be generated from the probability distribution map.

[0056] The image can be accessed or referenced by the segmentation component 212 for segmentation. In this regard, the segmentation component 212 can access or retrieve the image selected by the user via the data repository 202 and / or from a remote device (such as from a server or a user device). As another example, the segmentation component 212 can receive the image provided to the integrated segmentation system 204 via the user device.

[0057] Based on running an image segmentation method on an image (e.g., using segmentation component 212), integration engine 210 can be utilized to obtain an optimal segmentation of the image by integrating multiple segmentation methods. For example, an optimal segmentation can be generated by iteratively integrating a previous segmentation (e.g., using an image segmentation method performed by segmentation component 212) with a current segmentation (e.g., using the same or a different image segmentation method performed by segmentation component 212). As an example, as the user interacts with the image, integration engine 210 can integrate (multiple) previous segmentation methods into the current image segmentation.

[0058] In an embodiment, integration engine 210 can use one or more neural networks based on a convolutional recurrent neural network (RNN) architecture. For example, integration engine 210 can implement a convolutional RNN to maintain information when transitioning from one segmentation method to another, so that multiple interactive image segmentation methods can be incorporated into the integrated image segmentation process.

[0059] Integration engine 210 can be used to train a convolutional RNN to integrate information about (multiple) previous segmentations into the current segmentation. Various interactive image segmentation methods can be used to train the convolutional RNN. In one embodiment, two image segmentation methods (e.g., PhraseCut and DIOS) can be used to train the convolutional RNN. For example, when a spoken command is received, a language-based segmentation method (e.g., PhraseCut) can be used. When a click is received, a click-based segmentation method (e.g., DIOS) can be used.

[0060] Initially, to train the convolutional RNN, a first segmentation method can be performed on the image. To perform the first segmentation method, an interactive object selection indicating which segmentation should be used (e.g., a spoken command for PhraseCut and a click for DIOS) can be received. Then, the first segmentation method can be run to generate a probability distribution map. The probability distribution map can be fed into the convolutional RNN along with the image. The convolutional RNN can store a hidden state about the first segmentation method and output an updated probability mask. The loss between the updated probability mask and the ground truth mask can be used to update the convolutional RNN. For example, pixel loss can be used.

[0061] Integration engine 210 can also use one or more classifier neural networks. For example, integration engine 210 can implement a classifier neural network that can generate a segmentation mask from the updated probability distribution map. Such a classifier neural network can receive features (e.g., in the form of the updated probability distribution map) and generate a final output (e.g., in the form of an optimal segmentation mask). For example, the classifier neural network can include a decoder part that can extract features into a feature space that is not interpretable by humans and transform the features back to an image state.

[0062] In some embodiments, the integration engine 210 may train a classifier neural network to combine features from a feature map combined with a current probability distribution map and features from a feature map combined with a previous probability distribution map. For example, in embodiments where the convolutional RNN does not have knowledge of the image segmentation method being used, the classifier neural network may use this information to intelligently combine a first feature regarding a first image segmentation method (e.g., the combination of the feature map and the current probability distribution map) and a second feature regarding a second image segmentation method (e.g., the combination of the feature map and the previous probability distribution map). As an example, if the first image segmentation method is more accurate regarding the interior of an object but less accurate for the edges, and the second image segmentation method is less accurate regarding the interior of the object but more accurate for the edges, the classifier neural network may thus combine the first feature and the second feature (e.g., preferring the parts regarding the more reliable / accurate method).

[0063] As shown, the integration engine 210 may include an image analysis component 214, a removal component 216, and a mask component 218. The foregoing components of the integration engine 210 may be implemented, for example, in Figure 1A the operating environment 100 and / or Figure 1B the operating environment 112. In particular, these components may be integrated into any suitable combination of the user devices 102a and 102b to 102n and the (multiple) servers 106 and / or the user device 114. It should be understood that although the image analysis component, the removal component, and the mask component are depicted as separate components, in implementation, a single component and / or additional components may be used to perform the functionality of the engine. Additionally, in some embodiments, the removal component 216 and its associated functionality are optional and may be excluded from the integration engine 210.

[0064] Generally, the image analysis component 214 may be configured to analyze an image. In particular, the image analysis component 214 may be used to determine the features of the image. In one embodiment, such analysis may be performed using, for example, a neural network. Such a neural network may be based on a convolutional RNN architecture. For example, the convolutional RNN may be used to determine a feature map for the image. The feature map may indicate the features of the image. Such a feature map may assist in segmenting the image based on the selected / unselected objects. For example, if an interactive input indicates that a cat should be included in the segmentation, a representation of the feature map regarding the cat may be used to help generate a segmentation mask that includes the cat.

[0065] The removal component 216 can be used in instances where an interactive object selection indicates that an object, feature, or portion of an image should be excluded from a desired segmentation mask. The removal component can incorporate removal information from an interactive object selection (e.g., received and / or analyzed by the interaction analysis engine 206) into the convolutional RNN so that the objects, features, or portions of the image that should be excluded from the desired segmentation mask can be tracked. In particular, after the convolutional RNN determines a feature map for an image (e.g., using the image analysis component 214), the removal component 216 can incorporate the removal information into the feature map. For example, if the interactive object selection indicates the removal of an object, information about the object can be removed from the determined feature map.

[0066] Generally, the mask component 218 is configured to generate a segmentation mask for an image based on an optimal segmentation by integrating multiple segmentation methods. For example, the previous segmentation (e.g., using an image segmentation method) can be iteratively integrated with the current segmentation (e.g., using the same or a different image segmentation method) to generate an optimal segmentation mask. The mask component 218 can use the convolutional RNN to integrate information about the (multiple) previous segmentations when performing the current segmentation. In particular, the mask component 218 can receive information about the previous segmentation of an image and combine the information about the previous segmentation of the image with the information about the current segmentation of the image.

[0067] As shown, the mask component 218 can include a probability map element 220 and a generation element 222. The foregoing elements of the mask component 218 can be implemented, for example, in Figure 1A the operating environment 100 and / or Figure 1B the operating environment 112. In particular, these elements can be integrated into any suitable combination of the user devices 102a and 102b to 102n and the (multiple) servers 106 and / or the user device 114. It should be understood that although the probability map element and the generation element are depicted as separate elements, in an implementation, a single element and / or additional elements can be used to perform the functionality of the elements.

[0068] The probability map element 220 can be used to integrate information about the (multiple) previous segmentations when performing the current segmentation. Such information can at least include a probability distribution map. The probability distribution map can generally be an output of a segmentation method (e.g., a segmentation mask, a heat map, etc.). For example, a probability distribution map from the current image segmentation method can be received from the segmentation component 212. Based on the received probability distribution map, the convolutional RNN can generate a hidden state based on the information about the probability distribution map.

[0069] Subsequently, when an image is segmented using a subsequent segmentation method (e.g., using the segmentation component 212), the current probability distribution map can be received by the probability map element 220. The probability map element 220 can update the current probability distribution map using the hidden state of the convolutional RNN (e.g., based on the hidden state of the probability distribution map regarding the previous image segmentation method). The probability map element 220 can generate an updated probability distribution map that incorporates information about the previous segmentation into the current segmentation.

[0070] Generally, the convolutional RNN can update the hidden state (e.g., S t ), the updated probability distribution map (e.g., Y t ), and the current segmentation mask (e.g., M t ) as follows.

[0071]

[0072]

[0073]

[0074] In such an equation, and can be learnable convolutional parameters, and can be a bias. σ(x) = 1 / (1 + e -x ) can be a double sigmoid function. At each stage, the hidden state can be updated based on the previous state (e.g., S t-1 ) and the new probability map (e.g., P t ).

[0075] The generation element 222 can use the generated updated probability distribution map to generate an updated segmentation mask (e.g., combining the previous segmentation and the current segmentation). The segmentation mask can be presented to the user in various ways. The generation unit 222 can run a classifier neural network to generate a segmentation mask from the updated probability distribution map.

[0076] Reference Figure 3 provides a processing flow showing an embodiment of a method 300 for performing integration of multiple segmentation methods according to an embodiment of the present disclosure. As illustrated in Figure 2 , the method 300 can be executed, for example, by the integrated segmentation system 204.

[0077] At block 302, an image is received. Such an image can be from a database stored (such as Figure 2A set of images or pictures in the data repository 202) is received. In particular, the user may select or input the received images. Such images may be selected or input in any way. For example, the user may take a photo using the camera on the device. As another example, the user may select a desired image from a repository (e.g., stored in a data repository accessible via a network or locally stored at the user device).

[0078] At block 304, an interaction is received. The interaction may be an interactive object selection of an image by the user. Such interactive object selection may be based on interactive operations (e.g., click, scribble, bounding box, and / or language). Based on the received interactive object selection, an image segmentation method for segmenting the image (e.g., received at block 302) may be selected. In some embodiments, the user may select a particular image segmentation method. For example, the user may explicitly select the method by selecting an image segmentation method (e.g., via a graphical user interface interaction). In other embodiments, the image segmentation method for segmenting the image is based on the received interaction. In this way, the user may implicitly select the method (e.g., the method may be selected based on the interactive object selection).

[0079] At block 306, segmentation of the image is performed. Image segmentation is the process of partitioning an image into at least one segment. In particular, segments may be created along boundaries within the image, and / or segments may be used to specify objects and / or features within the image. For example, when a segmentation method is used to segment an image, a probability distribution map may be generated. The probability distribution map may generally be the output of the segmentation method (e.g., a segmentation mask, a heat map, etc.). For example, when segmenting an image, the probability distribution map may be the information generated by the segmentation method.

[0080] This segmentation may be performed using any number of techniques. Such techniques include deep learning techniques, color range or saliency detection, thresholding, clustering methods, compression-based methods, histogram-based methods, edge detection, dual clustering methods, region growing methods, partial differential equation-based methods, variational methods, graph partitioning methods, watershed transform, model-based segmentation, multi-scale segmentation, and semi-automatic segmentation. Specifically, deep learning techniques may include instance-level semantic segmentation, automatic boundary-aware human parsing, object detection using cascaded convolutional neural networks, general segmentation algorithms such as R-CNN and / or Mask R-CNN.

[0081] At block 308, probability map integration is performed. Probability map integration is the process of combining a previous probability distribution map (e.g., from a previous segmentation method) and a current probability distribution map (e.g., from a current segmentation method).

[0082] In particular, a convolutional RNN can be used to receive information about a previous probability distribution map. Using this information, the convolutional RNN can generate a hidden state based on the information about the previous probability distribution map. Subsequently, when the current segmentation method is used to segment an image, this hidden state can be used to update the current probability distribution map (e.g., determined using the current segmentation method) to generate an updated probability distribution map. In this way, the hidden state of the convolutional RNN can be used to incorporate information about previous segmentations into the current segmentation.

[0083] At block 310, a segmentation mask is generated. The generated segmentation mask can be generated using the resulting updated probability distribution map. The segmentation mask can combine the previous segmentation and the current segmentation.

[0084] At block 312, the segmentation mask can be presented. The presentation of the segmentation mask allows a user to view and visualize the (multiple) segmented regions of the image. The user can further interact with the image and the displayed segmentation mask with additional (multiple) interactive object selections. Such interactive object selections can indicate further refinements that the user desires to make to the displayed segmentation mask. From these additional (multiple) interactive object selections, an updated segmentation mask (e.g., an optimized segmentation mask) can be displayed to the user.

[0085] Reference Figure 4 provides a process flow diagram illustrating an embodiment of a method 400 for integrating multiple segmentation methods in accordance with an embodiment of the present disclosure. As illustrated in Figure 2 , method 400 can be performed, for example, by an integrated segmentation system 204.

[0086] At block 402, an image is received. Such an image can be received from a set of images or pictures stored in a database (e.g., Figure 2 data repository 202). In particular, a user can select or input the received image. Such an image can be selected or input in any manner. For example, a user can take a photo using a camera on the device. As another example, a user can select a desired image from a repository (e.g., stored in a network-accessible data repository or locally stored at the user device). At block 404, a feature map is generated for the image (e.g., the image received at block 402). Such a feature map can generally relate to information about objects, individual features, and / or portions of the image. The feature map can be generated using a convolutional RNN.

[0087] At block 406, probability map integration is performed. The probability distribution map can generally be the output of a segmentation method (e.g., a segmentation mask, a heatmap, etc.). Probability map integration is the process of combining a previous probability distribution map (e.g., from a previous segmentation method) with a current probability distribution map (e.g., from a current segmentation method). In particular, the feature map of an image can be combined with the current probability distribution map (e.g., from a current segmentation method) and the previous probability distribution map (e.g., from a previous segmentation method). For example, the feature map can be combined with the current probability distribution map to generate a first feature (e.g., the combination of the feature map and the current probability distribution map). Additionally, the feature map can be combined with the previous probability distribution map to generate a second feature (e.g., the combination of the feature map and the previous probability distribution map). Then, the first feature and the second feature can be concatenated. The concatenated first and second features can be used to generate an updated probability distribution map. The updated probability distribution map incorporates information about the current and previous segmentation methods.

[0088] More specifically, probability map integration can be performed by a convolutional RNN that receives information about the previous probability distribution map. Using this information, the convolutional RNN can generate a hidden state based on the information about the previous probability distribution map. Subsequently, when a subsequent segmentation method (e.g., the current segmentation method) is used to segment the image, this hidden state can be used to update the current probability distribution map (e.g., determined using the subsequent segmentation method) to generate an updated probability distribution map. In this way, the hidden state of the convolutional RNN can be used to incorporate information about the previous segmentation into the current segmentation.

[0089] At block 408, a segmentation mask is generated. In particular, the updated probability distribution map can be used to generate a segmentation mask (e.g., combining the previous and current segmentations). The user can further interact with the image and the generated segmentation mask using additional (multiple) interactive object selections. Such interactive object selections can indicate further refinements that the user desires to make to the segmentation mask. Based on these additional interactive object selections, an optimized segmentation mask can be generated.

[0090] Reference Figure 5 provides a process flow diagram illustrating an embodiment of a method 500 for integrating multiple segmentation methods to generate an optimized segmentation mask according to an embodiment of the present disclosure. Method 500 can be performed, for example, by an integrated segmentation system 204 as illustrated in Figure 2 .

[0091] At block 502, an image is received. Such an image can be retrieved from a database (such as Figure 2A set of images or pictures in the data repository 202) is received. In particular, the user can select or input the received images. Such images can be selected or input in any way. For example, the user can take a photo using the camera on the device. As another example, the user can select a desired image from a repository (e.g., stored in a data repository accessible via a network or locally stored at the user device).

[0092] At block 504, an interaction is received. The interaction can be an interactive object selection of an image by the user. Such interactive object selection can be based on interactive operations (e.g., click, scribble, bounding box, and / or language). Based on the received interactive object selection, an image segmentation method for segmenting the image (e.g., received at block 502) can be selected. In some embodiments, the user can select a specific image segmentation method. For example, the user can explicitly select the method by selecting the image segmentation method (e.g., via a graphical user interface interaction). In other embodiments, the image segmentation method for segmenting the image is based on the received interaction. In this way, the user can implicitly select the method (e.g., the method can be selected based on the interactive object selection).

[0093] At block 506, segmentation of the image is performed. Image segmentation is the process of partitioning an image into segments. In particular, segments can be created along boundaries within the image, and / or segmentation can be used to indicate objects and / or features within the image. For example, image segmentation can generate a probability distribution map. The probability distribution map can generally be the output of a segmentation method (e.g., a segmentation mask, a heat map, etc.). Such segmentation can be performed using any number of techniques. Such techniques include deep learning techniques, color range or saliency detection, thresholding, clustering methods, compression-based methods, histogram-based methods, edge detection, double clustering methods, region growing methods, partial differential equation-based methods, variational methods, graph partitioning methods, watershed transform, model-based segmentation, multi-scale segmentation, and semi-automatic segmentation. Specifically, deep learning techniques can include instance-level semantic segmentation, automatic boundary-aware human parsing, object detection using cascaded convolutional neural networks, general segmentation algorithms such as R-CNN and / or Mask R-CNN.

[0094] At block 508, it is determined whether a previous segmentation has been performed on the image. If there is no previous segmentation, the process can proceed to block 510. If there has been a previous segmentation, the process can proceed to block 516, which is described in further detail below.

[0095] At block 510, a probability distribution map is generated from the current segmentation. The probability distribution map can generally be the output of a segmentation method (e.g., a segmentation mask, a heatmap, etc.). At block 512, the hidden state is stored. In particular, a convolutional RNN can receive information about the probability distribution map. Using this information, the convolutional RNN can generate a hidden state. This hidden state can be used to maintain information about the probability distribution map. By maintaining this information, the convolutional RNN can combine information about the probability distribution map with subsequently generated probability distribution maps (e.g., combine a previous image segmentation method with a subsequent image segmentation method).

[0096] At block 514, a segmentation mask is generated. The generated segmentation mask can be generated using the probability distribution map. The user can further interact with the image using additional interactive object selections (e.g., leveraging subsequent interactions at block 504). Such interactive object selections can indicate that the user desires further refinement of the displayed segmentation mask. Based on these additional interactive object selections, as further discussed with reference to blocks 516 to 520, an updated segmentation mask (e.g., an optimized segmentation mask) can be generated.

[0097] At block 516, the current probability distribution map is received. The current probability map can be based on the current segmentation. At block 518, the previous probability distribution map is received. The previous probability distribution map can be based on a previous segmentation. For example, the previous probability distribution map can be received using the hidden state. In particular, a convolutional RNN can be used to receive information about the previous probability distribution map. Using this information, the convolutional RNN can generate a hidden state based on the information about the previous probability distribution map.

[0098] At block 520, the current probability distribution map is integrated with the previous probability distribution map. In particular, the current probability distribution map (e.g., determined using the current segmentation method) can be updated using the hidden state to generate an updated probability distribution map. In this way, the hidden state of the convolutional RNN can be used to incorporate information about the previous segmentation into the current segmentation.

[0099] At block 514, a segmentation mask is generated. In particular, the resulting updated probability distribution map can be used to generate an updated segmentation mask (e.g., combining the previous segmentation and the current segmentation). The generated segmentation mask can be presented to the user. The presentation of the segmentation mask allows the user to view and visualize the (multiple) segmented regions of the image. As described above, the user can use additional (multiple) interactive object selections to further interact with the image and the displayed segmentation mask.

[0100] Reference Figure 6 , a process flow is provided showing an embodiment of a method 600 for integrating multiple segmentation methods including a removal segmentation method according to an embodiment of the present disclosure. Method 600 can be performed, for example, by, as inFigure 2 performed by the integrated segmentation system 204 illustrated in the figure.

[0101] At block 602, an image is received. Such an image can be received from a set of images or pictures stored in a database (e.g., Figure 2 data repository 202). In particular, the user can select or input the received image. Such an image can be selected or input in any manner. For example, the user can take a photo using a camera on the device. As another example, the user can select a desired image from a repository (e.g., stored in a data repository accessible via a network or locally stored at the user device).

[0102] At block 604, it is determined whether there is a removal interaction. The removal interaction can be an interactive object selection that indicates which objects, features, or parts of the image should be excluded from the desired segmentation mask. When there is a removal interaction, the process can proceed to block 606. When there is no removal interaction, the process can proceed to block 608.

[0103] At block 606, the removal is performed. The removal can be performed by incorporating the removal information into the convolutional RNN such that the system can maintain track of which objects, features, or parts of the image should be excluded from the desired segmentation mask.

[0104] At block 608, a feature map is generated. Such a feature map can generally involve information about the objects, features, and / or parts of the image. The feature map can be generated using a convolutional RNN. When the removal is performed at block 606, the feature map can incorporate the removal information (e.g., from block 606). For example, if the interactive object selection indicates the removal of an object, the information about that object can be removed from the feature map.

[0105] At block 610, probability map integration is performed. The probability distribution map can generally be an output of a segmentation method (e.g., a segmentation mask, a heat map, etc.). Probability map integration is a process of combining a previous probability distribution map (e.g., from a previous segmentation method) with a current probability distribution map (e.g., from a current segmentation method). For example, a feature map (e.g., having information about the object removed when the removal occurs at block 606) can be combined with the current probability distribution map (e.g., from the current segmentation method) and the previous probability distribution map (e.g., from a previous segmentation method).

[0106] In particular, a convolutional RNN can be used to receive information about a previous probability distribution map. Using this information, the convolutional RNN can generate a hidden state based on the information about the previous probability distribution map. Subsequently, when segmenting an image using a subsequent segmentation method (e.g., the current segmentation method), this hidden state can be used to update the current probability distribution map (e.g., determined using the subsequent segmentation method) to generate an updated probability distribution map. In this way, the hidden state of the convolutional RNN can be used to incorporate information about a previous segmentation into the current segmentation. Then, the resulting updated probability distribution map can be used to generate an updated segmentation mask (e.g., combining the previous segmentation and the current segmentation).

[0107] At block 612, a segmentation mask is generated. The generated segmentation mask can be presented to the user. The presentation of the segmentation mask allows the user to view and visualize the segmented regions of the image. The user can further interact with the image and the generated segmentation mask using additional (multiple) interactive object selections. Such interactive object selections can indicate further refinements that the user desires to make to the segmentation mask. Based on these additional (multiple) interactive object selections, an optimized segmentation mask can be generated.

[0108] Figure 7 An example environment 700 in accordance with an embodiment of the present disclosure is illustrated, which can be used to integrate information about (multiple) previous segmentations when performing a current segmentation. In particular, a convolutional RNN can be used to integrate this information. The convolutional RNN can combine (multiple) previous segmentations with the current segmentation without any information about the image segmentation method used to generate the segmentation.

[0109] An image 702 can be received for segmentation. Such an image can be received from a set of images or pictures stored in a database (such as Figure 2 the data repository 202). In particular, the user can select or input the received image. Such an image can be selected or input in any way. For example, the user can take a photo using a camera on the device. As another example, the user can select a desired image from a repository (e.g., stored in a network-accessible data repository or locally stored at the user device).

[0110] Various image segmentation methods (e.g., method 708a, method 708b,..., method 708n) can be integrated into a unified library 706. The unified library allows the various image segmentation methods to build on each other in response to an interaction (e.g., click, swipe, bounding box) indicating a desired segmentation for the image.

[0111] Interaction 704 can be received. The interaction can be an interactive object selection of the image by the user. Such interactive object selection can be based on interactive actions (e.g., click, scribble, bounding box, and / or language). Based on the received interactive object selection, an image segmentation method (e.g., method 708a, method 708b, …, method 708n) can be selected for segmenting the image 702. In some embodiments, the user can select the image segmentation method (e.g., method 708n). For example, the user can explicitly select the method by selecting method 708n (e.g., via a graphical user interface interaction). In other embodiments, the image segmentation method for segmenting the image is based on the interaction 704. In this way, the user can implicitly select method 708n (e.g., based on the interactive object selection).

[0112] Method 708n can be used to segment the image 702 based on the interaction 704. According to this segmentation, a current probability distribution map 714 can be generated. The probability distribution map can generally be the output of the segmentation method (e.g., segmentation mask, heatmap, etc.).

[0113] The image 702 can also be input into the CNN 710. The CNN 710 can be a convolutional recurrent neural network that is used to integrate information about (multiple) previous segmentations when performing the current segmentation. According to the image 702, the CNN 710 can generate a feature map 712. The feature map 712 can generally involve information about the objects, features, and / or parts of the image 702.

[0114] The CNN 710 can receive the current probability distribution map 714 from the unified library 706. The CNN 710 can also have information about a previous probability distribution map 716. This information about the previous probability distribution map 716 can be stored as the hidden state of the CNN 710. The CNN 710 can combine the current probability distribution map 714 and the previous probability distribution map 716 with the feature map 712. In particular, the CNN 710 can combine the current probability distribution map 714 with the feature map 712 to generate a first feature 718. The CNN 710 can also combine the previous probability distribution map 716 and the feature map 712 to generate a second feature 720. The first feature 718 and the second feature 720 can be concatenated to generate an updated feature 722. The updated feature 722 can be input into a classifier 724. The classifier 724 can be a classifier neural network capable of generating a segmentation mask from the updated feature 722. For example, the classifier neural network can include a decoder part that can extract features into a feature space that is not interpretable by humans and transform the features back into an image state. In this way, the classifier 724 can generate an updated segmentation mask 726.

[0115] Figure 8FIG. 800 illustrates an example environment that can be used to integrate information about one or more previous segmentations when performing a current segmentation that includes a removal action, in accordance with an embodiment of the present disclosure. In particular, a convolutional RNN can be used to integrate this information. The convolutional RNN can combine the one or more previous segmentations with the current segmentation without any information about the one or more image segmentation methods used to generate the segmentations.

[0116] An image 802 can be received for segmentation. Such an image can be received from a set of images or pictures stored in a database (such as Figure 2 the data repository 202). In particular, a user can select or input the received image. Such an image can be selected or input in any manner. For example, the user can take a photo using a camera on the device. As another example, the user can select a desired image from a repository (e.g., stored in a network-accessible data repository or locally stored at the user device).

[0117] Various image segmentation methods (e.g., method 808a, method 808b, …, method 808n) can be integrated in a unified library 806. The unified library allows the various image segmentation methods to build on each other in response to an interaction (e.g., click, swipe, bounding box) indicating a desired segmentation for the image.

[0118] An interaction 804 can be received. The interaction can be an interactive object selection of the image by the user. Such an interactive object selection can be based on an interactive action (e.g., click, scribble, bounding box, and / or language). Based on the received interactive object selection, an image segmentation method (e.g., method 808a, method 808b, …, method 808n) can be selected for segmenting the image 802. In some embodiments, the user can select the image segmentation method (e.g., method 808n). For example, the user can explicitly select the method by selecting method 808n (e.g., via a graphical user interface interaction). In other embodiments, the image segmentation method used to segment the image is based on the interaction 804. In this way, the user can implicitly select method 808n (e.g., based on the interactive object selection).

[0119] Based on the interaction 804, the method 808n can be used to segment the image 802. Based on this segmentation, a current probability distribution map 818 can be generated. The probability distribution map can generally be an output of the segmentation method (e.g., a segmentation mask, a heatmap, etc.).

[0120] The image 802 can also be input into the CNN 810. The CNN 810 can be a convolutional recurrent neural network that is used to integrate information about the (multiple) previous segmentations when performing the current segmentation. Based on the image 802, the CNN 810 can generate a feature map 812. The feature map 812 can generally involve information about the objects, features, and / or parts of the image 802.

[0121] When the interaction 804 indicates an object, feature, or part of the image that should be excluded (e.g., removed) from the desired segmentation mask, the removal 814 can be incorporated into the CNN 810. The removal 814 can incorporate the removal information from the interaction 804 into the CNN 810 so that the objects, features, or parts of the image 802 that should be excluded from the desired segmentation mask can be tracked. In particular, the removal 814 can be combined with the feature map 812 by the CNN 810 to generate a feature map 816. For example, if the interaction 804 indicates the removal of an object, the removal 814 can include information about the object such that the object is removed from the feature map 8,16.

[0122] The CNN 810 can receive the current probability distribution map 818 from the unified library 806. The CNN 810 can also have information about the previous probability distribution map 820. The information about the previous probability distribution map 820 can be stored as the hidden state of the CNN 810. The CNN 810 can combine the current probability distribution map 818 and the previous probability distribution map 820 with the feature map 816. In particular, the CNN 810 can combine the current probability distribution map 818 with the feature map 816 to generate a first feature 822. The CNN 810 can also combine the previous probability distribution map 820 with the feature map 816 to generate a second feature 824. The first feature 822 and the second feature 824 can be concatenated to generate an updated feature 826. The updated feature 826 can be input into a classifier 828. The classifier 828 can be a classifier neural network that is capable of generating a segmentation mask from the updated feature 826. For example, the classifier neural network can include a decoder part that can extract features into a feature space that is not interpretable by humans and transform the features back into an image state. In this way, the classifier 828 can generate an updated segmentation mask 830.

[0123] Figure 9 An example environment 900 according to an embodiment of the present disclosure is illustrated. The example environment 900 can be used for joint embedding supervision of an integrated segmentation system for allowing optimal segmentation of an image based on multiple segmentations. In particular, a convolutional RNN can be used for this joint embedding supervision. The convolutional RNN can combine the (multiple) previous segmentations with the current segmentation without any information about the (multiple) image segmentation methods used to generate the segmentations.

[0124] An image 902 may be received for segmentation. Such an image may be received from a set of images or pictures stored in a database (such as Figure 2 data repository 202). In particular, a user may select or input the received image. Such an image may be selected or input in any manner. For example, a user may take a photo using a camera on a device. As another example, a user may select a desired image from a repository (e.g., stored in a network-accessible data repository or locally stored at the user device).

[0125] (Multiple) interactions 904 may be received. (Multiple) interactions 904 may be interactive inputs from a user. These interactions may be an interactive object selection of the image 902 by the user. Such an interactive object selection may be based on an interactive action (e.g., click, scribble, bounding box, and / or language). Based on (multiple) interactions 904, a unified library 906 may select an image segmentation method for segmenting the image 902. In some embodiments, a user may select an image segmentation method. For example, a user may explicitly select a method (e.g., via a graphical user interface interaction). In other embodiments, the image segmentation method used to segment the image may be selected by the unified library 906 based on the interaction 904.

[0126] Using the image segmentation method, the unified library 906 may generate (multiple) probability distribution maps 908 (e.g., P0) based on the image 902 and (multiple) interactions 904. (Multiple) probability distribution maps 908 may be used to generate (multiple) segmentation masks 914 (e.g., M0). Next, a user may update (multiple) interactions 904 based on (multiple) segmentation masks 914. The updated (multiple) interactions 904 may be any type of interaction (e.g., the same as or different from the initial image segmentation method). The unified library 906 may select an image segmentation method for updating (multiple) probability distribution maps 908. A convolutional RNN may be used to generate (multiple) updated probability distribution maps 908. For example, the image segmentation method may update the probability distribution map (e.g., P t ) as the current output, and together with (multiple) hidden states 910 (e.g., S t-1 ), to infer an updated probability distribution map 910 (e.g., Y t ). These steps may continue until the user is satisfied with the generated segmentation mask 914 (e.g., M t ).

[0127] After describing embodiments of the present disclosure, an example operating environment in which embodiments of the present disclosure may be implemented is described below to provide a general context for various aspects of the present disclosure. Referring to Figure 10, an illustrative operating environment for implementing embodiments of the present disclosure is shown and is generally designated as computing device 1000. Computing device 1000 is only one example of a suitable computing environment and is not intended to imply any limitation as to the scope of use or functionality of the present disclosure. Nor should computing device 1000 be construed as having any dependency or requirement related to any one component or combination of components shown.

[0128] Embodiments of the present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, executed by a computer or other machine, such as a smart phone or other handheld device. Generally, program modules (or engines), including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. Embodiments of the present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. Embodiments of the present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.

[0129] Referring to Figure 10 , computing device 1000 includes a bus 1010 that directly or indirectly couples the following devices: a memory 1012, one or more processors 1014, one or more presentation components 1016, an input / output port 1018, input / output components 1020, and an illustrative power supply 1022. Bus 1010 represents one or more buses, such as an address bus, a data bus, or a combination thereof. Although, for clarity, Figure 10 the various boxes of Figure 10 are shown with clearly depicted lines, in reality, such depictions are not so clear and these lines may overlap. For example, a presentation component, such as a display device, may also be considered an I / O component. Additionally, a processor typically has memory in the form of a cache. We recognize this as being in the nature of the art and reiterate Figure 10 that the schematic of

[0130] The computing device 1000 generally includes various non-transitory computer-readable media. The non-transitory computer-readable media can be any available media that can be accessed by the computing device 1000, and the non-transitory computer-readable media can include volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, the non-transitory computer-readable media can include non-transitory computer storage media and communication media.

[0131] Non-transitory computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Non-transitory computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by the computing device 1000. Non-transitory computer storage media does not itself include a signal.

[0132] Communication media typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and communication media typically includes any information delivery media. The term "modulated data signal" means a signal that has one or more of the following characteristics: the characteristics of the signal are set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.

[0133] The memory 1012 includes computer storage media in the form of volatile and / or non-volatile memory. As depicted, the memory 1012 includes instructions 1024. When executed by the processor(s) 1014, the instructions 1024 are configured to cause the computing device (referring to the drawings discussed above) to perform any of the operations described herein, or to implement any of the program modules described herein. The memory can be removable, non-removable or a combination thereof. Illustrative hardware devices include solid state memory, hard disk drives, optical disk drives, etc. The computing device 1000 includes one or more processors that read data from various entities such as the memory 1012 or the I / O component 1020. The presentation component(s) 1016 presents data indications to the user or other device. Illustrative presentation components include display devices, speakers, printing components, vibration components, etc.

[0134] The I / O port 1018 allows the computing device 1000 to be logically coupled to other devices including I / O components 1020, and some of the I / O components 1020 may be built-in. Exemplary components include microphones, joysticks, game pads, satellite dishes, scanners, printers, wireless devices, and the like.

[0135] The embodiments presented herein have been described with respect to specific embodiments that are intended to be illustrative in all respects and not restrictive. Alternative embodiments will be apparent to those of ordinary skill in the art to which this disclosure pertains without departing from the scope of the disclosure.

[0136] From the foregoing, it will be seen that an advantage of the present disclosure is the attainment of all the above objects and aims, as well as the accomplishment of other advantages which are obvious and inherent in the construction.

[0137] It will be understood that certain features and subcombinations are useful and may be employed without reference to other features or subcombinations. This is contemplated by the claims and is within the scope of the claims.

[0138] In the foregoing detailed description, reference has been made to the accompanying drawings which form a part hereof, and in which are shown by way of illustration embodiments in which the disclosure may be practiced. It is to be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the disclosure. Accordingly, the foregoing detailed description is not to be taken in a limiting sense, and the scope of embodiments is defined by the appended claims and their equivalents.

[0139] The various aspects of the illustrative embodiments have been described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art. However, it will be apparent to those skilled in the art that alternative embodiments may be practiced using only some of the described aspects. For purposes of illustration, specific numbers, materials, and configurations have been set forth in order to provide a thorough understanding of the illustrative embodiments. However, it will be apparent to those skilled in the art that alternative embodiments may be practiced without the specific details. In other instances, well-known features have been omitted or simplified in order not to obscure the illustrative embodiments.

[0140] The various operations have been described serially in a manner that is most helpful for understanding the illustrative embodiments; however, the order of the description should not be construed as implying that these operations necessarily depend on the order. In particular, these operations need not be performed in the order presented. Additionally, the description of the operations as separate operations should not be construed as requiring that the operations be independent and / or performed by separate entities. The description of entities and / or modules as separate modules should likewise not be construed as requiring that the modules be separate and / or perform separate operations. In various embodiments, the illustrated and / or described operations, entities, data, and / or modules may be combined, broken down into further sub-parts, and / or omitted.

[0141] The phrase "in one embodiment" or "in an embodiment" is repeated. This phrase generally does not refer to the same embodiment; however, it may refer to the same embodiment. Unless the context dictates otherwise, the terms "comprising," "having," and "including" are synonyms. The phrase "A / B" means "A or B." The phrase "A and / or B" means (A), (B), or (A and B). The phrase "at least one of A, B, and C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).

Claims

1. A computer-implemented method, the method comprising: Receiving a first interaction with an image; Receiving a second interaction with the image after the first interaction; Based on the second interaction with the image, performing segmentation of the image using a first segmentation method to generate a current probability distribution map; Combining the feature map of the image with the current probability distribution map, and combining the feature map with a previous probability distribution map, the previous probability distribution map being generated based on a previous segmentation of the image, the previous segmentation being performed based on the first interaction with the image and using a second segmentation method different from the first segmentation method; And Generating a segmentation mask based on a concatenation of the feature map combined with the current probability distribution map and the feature map combined with the previous probability distribution map.

2. The computer-implemented method according to claim 1, further comprising: Generating the feature map for the image; And Incorporating removal information into the feature map, the removal information being based on the second interaction being a removal interaction, the removal interaction indicating objects, features, and parts of the image to be excluded from the segmentation mask.

3. The computer-implemented method according to claim 1, further comprising: Determining an image segmentation method to be used for the segmentation, wherein the image segmentation method is based on the first interaction.

4. The computer-implemented method according to claim 1, wherein a neural network is used to combine the feature map of the image with the current probability distribution map and with the previous probability distribution map to maintain a hidden state regarding the previous probability distribution map; wherein the neural network is trained to perform the combination of the feature map of the image with the current probability distribution map and to perform the combination of the feature map with the previous probability distribution map without knowledge of which segmentation method was used to generate the current probability distribution map and without knowledge of which segmentation method was used to generate the previous probability distribution map.

5. The computer-implemented method according to claim 1, wherein the segmentation mask is generated using a classification neural network to convert the feature map combined with the current probability distribution map and the previous probability distribution map into an image form.

6. The computer-implemented method according to claim 5, wherein the classification neural network is trained to intelligently concatenate the feature map combined with the current probability distribution map and the previous probability distribution map.

7. The computer-implemented method according to claim 6, wherein the neural network is trained to integrate various image segmentation methods, the training comprising: Receiving a first interaction with a first image; Based on the first interaction, performing a first segmentation of the first image to generate a first probability distribution map; Storing the first probability distribution map using a hidden state; Receiving a second interaction with the first image; Based on the second interaction, performing a second segmentation of the first image to generate a second probability distribution map; Combine the image feature map of the first image with the second probability distribution map and the first prior probability distribution map, where the first prior probability distribution map is represented using the hidden state; And Generate an optimized segmentation mask based on the image feature map combined with the second probability distribution map and the first probability distribution map.

8. The computer-implemented method according to claim 7, wherein the training further comprises: Compare the optimized segmentation mask with the ground truth segmentation mask to determine an error; And Update the neural network based on the determined error.

9. One or more non-transitory computer-readable media having embodied thereon a plurality of executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising: Receive a first interaction with an image; Based on the first interaction with the image, perform a first segmentation of the image using a first segmentation method to generate a prior probability distribution map; Receive a second interaction with the image after the first interaction; Based on the second interaction with the image, perform a second segmentation of the image using a second segmentation method different from the first segmentation method to generate a current probability distribution map; Generate a feature map of the image using a neural network; Use the neural network to combine the feature map with the current probability distribution map to generate a first feature, and combine the feature map with the prior probability distribution map to generate a second feature, where the prior probability distribution map is represented using the hidden state of the neural network, and where the neural network is trained to perform the combination of the feature map of the image with the current probability distribution map, and to perform the combination of the feature map with the prior probability distribution map, without knowledge of which segmentation method is used to generate the current probability distribution map and the prior probability distribution map; and Generate a segmentation mask based on the concatenation of the first feature and the second feature.

10. The medium according to claim 9, wherein the method further comprises: Incorporate removal information into the feature map, where the removal information is based on the second interaction being a removal interaction that indicates objects, features, and portions of the image to be excluded from the segmentation mask.

11. The medium according to claim 9, wherein the method further comprises: Determine the first segmentation method to be used for the first segmentation, where the first segmentation method is based on the first interaction; And Determine the second segmentation method to be used for the second segmentation, where the second segmentation method is based on the second interaction.

12. The medium according to claim 9, wherein the method further comprises: Receive an indication of an image segmentation method to be used for the segmentation, where the indication is input by a user.

13. The medium according to claim 9, wherein the method further comprises: Receive a further interaction with the image; Based on the further interaction, perform a subsequent segmentation of the image to generate a subsequent probability distribution map; Use the neural network to combine the feature map with the subsequent probability distribution map to generate an updated first feature, and combine the feature map with the current probability distribution map to generate an updated second feature, where the current probability distribution map is represented using an updated hidden state of the neural network; And Generate an optimized segmentation mask based on a new concatenation of the updated first feature and the updated second feature.

14. The medium according to claim 9, wherein the neural network system is trained by: Receiving the first interaction with the first image; Based on the first interaction, performing the first segmentation of the first image to generate a first probability distribution map; Storing the first probability distribution map using a first hidden state; Receiving the second interaction with the first image; Based on the second interaction, performing the second segmentation of the first image to generate a second probability distribution map; Combining the image feature map of the first image with the second probability distribution map and a first previous probability distribution map, the first previous probability distribution map being represented using the first hidden state; And Generating an optimized segmentation mask based on the image feature map combined with the second probability distribution map and the first probability distribution map.

15. The medium according to claim 14, wherein the training further comprises: Comparing the optimized segmentation mask with a ground truth segmentation mask to determine an error; And Updating the neural network based on the determined error.

16. The medium according to claim 9, wherein the method further comprises: Generating a unified library of image segmentation methods, wherein the neural network is trained using one or more of the image segmentation methods in the image segmentation method.

17. A computing system, comprising: Components for receiving a first interaction with an image; Components for receiving a second interaction with the image after the first interaction; Components for performing segmentation of the image using a first segmentation method based on the second interaction with the image to generate a current probability distribution map; Components for combining the feature map of the image with the current probability distribution map to generate a first feature and for combining the feature map with a previous probability distribution map to generate a second feature, the previous probability distribution map being generated based on a previous segmentation of the image, the previous segmentation being performed based on the first interaction with the image and using a second segmentation method different from the first segmentation method; And Components for generating a segmentation mask based on a concatenation of the first feature and the second feature.

18. The system according to claim 17, further comprising: Components for incorporating removal information into the feature map, the removal information being a removal interaction based on the second interaction, the removal interaction indicating objects, features, and parts of the image to be excluded from the segmentation mask.

19. The system according to claim 17, further comprising: Components for determining an image segmentation method to be used for the segmentation, wherein the image segmentation method is based on the interaction.

20. The system according to claim 17, further comprising: a component for generating a unified library of image segmentation methods, wherein the neural network is trained using one or more of the image segmentation methods in the image segmentation method.

Citation Information

Patent Citations

  • A system and computer-implemented method for segmenting an image

    WO2018229490A1