Semantic Category Localization in Digital Environments

Through the combination of sequential neural networks and vector representations, the neural network is trained using image-level labels, bounding boxes and segmentation masks, and the problem of limited training data is solved, and efficient refinement and recognition of semantic categories in digital images is achieved.

CN110232689BActive Publication Date: 2025-07-22ADOBE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910034225.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-03-06
Filing Date
2019-01-14
Publication Date
2025-07-22
Estimated Expiration
2039-01-14

AI Technical Summary

Technical Problem

Traditional semantic segmentation techniques are limited by the limited availability of training data and label positioning accuracy, making it difficult to effectively identify and refine the location of millions of potential semantic categories in digital images.

Method used

By using sequential neural networks to train neural networks, first use image-level labels to generate attention maps, then use bounding boxes and segmentation masks to refine them, and combine vector representations to perform zero-sample learning to achieve refinement of semantic categories.

Benefits of technology

It effectively overcomes the limitations of training data, can identify and refine multiple semantic categories in digital images, and improves the accuracy and wide applicability of semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110232689B_ABST
    Figure CN110232689B_ABST
Patent Text Reader

Abstract

Describes semantic segmentation techniques and systems that overcome the challenge of the limited availability of training data for the potentially millions of labels that can be used to describe semantic classes in digital images. In one example, these techniques are configured to train a neural network to utilize different types of training datasets using a sequential neural network and to represent different semantic classes using a vector representation.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Semantic segmentation has made great progress with the advancement of neural networks to locate portions of a digital image corresponding to semantic classes. For example, a computing device can use machine learning to train a neural network based on training digital images and labels identifying the semantic classes presented by the digital images. Semantic classes can be used to identify specific objects included in a digital image, the feelings evoked by the digital image, etc. Once trained, the model is configured for use by the computing device to identify the locations in the digital image corresponding to the semantic classes.

[0002] However, traditional techniques require examples of labels and associated digital images for each semantic class to be trained. Thus, traditional techniques are challenged by the limited availability of training data, which is further exacerbated by the multiple labels that can be used to identify the same and similar semantic classes. For example, a traditional model trained by a computing device using machine learning for the semantic concept "human" may fail for the semantic concept "person" since the traditional model cannot recognize the correlation of the two semantic classes with each other. Summary of the Invention

[0003] Semantic segmentation techniques and systems are described that overcome the challenge of the limited availability of training data for the potentially millions of labels that can be used to describe semantic classes in digital images. In one example, labels defining semantic concepts presented by digital images used to train a neural network are converted to a vector representation. The vector representation and the corresponding digital images are then used to train the neural network to recognize the corresponding semantic concepts.

[0004] To this end, the techniques described herein are configured to train a neural network to use a sequential neural network to leverage different types of training datasets. In one example, an embedding neural network is first trained by a computing device using a first training dataset. The first training dataset includes digital images and corresponding image-level labels. Once trained, the embedding neural network is configured to generate an attention map defining a rough location of the label within the digital image.

[0005] Then, a refinement system is trained by the computing device to refine the attention map, i.e., the location of the semantic class within the digital image. For example, the refinement system can include a refinement neural network trained using bounding boxes and segmentation masks that define different levels of accuracy in identifying the semantic class. Once the embedding neural network and the refinement neural network of the refinement system are trained, the digital image segmentation system of the computing device can sequentially use these networks to generate and further refine the location of the semantic class in the input digital image. Additionally, by using the vector representation, this can also be performed for "new" semantic classes that have not been used as a basis for training the neural network by leveraging the similarity of the new semantic class to the semantic classes used to train the network.

[0006] This "Summary of the Invention" introduces some concepts in a simplified form, which will be further described in the following "Detailed Description". Thus, this "Summary of the Invention" is not intended to identify the essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Brief Description of the Drawings

[0007] The detailed description is described with reference to the accompanying drawings. Entities represented in the figures may represent one or more entities, and thus entities in the singular or plural form may be interchangeably referred to in the discussion.

[0008] Figure 1 is a diagram of an environment in an example implementation operable to employ the semantic category localization techniques described herein;

[0009] Figure 2 depicts a system in an example implementation that more particularly illustrates Figure 1 the operation of a digital image segmentation system;

[0010] Figure 3 is a flowchart depicting a process in an example implementation of generating an attention map by embedding a neural network and refining the attention map using a refinement system;

[0011] Figure 4 depicts a system showing the Figure 2 training of an embedding neural network of a digital image segmentation system for image-level labels;

[0012] Figure 5 depicts a system showing the Figure 2 training of a refinement neural network of a refinement system of a digital image segmentation system for labels with localization using a specified bounding box;

[0013] Figure 6 depicts a system showing the Figure 2 training of a refinement neural network of a refinement system of a digital image segmentation system for labels with localization using a specified segmentation mask;

[0014] Figure 7 depicts a system in an example implementation of a refinement system including a refinement neural network trained on both labels with localization defining a bounding box and labels with localization defining a segmentation mask for sequentially refining an attention map of an embedding neural network;

[0015] Figure 8 depicts an exemplary architecture of a subsequent refinement neural network of Figure 7 a segmentation network; and

[0016] Figure 9illustrates an example system that includes various components of an example device, which can be implemented as any type of computing device described and / or utilized to implement embodiments of the techniques described herein. Figure 1-8 Description and / or utilization to implement embodiments of the techniques described herein. Detailed Description

[0017] Overview

[0018] Semantic segmentation has made great progress with the advancement of neural networks. However, this progress is hindered by traditional techniques for training neural networks. For example, due to the complexity caused by semantic class overlap and lack of training data, traditional semantic segmentation techniques are limited to a few semantic classes.

[0019] For example, the labels of semantic classes can be considered to form branches in a hierarchy with complex spatial correlations, which may challenge semantic segmentation techniques. For example, for a person's face, both the fine-level annotation of "face" and the higher-level annotation of "person" are correct, and for the "clothing" area on a human body, it can also be annotated as "person" or "body". Since different semantic classes are used to describe similar and overlapping concepts, this introduces substantial challenges in training semantic segmentation techniques.

[0020] In addition, as mentioned above, there is a limited availability of training data for training neural networks to perform segmentation. This availability is further limited by the accuracy of the localization of labels within the digital images included as part of this training. For example, compared to training data items with labels that define localization by using bounding boxes, there are fewer training data items with available labels that define pixel-level localization by using segmentation masks, which are even more limited than training data items with image-level labels that do not support localization but rather refer to the digital image as a whole.

[0021] Accordingly, semantic segmentation techniques and systems are described that overcome the challenge of the limited availability of training data for describing potentially millions of labels that can be used to describe semantic classes in digital images. In one example, labels that define semantic concepts exhibited by digital images used for training neural networks are transformed into a vector representation. For example, the vector representation can be transformed from the text of the label into a word embedding by using a machine learning model, such as a two-layer neural network as part of "word2vec". The model is trained to reconstruct the linguistic context of the label and can thus be used to determine the similarity between labels by comparing the vector representations to determine "how close" these representations are to each other in the vector space.

[0022] A neural network is then trained to recognize corresponding semantic concepts using vector representations and corresponding digital images. However, as previously mentioned, there is a limited availability of training datasets with labels that reference semantic concepts. This is further limited by the accuracy of locating semantic concepts within digital images, e.g., different amounts of "supervision" from image-level to bounding boxes to segmentation masks.

[0023] Accordingly, the techniques described herein are configured to train a neural network to utilize these different types of training datasets using a sequential neural network. In one example, an embedding neural network is first trained by a computing device using a first training dataset. The first training dataset includes digital images and corresponding image-level labels. Once trained, the embedding neural network is configured to generate an attention map that defines a rough location of the label within the digital image.

[0024] Then, a refinement system is trained by the computing device to refine the attention map, i.e., the location of the semantic class within the digital image. For example, the refinement system can include an initial refinement neural network that is trained to generate an initial refinement location using localization labels that are located using corresponding bounding boxes. The refinement system can also include a subsequent refinement neural network that is trained to generate subsequent refinement locations based on the initial refinement location using localization labels that are located using corresponding segmentation masks that localize the semantic class at the pixel level.

[0025] Once the embedding neural network and the refinement neural network of the refinement system are trained, the digital image segmentation system of the computing device can sequentially use these networks to generate and further refine the location of the semantic class within the input digital image. For example, neural networks can be sequentially used that are trained from the image level to the localization from bounding boxes to the pixel level. Additionally, by using vector representations, this can also be performed for "new" semantic classes that have not been used as a basis for training a neural network by leveraging the similarity of the new semantic class to the semantic classes used to train the network, which is referred to as "zero-shot" learning in the following discussion and is not possible using traditional techniques. In this way, by the use of vector representations and the successive refinement of attention maps, the digital image segmentation system can overcome the limitations of traditional systems that involve a lack of training data to address the millions of potential labels that can be used to describe the semantic classes presented by digital images. Additionally, these techniques can be used jointly by processing multiple labels simultaneously. Further discussion of these and other examples is included in the following sections and illustrated in the corresponding figures.

[0026] In the following discussion, example environments in which the techniques described herein can be employed are described. Example processes that can be performed in the example environments as well as other environments are also described. Accordingly, the execution of the example processes is not limited to the example environments, and the example environments are not limited to the execution of the example processes.

[0027] Example Environment

[0028] Figure 1 This is an illustration of a digital media environment 100 in an example implementation that is operable to employ the semantic category localization techniques described herein. The illustrated environment 100 includes a computing device 102 that can be configured in various ways.

[0029] For example, the computing device 102 can be configured as a desktop computer, a laptop computer, a mobile device (e.g., assuming a handheld configuration such as the tablet computer or mobile phone shown), etc. Thus, the computing device 102 can range from a full-resource device with substantial memory and processor resources (e.g., a personal computer, a gaming console) to a low-resource device with limited memory and / or processing resources (e.g., a mobile device). Additionally, although a single computing device 102 is shown, the computing device 102 can represent multiple different devices, such as multiple servers used by an enterprise to perform operations "in the cloud", as Figure 9 described therein.

[0030] The computing device 102 is shown as including an image processing system 104. The image processing system 104 is implemented at least in part with the hardware of the computing device 102 to process and transform a digital image 106, which is shown as being stored in the storage device 108 of the computing device 102. Such processing includes the creation of the digital image 106, the modification of the digital image 106, and the rendering of the digital image 106 in a user interface 110 for output by, for example, a display device 112. Although shown as being implemented locally at the computing device 102, the functionality of the image processing system 104 can also be implemented in whole or in part via functionality available through a network 114 (such as part of a web service or "in the cloud").

[0031] An example of the functionality included by the image processing system 104 to process the image 106 is shown as a digital image segmentation system 116. The digital image segmentation system 116 is implemented at least in part with the hardware of the computing device (e.g., by using Figure 9processing system and computer-readable storage medium) to process a digital image 106 and a label 118 indicating a semantic category 120 to be identified in the digital image 106. The processing is performed to generate an indication 122 as an attention map 124 that describes "where" the semantic category 120 is located in the digital image 106. For example, the attention map 124 can be configured to indicate a relative probability at each pixel by using a grayscale between white ("yes" is included in the semantic category) and black (e.g., "no" is included in the semantic category). Thus, in this way, the attention map 124 can be used as a heat map to specify the location of the semantic category 120 as a segmentation mask 128 included in the digital image 126 "where". This can be used to support various digital image processing performed by the image processing system 104, including hole filling, object replacement, and other techniques that can be used to transform the digital image 106, as further described in the following sections.

[0032] Generally, the functions, features, and concepts described in the examples above and below can be used in the context of the example processes described in this section. Additionally, the functions, features, and concepts described with respect to the different figures and examples in this document can be interchanged with each other and are not limited to implementation in the context of a specific figure or process. Further, the blocks associated with the different representative processes and corresponding figures in this document can be applied and / or combined in different ways. Thus, the individual functions, features, and concepts described with respect to the different example environments, devices, components, figures, and processes in this document can be used in any suitable combination and are not limited to the specific combinations represented by the examples enumerated in this specification.

[0033] Semantic Category Localization Digital Environment

[0034] Figure 2 depicts a more detailed illustration of Figure 1 System 200 in an example implementation of the operation of a digital image segmentation system 116. Figure 3 depicts a process 300 in an example implementation of generating an attention map by embedding a neural network and refining the attention map using a refinement system. Figure 4 depicts an illustration of training Figure 2 The embedded neural network of the digital image segmentation system 116 based on image-level labels.

[0035] Figure 5 depicts an illustration of training Figure 2 The refinement neural network of the refinement system of the digital image segmentation system 116 based on localization labels as bounding boxes. Figure 6 depicts an illustration of training Figure 2An example of a refinement neural network of a refinement system of the digital image segmentation system 116, system 600. Figure 7 Depicts system 700 in an example implementation of a refinement system that includes a refinement neural network that is trained on both labels that define the location of a bounding box and labels that define the location of a segmentation mask for sequentially refining an attention map of an embedding neural network.

[0036] The following discussion describes techniques that may be implemented using the previously described systems and devices. Aspects of each process may be implemented in hardware, firmware, software, or a combination thereof. The processes are shown as a set of blocks specifying operations to be performed by one or more devices, and are not necessarily limited to the order shown for performing the operations by the corresponding blocks. In parts of the following discussion, reference will be made to Figure 1-7 .

[0037] Starting with this example, an input is received by the digital image segmentation system 116 that includes a label 118 specifying a semantic category 120 to be located in the digital image 106. The digital image 106 can take various forms, including a single "static" image, a frame of digital video or animation, etc. As previously mentioned, the semantic category 120 specified by the label 118 can also take various forms, such as specifying an object included in the digital image, a feeling evoked by a user when viewing the digital image, etc.

[0038] The label 118 (e.g., dog in the example shown) specifying the semantic category 120 is received by the vector representation conversion module 202. This module is implemented at least in part with the hardware of the computing device 102 to convert the label 118 (i.e., the text included in the label 118 to define the semantic category 120) into a vector representation 204 (block 302). Various techniques may be employed by the vector representation conversion module 202 to perform this operation, one example of which is referred to as "word2vec".

[0039] For example, the vector representation conversion module 202 can be used to generate the vector representation 204 as a word embedding through a set of machine learning models. The machine learning models are trained to construct the vector representation 204 (e.g., using a two-layer neural network) to describe the linguistic context of the word. To this end, a text corpus is used to train the machine learning models to define a vector space for the linguistic context of the text in the corpus. Then, the vector representation 204 describes the corresponding location of the semantic category 120 within this vector space.

[0040] Accordingly, the vector representations generated using this technique that shares a common context are close to each other in the vector space, e.g., based on Euclidean distance. As a result, the labels 118 that are not used to train the underlying machine learning model can be input and processed by the digital image segmentation system 116. This is because the digital image segmentation system 116 is able to determine the similarity of these labels to the labels used to train the model, which is not possible with traditional techniques. Further discussion of this feature continues in the "Implementation Example" section with respect to the "zero-shot" learning example.

[0041] Then, the vector representation 204 and the digital image 106 are received by the embedding module 206, e.g., via corresponding application programming interfaces. The embedding module 206 is configured to use the embedding neural network 208 to generate an attention map 210 (block 304) that describes the location of the semantic class 120 specified by the label 118 within the digital image 106. For example, the attention map 210 can be configured as a heatmap to indicate the relative probability at each pixel by using grayscale between white (e.g., included in the semantic class) and black (e.g., not included in the semantic class). In this way, the attention map 210 specifies the possible location of the semantic class 120 within the digital image 126. This can be used to support various digital image processing performed by the image processing system 104, including hole filling, object replacement, semantic class (e.g., object) recognition, and other techniques that can be used to transform the digital image 106.

[0042] As Figure 4 shown in the example implementation of, e.g., the embedding module 206 includes an embedding neural network 208 that is configured to train a machine learning model 402 by using a loss function 404 from the digital image 406 and the associated image-level label 408. The image-level label 408 is not located at a specific location within the digital image 406 but defines the semantic class that is included as a whole within the digital image 406. For example, in the illustrated example, the image-level label 408 "Eiffel Tower" is used to specify the object included within the digital image 406 but does not specify the location of the object within the image.

[0043] As used herein, the term "machine learning model" 402 refers to a computer representation that can be adjusted (e.g., trained) based on an input by using a loss function 404 to approximate an unknown function. Specifically, the term "machine learning model" 402 can include a model that utilizes an algorithm to learn from known data and make predictions on the known data by analyzing the known data to learn to generate an output that reflects the patterns and attributes of the known data based on the loss function 404. Thus, in this example, the machine learning model 402 performs a high-level abstraction of the data by generating data-driven predictions or decisions from known input data, namely, digital images 406 and image-level labels 408 that serve as a training data set.

[0044] As Figure 4 shown, multiple digital images 406 including various different semantic classes can be used to train the machine learning model 402 described herein. Accordingly, the machine learning model 402 learns how to identify the semantic classes and the locations of the pixels corresponding to the semantic classes in order to generate the attention map 210. Thus, "training digital images" can be used to refer to the digital images used to train the machine learning model 402. Additionally, as used herein, "training labels" can be used to refer to the labels corresponding to the semantic classes used to train the machine learning model 402.

[0045] In fact, digital images 406 with image-level labels 408 are easier to use for training than localized labels. However, digital images 406 can also be used for a larger number of semantic classes than localized labels. In one implementation, an embedding neural network 208 is trained using six million digital images 406 that have corresponding one thousand eight hundred labels for the respective semantic classes. Thus, the embedding module 206 can process the digital images 106 and the corresponding labels 118 to generate an attention map 210 for multiple image labels indicating the approximate location of the semantic class 120 (e.g., corner) in the digital image 106.

[0046] Then, the refinement system 212 uses a refinement neural network 214 trained with localized labels for the respective semantic classes to refine the location of the semantic class 120 in the attention map 210 (block 306). The localized labels can be configured in various ways to indicate which parts of the digital image correspond to the semantic class and thus also which parts do not.

[0047] As Figure 5As shown, for example, the refinement system 212 includes a refinement neural network 214 configured to train a machine learning model 502 by using a loss function 404 from a digital image 506 and a localized label 508. In this case, the localized label 508 is localized by using a bounding box 510 to identify the location of a semantic category within the digital image 506. The bounding box 510 can be defined as a rectangular region of the digital image that includes the semantic category, but can also include pixels that are not included in the semantic category. In the example shown, this allows a person to be localized to a region of the digital image 506 that does not include a laptop computer and can thus be used to achieve a higher accuracy than image-level labels.

[0048] In another example, as Figure 6 shown, the refinement system 212 also includes a refinement neural network 214 configured to train a machine learning model 602 by using a loss function 604 from a digital image 606 and a localized label 608. However, in this example, the localized label 608 is localized at the "pixel level" by using a segmentation mask 610. Thus, the segmentation mask 610 specifies for each pixel whether that pixel is part of a semantic category, such as "corner" in the example shown. As a result, the segmentation mask 610 provides a higher accuracy than Figure 5 the bounding box example.

[0049] The segmentation mask 610 for the localized label 608 provides a higher accuracy than the localized label 508 by using a bounding box, and the bounding box provides an increased accuracy of the image-level label 406 when defining the location of the semantic category with respect to the digital image. However, in practice, the training dataset for the segmentation mask 610 can be used for a smaller number of semantic categories (e.g., eighty semantic categories) than the training dataset for the bounding box 510 (e.g., seven hundred and fifty semantic categories), while the training dataset for the bounding box 510 is smaller than the training dataset for the image-level label 408 (e.g., eighteen thousand).

[0050] Thus, in one example, the refinement system is configured to employ both a refinement neural network trained using a bounding box and a refinement neural network trained using a segmentation mask to leverage different levels of accuracy and availability of semantic labels. As Figure 7 shown, for example, the system 700 includes Figure 1 the embedding module 206 and the embedding neural network 208 as described above, and accepts as input a digital image 106 and a label 118 that specify the semantic category 120 "corner".

[0051] Then, the embedding module 206 uses an embedding neural network 208 trained with image-level labels to generate an attention map 702 that defines the approximate location of semantic class 120 within the digital image 106. Then, the refinement system 212 uses an initial refinement neural network 704 and a subsequent refinement neural network 706 to refine this location.

[0052] As described with respect to Figure 5 the initial refinement neural network 704 is trained using a bounding box 710. The initial refinement neural network 708 is thus trained to refine the location of semantic class 120 within the attention map 702 to generate an initial refinement location as part of an initial refinement attention map 712 (block 308).

[0053] Then, the initial refinement attention map 712 is passed as an input to the subsequent refinement neural network 706. The subsequent refinement neural network 706 is trained using a segmentation mask 716, as described with respect to Figure 6 the segmentation mask 716 defines the pixel-level accuracy of the localization of semantic class 120 within the digital image 106. The subsequent refinement neural network 706 is thus configured to further refine the initial refinement location of the initial refinement attention map 712 to a subsequent refinement location within a subsequent refinement attention map 718 (block 310). Thus, as shown, the location of semantic class 120 "corner" defined within the attention map 702 is further sequentially refined by the initial refinement attention map 712 and the subsequent refinement attention map 718. Other examples are also contemplated where the initial or subsequent refinement neural networks 704, 706 are used alone to refine the attention map 702 output by the embedding neural network 208.

[0054] Regardless of how it is generated, the refined attention map 216 output by the refinement system 212 can then be used to indicate the refined location of the semantic class within the digital image (block 312). For example, neural networks can be used sequentially that are trained from the image level to the localization from the bounding box to the pixel level. Additionally, by using vector representations, this can also be performed for "new" semantic classes that have not been used as a basis for training neural networks by exploiting the similarity of the new semantic classes to the semantic classes used to train the networks, which is referred to as "zero-shot" learning in the implementation example below and is not possible using traditional techniques. In this way, through the use of vector representations and the successive refinement of attention maps, the digital image segmentation system can overcome the limitations of traditional systems that involve a lack of training data to address the millions of potential labels that can be used to describe the semantic classes presented by digital images. Further discussion of this example and other examples is included in the implementation example section below.

[0055] Implementation Example

[0056] As described above, the semantic category localization technique utilizes different datasets with different levels of supervision to train corresponding neural networks. For example, the first training dataset may include six million digital images with eighteen thousand labels of different semantic categories. The second training dataset is configured as bounding boxes of seven hundred and fifty different semantic categories based on the localized labels. The third training dataset is configured as segmentation masks of eighty different semantic categories based on the localized labels.

[0057] Given these datasets, the digital image segmentation system 116 adopts a semi-supervised training technique as an incremental learning framework. This framework includes three steps. First, a deep neural network is trained on the first dataset described above to learn large-scale visual semantic embeddings between digital images and eighteen thousand semantic categories. By running the embedding network in a fully convolutional manner, a rough attention (heat) map can be calculated for any given semantic category.

[0058] Next, two fully connected layers are attached to the embedding neural network as the initial refinement neural network 704 of the refinement system 212. Then, the neural network is trained at a low resolution using the second dataset of seven hundred and fifty semantic categories with bounding box annotations to refine the attention map. In one implementation, multi-task training is used to learn from the second dataset without affecting the knowledge previously learned from the first dataset.

[0059] Finally, the subsequent refinement neural network 706 is trained as a label-agnostic segmentation neural network. The label-agnostic segmentation neural network takes the initial refinement attention map 712 and the original digital image 106 as inputs and predicts a high-resolution segmentation mask as the subsequent refinement attention map 718 without significant knowledge of the semantic category 120 of interest. The segmentation network is trained using pixel-level supervision on eighty concepts of the third dataset, but can be generalized to attention maps calculated for any semantic concept.

[0060] As Figure 7 shown, the overall framework of the large-scale segmentation system implemented by the digital image segmentation system 116 includes an embedding module 206. The embedding module 206 has an embedding neural network 208 that generates an attention map 702 based on the digital image 106 and the semantic category 120 specified by the label 118. The refinement system 212 includes an initial refinement neural network 704 that generates an initial refinement attention map 712 as a "low-resolution attention map", and the "low-resolution attention map" is then refined by the subsequent refinement neural network 706 to generate a subsequent refinement attention map 718 as a segmentation mask, for example, at the pixel level.

[0061] Embedded Neural Network 208

[0062] Train an embedding neural network 208 using a first training dataset with image-level labels to learn large-scale visual semantic embeddings. The first dataset has six million images, and each image has an annotated label from a set of eighteen thousand semantic classes. The first training set is denoted as D = {(I, (w1, w2,..., w n ))}, where I is the image and w i is the word vector representation of its associated ground truth label.

[0063] Use pointwise mutual information (PMI) to generate the word vector representation for each label w in the vocabulary. PMI is a measure of the degree of association used in information theory and statistical data. Specifically, calculate the PMI matrix M, where the (i, j) element is:

[0064]

[0065] where p(w i , w j ) represents the co-occurrence probability between w i and w j , and p(w i ) and p(w j ) represent the occurrence frequencies of w i and w j respectively. The size of the matrix M is V×V, where V is the size of the label vocabulary W. The value M illustrates the co-occurrence of labels in the training corpus. Then, apply eigenvector decomposition to decompose the matrix M into M = USU T . Assume then use each row of the column-truncated submatrix W :,1:D as the word vector for the corresponding label.

[0066] Since each image is associated with multiple labels, to obtain a single vector representation for each label, calculate the weighted average over each associated label. where α = -log(p(w i )) is the inverse document frequency (idf) of the word w i . The weighted average is called the soft topic embedding.

[0067] Learn the embedding neural network 208 to map the image representation and vector representation of its associated labels into a common embedding space. In one example, each image I is extracted by a CNN feature extractor, such as a ResNet-50 extractor. After global average pooling (GAP), the visual features from the digital image 106 are then fed into a 3-layer fully connected network, where each fully connected layer is followed by a batch normalization layer and a ReLU layer. The output is the visual embedding e = embd_net(I), and it is aligned with the soft topic word vector through cosine similarity loss as follows:

[0068]

[0069] After the embedded neural network 208 is trained, the global average pooling layer is removed to obtain an attention map for a given semantic class, thereby transforming the network into a fully convolutional network. This is performed by converting the weights of the fully connected layer into 1×1 convolutional kernels and converting the batch normalization layer into a spatial batch normalization layer. After this transformation, given a digital image 106 and a vector representation 204, a dense embedded map can be obtained, where the value at each position is the similarity between the semantic class 120 and the image region around that position. Thus, the embedded map is also referred to as the attention map for that word.

[0070] Formally, the attention map for a given semantic class w can be computed as:

[0071] Att (i,j ) = <e i,j , w>

[0072] where (i, j) is the position index of the attention map.

[0073] For unseen semantic classes that were not part of the training of the image word embeddings, as long as a vector representation (i.e., word vector) w can be generated, the above equation can still be used to obtain its attention map. Thus, the embedded neural network 208 can be generalized to any arbitrary semantic class, which is not possible using traditional techniques.

[0074] Although the embedded network trained on image-level annotations is able to predict attention maps for large-scale concepts, due to the lack of annotations with spatial information, the resulting attention maps are still coarse.

[0075] Refinement System 212

[0076] To improve the quality of the attention map 210, a refinement system 212 is used to utilize a finer level of labels, namely the object bounding box labels available in the second dataset, such as using seven hundred and fifty semantic classes as part of the curated OIVG-750 dataset.

[0077] The refinement neural network 214 is attached at the end of the embedding neural network 208 and includes two convolutional layers with 1×1 kernels, followed by a sigmoid layer. By treating the eighteen thousand word embeddings as convolutional kernels, the embedding neural network 208 can output eighteen thousand rough attention maps 210. The two-layer refinement neural network 214 of the refinement system 212 then takes these eighteen thousand rough attention maps as input and learns non-linear combinations of these concepts to generate refined attention maps 216 for seven hundred and fifty semantic classes. Thus, the refinement neural network 214 considers the relationships between different semantic classes during its training.

[0078] For a given semantic class, the training signal for its attention map is a binary mask based on the ground truth bounding box, and the sigmoid cross-entropy loss is used. For better performance, the embedding neural network 208 is also fine-tuned. However, since the bounding box labels are restricted to a smaller number of semantic classes (e.g., 750) in this example, the refinement neural network 214 is only trained on these classes. To preserve the knowledge learned from the remaining semantic classes out of the eighteen thousand semantic classes, an additional matching loss is added. For example, the attention map 210 generated by the embedding neural network 208 is thresholded into a binary mask, and the sigmoid cross-entropy loss is imposed on the refined attention map 216 to match the attention map 210 from the embedding module 206. Thus, the multi-task loss function is as follows:

[0079]

[0080] where L xe (p, q) is the cross-entropy loss between the distributions p and q. Att is the attention map for a given concept, GT is the ground truth mask. B(Att) is the binary mask after thresholding the attention map. Att_ori j and Att j are the original attention map and the refined attention map respectively. ψ N is a set of indices of the top N activated original attention maps. Thus, the matching loss is only imposed on the highly activated attention maps. α is the weight that measures the loss. In one example, the value N = 800, and α = 10 -6 .

[0081] In one implementation, during training, the sigmoid cross-entropy loss is used instead of the softmax loss as in semantic segmentation to address semantic classes with masks that overlap with each other, which is particularly common for objects and their parts. For example, the mask of a face is always covered by the mask of a person. Thus, using the softmax loss would somehow prevent the mask prediction for these concepts. At the same time, in many cases, the masks of two semantic classes never overlap. To utilize this information and make the training of the attention map more discriminative, an auxiliary loss is added for those non-overlapping concepts to prevent high responses for two co-occurring concepts.

[0082] Specifically, the mask overlap rate is calculated between each pair of co-occurring concepts in the training data as follows:

[0083]

[0084] where a n (i) is the mask of the i-th concept in image n, and o(·, ·) is the overlapping region between two concepts. Note that the mask overlap rate is asymmetric.

[0085] Using the overlap rate matrix, the training example of concept i can be used as the negative training example of its non-overlapping concept j, i.e., for a specific location in the image, if the ground truth of concept i is 1, the output of concept j should be 0. To soften the constraint, the auxiliary loss is further weighted based on the overlap rate, where the weight γ is calculated as:

[0086]

[0087] The refinement system 212 can now use its vector representation 204 to predict a low-resolution attention map for any concept. To further obtain the mask of a concept with higher resolution and better boundary quality, the label-agnostic segmentation network is trained as the subsequent refinement neural network 706, which takes the original digital image 106 and the attention map as inputs and generates a segmentation mask without knowing the semantic class 120. Since the subsequent refinement neural network 706 is configured to generate a segmentation mask given the prior knowledge of the initial refined attention 712, the segmentation network can generalize to unseen concepts, even if it is fully trained on the third training dataset with eighty semantic classes in this example.

[0088] To perform mask segmentation of concepts at different scales, multiple attention maps are generated by feeding different input image sizes (e.g., 300 and 700 pixel sizes) to the embedding neural network 208. Then, the resulting attention maps are upsampled and used as additional input channels to the refinement system 212 together with the digital image 106.

[0089] To enable the refinement system 212 to focus on generating accurate masks rather than having the additional burden of predicting the presence of concepts in the predicted image, the attention map can be normalized to [0,1] to improve computational efficiency. This means that assuming the semantic classes of interest appear in the digital image during the third stage of training of the subsequent refinement neural network 706, this results in the segmentation network having improved accuracy.

[0090] Figure 8 An example architecture 800 of the subsequent refinement neural network 706 as a segmentation network is depicted. The example architecture 800 has a Y shape and includes three parts: a high-level stream 802 that uses a conventional encoder network to extract visual features to generate a two-channel low-resolution feature map as output; a low-level stream 804 that extracts a full-resolution multi-channel feature map through a shallow network module; and a boundary refinement 806 module that combines the high-level and low-level features to generate a full-resolution segmentation mask 808 as the subsequent refinement attention map 718. The boundary refinement 806 module concatenates the outputs of the low-level and high-level streams and passes it to several densely connected units, where the output of each dense unit is part of the input of any other dense unit.

[0091] The high-level stream 802 can be implemented as a deep CNN encoder network except that the input to the network has two additional channels of the attention map obtained in the attention network (e.g., one channel from an input image of size 300×300 and one channel from an input image of size 700×700). For the segmentation model, a version of Inception-V2 can be used, where the last three layers are removed, namely, pooling, linear, and softmax. The input is a 5-channel digital image 108 of 244×244 plus the initial attention map 702, and the output of the truncated Inceptions-V2 is a 1024-channel feature map of 7×7. To obtain a 14×14 feature map, dilated convolutions are used for the last two Inception modules. Finally, a convolutional layer is added to generate a 2-channel 14×14 feature map.

[0092] The low-level stream 804 is implemented as a shallow network. The input to the shallow network is the 3-channel digital image 108 and two additional channels of the initial attention map 702. Specifically, a single 7×7 convolutional layer with a stride of 1 can be used. The output of this stream is a 64-channel 224×224 feature map.

[0093] The boundary refinement 806 module takes low-level and high-level features as input and outputs the final result as a segmentation mask 808. More specifically, the size of the high-level feature map is adjusted to the original resolution (in this case 224×224) by bilinear upsampling. Then, the upsampled high-level feature map is concatenated with the low-level feature map, and then passed to a densely connected layer unit. Each dense unit includes a convolutional layer, and the output is concatenated with the input of the unit. This densely connected structure allows for more efficient training to improve the boundary quality.

[0094] Zero-Shot Learning

[0095] As mentioned before, only 18,000 semantic classes are trained using image-level supervision by using image-level labels 408 on the embedding neural network 208. However, localization labels 608 are used to train the refinement system 212, for example, at the bounding box level or the segmentation mask (pixel) level. Therefore, the difference between the lower-quality attention maps of the embedding neural network 208 and the higher-quality attention maps of the refinement system 212 (e.g., 750 semantic classes) may affect the segmentation performance of the 18,000 semantic classes.

[0096] Therefore, for semantic class q among the 18,000 semantic classes that only use image-level supervision, its nearest neighbor concept p is found in the embedding space from the semantic classes (e.g., 750 semantic classes) used to train the refinement neural network 214. Then, a linear combination of the attention maps from the two concepts is used as the input attention map 210 of the refinement system 212.

[0097] Att = θAtt q +(1 - θ)Att p

[0098] where θ is determined on the validation set.

[0099] For zero-shot learning, the embedding maps and attention maps of semantic classes are obtained as described above. To predict the segmentation of a semantic class, the same technique is used, using a linear combination of the attention maps of the semantic class and its nearest neighbor for the refinement system 212. In this way, the digital image segmentation system 116 can handle semantic classes even if these classes have not been used to train the neural network of the system, which is not possible using traditional techniques.

[0100] Example Systems and Devices

[0101] Figure 9An example system 900 is shown, which includes an example computing device 902 representative of one or more computing systems and / or devices that can implement the various techniques described herein. This is illustrated by including a digital image segmentation system 116. For example, the computing device 902 can be a server of a service provider, a device associated with a client (e.g., a client device), a system-on-chip, and / or any other suitable computing device or computing system.

[0102] The example computing device 902 shown in the figure includes a processing system 904, one or more computer-readable media 906, and one or more I / O interfaces 908 that are communicatively coupled to each other. Although not shown, the computing device 902 may also include a system bus or other data and command transmission system that couples the various components to each other. The system bus can include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any of a variety of bus architectures. Various other examples are also contemplated, such as control lines and data lines.

[0103] The processing system 904 represents the functionality to perform one or more operations using hardware. Thus, the processing system 904 is shown as including hardware elements 910 that can be configured as a processor, functional blocks, etc. This can include being implemented in hardware as an application-specific integrated circuit or other logic device formed using one or more semiconductors. The hardware elements 910 are not limited by the materials from which they are formed or the processing mechanisms used therein. For example, a processor can include (multiple) semiconductors and / or transistors (e.g., an electronic integrated circuit (IC)). In such a case, the processor-executable instructions can be electronically executable instructions.

[0104] The computer-readable storage medium 906 is shown as including a memory / storage device 912. The memory / storage device 912 represents the memory / storage capacity associated with one or more computer-readable media. The memory / storage component 912 can include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical discs, magnetic disks, etc.). The memory / storage component 912 can include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disc, etc.). As further described below, the computer-readable media 906 can be configured in various other ways.

[0105] (Multiple) input / output interfaces 908 represent functionality that allows a user to input commands and information into computing device 902 and also allows information to be presented to the user and / or other components or devices using a variety of input / output devices. Examples of input devices include keyboards, cursor control devices (e.g., mice), microphones, scanners, touch capabilities (e.g., capacitive or other sensors configured to detect physical touch), cameras (e.g., which may employ visible or non-visible wavelengths such as infrared frequencies to recognize motion as a gesture not involving touch), etc. Examples of output devices include display devices (e.g., monitors or projectors), speakers, printers, network cards, haptic response devices, etc. Thus, computing device 902 can be configured in a variety of ways, as further described below, to support user interaction.

[0106] Various technologies may be described herein in the general context of software, hardware elements, or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The terms "module", "function", and "component" as used herein generally represent software, firmware, hardware, or a combination thereof. The features of the technologies described herein are platform-independent, which means that these technologies can be implemented on a variety of commercial computing platforms having a variety of processors.

[0107] Implementations of the described modules and technologies can be stored on or transmitted via some form of computer-readable medium. Computer-readable media can include various media that can be accessed by computing device 902. By way of example and not limitation, computer-readable media can include "computer-readable storage media" and "computer-readable signal media".

[0108] "Computer-readable storage media" can refer to media and / or devices that implement persistent and / or non-transitory storage of information as opposed to mere signal transmission, carrier waves, or signals themselves. Thus, computer-readable storage media refers to non-signal-bearing media. Computer-readable storage media includes hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented in a method or technology suitable for storing information such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data. Examples of computer-readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical storage devices, hard disks, magnetic tape cartridges, tapes, magnetic disk storage devices or other magnetic storage devices or other storage devices, tangible media, or articles of manufacture suitable for storing the desired information and accessible by a computer.

[0109] "Computer-readable signal medium" can refer to a signal-bearing medium configured to transmit instructions to the hardware of computing device 902, such as via a network. A signal medium typically can include computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave, data signal, or other transmission mechanism. The signal medium also includes any information delivery medium. The term "modulated data signal" refers to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0110] As described above, hardware elements 910 and computer-readable media 906 represent modules, programmable device logic, and / or fixed device logic implemented in hardware that can be used in some embodiments to implement at least some aspects of the techniques described herein, such as executing one or more instructions. Hardware can include an integrated circuit or system-on-chip, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), and other implemented components of silicon or other hardware. In such a case, the hardware can execute as a processing device for executing program tasks defined by instructions and / or logic implemented in the hardware and as hardware for storing instructions for execution (e.g., the aforementioned computer-readable storage medium).

[0111] The foregoing combinations can also be used to implement the various techniques described herein. Thus, software, hardware, or executable modules can be implemented as one or more instructions and / or logic on some form of computer-readable storage medium and / or implemented by one or more hardware elements 910. Computing device 902 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, for example, by using the computer-readable storage medium of processing system 904 and / or hardware elements 910, the implementation of modules executable as software by computing device 902 can be implemented at least in part in hardware. The instructions and / or functions can be executable / operable by one or more articles (e.g., one or more computing devices 902 and / or processing systems 904) to implement the techniques, modules, and examples described herein.

[0112] The techniques described herein can be supported by various configurations of computing device 902 and are not limited to the specific examples of the techniques described herein. The functionality can also be implemented in whole or in part by using a distributed system, such as via platform 916 through "the cloud" 914, as described below.

[0113] Cloud 914 includes and / or represents a platform 916 for resources 918. The platform 916 abstracts the underlying functionality of the hardware (e.g., servers) and software resources of the cloud 914. The resources 918 can include applications and / or data that can be used when computer processing is performed on servers remote from the computing device 902. The resources 918 can also include services provided over the Internet and / or over a subscriber network such as a cellular or Wi-Fi network.

[0114] The platform 916 can abstract resources and functionality to connect the computing device 902 with other computing devices. The platform 916 can also be used to abstract the scaling of resources to provide a corresponding level of scaling to meet the requirements of the resources 918 implemented via the platform 916. Thus, in embodiments of interconnected devices, the implementation of the functionality described herein can be distributed across the system 900. For example, the functionality can be implemented partially on the computing device 902 and via the platform 916 that abstracts the functionality of the cloud 914.

[0115] Conclusion

[0116] Although the invention has been described in language specific to structural features and / or methodological acts, it is to be understood that the invention defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed invention.

Claims

1. A method implemented by at least one computing device in a digital media semantic category localization environment, the method comprising: Converting, by the at least one computing device, a label into a vector representation, the label defining a semantic category to be localized in a digital image; Generating, by the at least one computing device, an attention map based on the digital image and the vector representation through an embedding neural network, the attention map defining positions in the digital image corresponding to the semantic category, the embedding neural network being trained using image-level labels of the corresponding semantic category; Refining, by the at least one computing device, the positions of the semantic category in the attention map through a refinement neural network, the refinement neural network being trained using localized labels of the corresponding semantic category; And Indicating, by the at least one computing device, the refined positions of the semantic category in the digital image using the refined attention map.

2. The method according to claim 1, wherein the conversion of the vector representation uses an embedding neural network as part of machine learning.

3. The method according to claim 1, wherein the image-level label indicates the corresponding semantic category associated with the corresponding digital image as a whole used for training the embedding neural network.

4. The method according to claim 1, wherein the image-level label is not localized to a corresponding part of the digital image used for training the embedding neural network.

5. The method according to claim 1, wherein the localized labels of the semantic category are localized to corresponding parts of the digital image used for training the refinement neural network using corresponding bounding boxes.

6. The method according to claim 1, wherein the localized labels of the semantic category are localized to corresponding parts of the digital image used for training the refinement neural network using corresponding segmentation masks.

7. The method according to claim 1, wherein the number of semantic categories used for training the refinement neural network is less than the number of semantic categories used for training the embedding neural network.

8. The refinement of the refinement neural network according to claim 1, wherein the refinement comprises: Refining, by an initial refinement neural network, the positions of the semantic category in the attention map to generate initial refined positions, the initial refinement neural network being trained using localized labels localized using corresponding bounding boxes; And Refining, by a subsequent refinement neural network, the initial refined positions of the semantic category to generate subsequent refined positions, the subsequent refinement neural network being trained using localized labels localized using corresponding segmentation masks, and wherein the indication is based on the subsequent refined positions.

9. The method according to claim 1, wherein the label defining the semantic category to be localized in the digital image is not one of the image-level labels used for training the embedding neural network nor one of the localized labels used for training the refinement neural network.

10. The method according to claim 1, wherein the conversion is performed for a first tag and a second tag, and the generation, the refinement, and the indication are performed jointly based on the first tag and the second tag.

11. A system in a digital media semantic category localization environment, comprising: at least one processor; at least one memory storing instructions configured to cause the at least one processor to: convert a tag into a vector representation, the tag defining a semantic category to be localized in a digital image; implement an embedding neural network to generate an attention map based on the digital image and the vector representation, the attention map defining positions in the digital image corresponding to the semantic category, the embedding neural network being trained using image-level tags of corresponding semantic categories; and implement a refinement neural network to refine the positions of the semantic category in the attention map, the refinement neural network being trained using localization tags of the semantic category.

12. The system according to claim 11, wherein the image-level tags indicate corresponding semantic categories associated with the corresponding digital image as a whole for training the embedding neural network and not localized to corresponding parts of the digital image.

13. The system according to claim 11, wherein the localization tags of the semantic category are localized to corresponding parts of the digital image for training the refinement neural network using corresponding bounding boxes.

14. The system according to claim 11, wherein the localization tags of the semantic category are localized to corresponding parts of the digital image for training the refinement neural network using corresponding segmentation masks.

15. The system according to claim 11, wherein the instructions are configured to cause the at least one processor to: refine the positions of the semantic category in the attention map to initial refinement positions by an initial refinement neural network, the initial refinement neural network being trained using localization tags of the semantic category localized using corresponding bounding boxes; and refine the initial refinement positions of the semantic category to generate subsequent refinement positions by a subsequent refinement neural network, the subsequent refinement neural network being trained using localization tags localized using corresponding segmentation masks.

16. The system according to claim 11, wherein the tag defining the semantic category to be localized in the digital image is not one of the image-level tags for training the embedding neural network nor one of the localization tags for training the refinement neural network.

17. A method implemented by at least one computing device in a digital media semantic category localization environment, the method comprising: converting, by a processor of the computing device, a tag defining a semantic category to be localized in a digital image into a vector representation; generating, based on the digital image and the vector representation, an attention map by an embedding network as part of machine learning, the attention map defining positions in the digital image corresponding to the semantic category, the embedding network being trained using image-level tags of corresponding semantic categories; Refining the position of the semantic category in the attention map to an initial refinement position by an initial refinement neural network, the initial refinement neural network being trained using labels of the localization of the semantic category located by the corresponding bounding box; and Refining the initial refinement position of the semantic category to a subsequent refinement position by a subsequent refinement neural network, the subsequent refinement neural network being trained using labels of the localization of the semantic category located by the corresponding segmentation mask.

18. The method according to claim 17, wherein the image-level label indicates a corresponding semantic category associated with the corresponding digital image as a whole for training the embedding network and not localized to a corresponding part of the digital image.

19. The method according to claim 17, wherein the segmentation mask is a pixel-level segmentation mask.

20. The method according to claim 17, wherein: The number of the localization labels for training the subsequent refinement neural network is less than the number of the localization labels for training the initial refinement neural network; and The number of the localization labels for training the initial refinement neural network is less than the number of the image-level labels for training the embedding network.

Citation Information

Patent Citations

  • Modifying At Least One Attribute Of Image With At Least One Attribute Extracted From Another Image

    CN106560809A

  • Embedding space for images with multiple text labels

    CN106980868A