Segmentation Using Unsupervised Neural Network Training Techniques
Through self-supervised deep learning framework and constraint training neural network, the problem of difficulty in developing highly robust object transformation models in the existing technology is solved, and unsupervised component segmentation and robustness improvement are achieved.
Patent Information
- Application Number
- CN202010268572.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-08
- Filing Date
- 2020-04-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-07-07
AI Technical Summary
It is difficult to develop a model that can be robust to object transformations and deformations caused by camera poses, occlusions, changes in object appearance and changes in postures, especially when building a fully supervised model, a large amount of user-manual annotation of training data is required.
A self-supervised deep learning framework is adopted to train a neural network using a set of constraints (geometric concentration, equal variance and semantic consistency), automatically detect component fragments in the image, and generate component segmentation models in an unsupervised manner.
It realizes robustness for object changes, reduces dependence on user manual annotation data, improves the robustness and adaptability of the model, and enables component segmentation without the need for ground real-life data.
Smart Images

Figure CN111798450B_ABST
Abstract
Description
Background Art
[0001] In various scenarios such as computer vision, the main challenge in analyzing an object is to develop a model that is robust to various changes such as object transformations and deformations caused by camera pose, occlusion, changes in object appearance, and changes in pose. Parts can provide an intermediate representation of the object that is robust to various types of changes. Therefore, part-based representations can be used for various object analysis tasks such as 3D reconstruction, detection, fine-grained recognition, pose estimation, etc.
[0002] There are different types of 2D part representations, such as those using landmarks, bounding boxes, and part segmentation. However, these annotations pose many challenges, which can be computationally very expensive and / or require users to manually annotate a large amount of training data to generate a prediction model. Therefore, it may be difficult to establish a fully supervised model that can detect or segment object parts in an image. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Various techniques will be described with reference to the accompanying drawings, in which:
[0004] Figure 1 A system according to an embodiment is shown, in which a framework for self-supervised part co-segmentation is implemented;
[0005] Figure 2 A system according to an embodiment is shown, in which an isovariance constraint is adopted in part segmentation to encourage robustness to spatial variations;
[0006] Figure 3 A system according to an embodiment is shown, in which a computer system adopts a semantic consistency constraint as part of part segmentation to encourage robustness to object variations;
[0007] Figure 4 An illustrative example of a process for training a neural network in an unsupervised manner to determine one or more part segments of an image is shown;
[0008] Figure 5 An illustrative example of a process for detecting one or more part segments of one or more objects or object parts within one or more images based at least in part on a neural network trained in an unsupervised manner to infer one or more part segments is shown;
[0009] Figure 6 An example of a parallel processing unit (“PPU”) according to an embodiment is shown;
[0010] Figure 7Shows an example of a general processing cluster (“GPC”) according to one embodiment;
[0011] Figure 8 Shows an example of a memory partition unit according to one embodiment;
[0012] Figure 9 Shows an example of a streaming multiprocessor according to one embodiment; and
[0013] Figure 10 Shows a computer system in which various examples can be implemented according to one embodiment. Detailed Description
[0014] In one embodiment, the techniques described herein are implemented as systems and methods for implementing a self-supervised framework for part segmentation. In one embodiment, the model is given a set of images of the same object class. In one embodiment, the parts provide an intermediate representation of the object that is robust to camera, pose, and appearance variations. In one embodiment, a self-supervised deep learning framework for determining part segments utilizes one or more loss functions that help predict part segments based on a set of constraints for segment detection. In one embodiment, one or more loss functions are used to train a self-supervised or unsupervised neural network based on constraints from one or more of the following: geometric concentration; robustness to object variations; and semantic consistency across different object instances. In one embodiment, the part representation is robust to variations and can be used to aid high-level object understanding. In one embodiment, a set of images of a single object class can have high variability in terms of pose, object appearance, camera viewpoint, presence of multiple objects, occlusion, and other variations that make the detection of part segments challenging.
[0015] In one embodiment, an unsupervised deep learning framework for part segmentation is implemented, where the neural network is trained on part segmentation that is semantically consistent across different object types and can be applied to other types of rigid or non-rigid object classes. As used herein, an “object” can alternatively refer to an entire logical or physical object (e.g., an animal, such as a bird) or a part of an object (e.g., the head of a bird). In one embodiment, the neural network prediction provides a part segmentation that gives a richer intermediate object representation compared to landmarks or bounding boxes. In one embodiment, the neural network is used for part segmentation detection.
[0016] In one embodiment, the neural network is trained in an unsupervised manner to detect one or more boundaries of one or more objects in one or more images (e.g., boundaries determined from component segments) based at least in part on one or more loss functions that encode rules for how constraints determine component segments. In one embodiment, one or more constraints include at least one of the following: geometric density; spatial invariance; and semantic consistency. In one embodiment, the neural network is trained on a set of images of object categories and predicts component segments based on a single image of an object of the same category.
[0017] In one embodiment, Figure 1 FIG. 100 shows a system 100 according to one embodiment in which a framework for self-supervised component co-segmentation is implemented. In one embodiment, a computer system 102, such as an image processing system, includes a memory, one or more hardware sensors that capture one or more images, one or more processors that detect one or more component segments of one or more objects in one or more images based at least in part on a neural network trained in an unsupervised manner to infer one or more component segments, and one or more memories that store parameters associated with the one or more neural networks. In one embodiment, the parameters associated with the one or more neural networks are the weights of the neural network and are determined as part of training the neural network to segment images. In one embodiment, the hardware sensors that capture one or more images include video cameras, cameras, and other such devices that can be used to capture or generate video and / or still images. In one embodiment, one or more hardware sensors are used to obtain at least a portion of a set of images, where the one or more hardware sensors include a video capture device and at least a portion of the set of images is from video captured by the video capture device.
[0018] In one embodiment, the system (e.g., computer system 102) obtains a set of images {I} 104 having the same object category or classification. In one embodiment, the images are labeled as part of a particular category by one or more users. In one embodiment, the system obtains the set of images from an image repository organized by category via a network. In one embodiment, the set of images includes still images of video (e.g., recorded in real time by a camera). In one embodiment, the set of images is used to train a component segmentation network with parameters θ f of In one embodiment, the neural network is a fully convolutional neural network (FCN) with per-channel max software to generate a component response map: Where K represents the number of parts, and H×W is the image resolution. In one embodiment, the part segmentation network 106 predicts K+1 channels including K foreground channels and one background channel. In one embodiment, the channels are segments 108 as shown in Figure 1 In one embodiment, the final part segmentation result is obtained by normalizing each part map with the maximum response value in the spatial dimension and setting the background map to a constant with value T R . In one embodiment, the normalization is formulated as: In one embodiment, normalization is used to enhance weak part responses, and then the argmax function along the channel dimension is utilized to obtain part segmentation. In one embodiment, DeepLab-V2 with ResNet50 is used as the part for network segmentation.
[0019] According to at least one embodiment, ground truth segmentation annotations of images in the image set are not required or assumed. In one embodiment, the system formulates a set of constraints 110 into a differentiable loss function to encourage certain characteristics of part segmentation, including but not limited to: geometric concentration; equi-variance and semantic consistency. In one embodiment, the set of constraints encodes the attributes of good part segmentation as a set of loss functions. In one embodiment, contrary to other co-segmentation methods, the techniques described herein improve the operation of a computer system, at least because the techniques described in this disclosure do not require multiple images as input during inference at test time, but the network described in more detail above and below in the present invention can be configured to accept a single image as input during test time, enabling the trained model to be better transplanted to unseen test images. Other segmentation methods may rely on multiple images during inference to optimize segment predictions, but compared to the techniques described herein, either cannot generate part segment predictions or generate much poorer quality ones. In one embodiment, the loss function described herein is based on the combination of Figures 2 - 5 discussed ones.
[0020] In one embodiment, one or more loss functions are used to train the part segmentation network 106 and the semantic part basis 112. In one embodiment, the loss functions include a geometric concentration loss with orthonormal constraints, an equi-variance loss, and a semantic consistency loss. In one embodiment, the final objective function is a linear combination of these loss functions:
[0021]
[0022] In one embodiment, different weighting coefficients are applied to each of these loss functions. In one embodiment, (λ con , λ eqv, λ sc , λ on ) is set to (0.1, 10, 100, 0.1) or calculated based on such values and / or ratios. In one embodiment, the weights are obtained at least in part based on a coarse grid search of a subset of the dataset images.
[0023] In one embodiment, unless there is occlusion or multiple instances, geometric concentration refers to the tendency for pixels belonging to the same object part or segment to generally be spatially concentrated within the image and form connected components. In one embodiment, as a first loss function, the system applies geometric concentration on the part response map to shape the part segments. In one embodiment, the system utilizes a loss term that encourages pixels belonging to a part to be spatially close to the part center. In one embodiment, the part center of part k along axis u is calculated as: where z k = ∑ u,v R(k, u, v) is a normalization term that converts the part response map into a spatial probability distribution function. Thus, in one embodiment, the geometric concentration loss function is formulated as: And for R(k, u, v) and z k are differentiable.
[0024] As described above, the geometric loss function is a constraint imposed during training that encourages the geometric concentration of parts and attempts to minimize the variance of the spatial probability distribution function R(k, u, v) / z k . In one embodiment, the geometric loss function penalizes part responses based on the distance to the part center. In one embodiment, the system implements a loss function based on a separation (difference) loss that maximizes the distance between different landmarks. In one embodiment, such a constraint will result in part segments separated by background pixels therebetween.
[0025] Figure 2 Shows a system 200 according to one embodiment, where an isotropic variance constraint is employed in part segmentation to encourage robustness to spatial variations. In one embodiment, the system is a computer system 202 or includes a computer system 202. In one embodiment, computer system 202 includes a processor that includes one or more arithmetic logic units (ALUs) for calculating loss functions to implement the isotropic variance constraint. In one embodiment, the system obtains input image 204 from a set of input images during the training of part segmentation network 206. In one embodiment, part segmentation network 206 is as described in other parts of the present disclosure, such as in combination with Figure 1 and Figure 4As described above. In one embodiment, for each training image, the system performs one or more transformations 208, which are operations that manipulate or otherwise change the input image 204 to generate a transformed image 210. In one embodiment, the computer system applies a spatial transformation T s () and an appearance perturbation T a () from a predefined range of parameters. In one embodiment, the spatial transformation refers to a transformation of the spatial characteristics of the input image 204, such as operations like translation, rotation, reflection, etc. In one embodiment, the appearance perturbation is an operation that changes the color or gradient of the picture (e.g., converting the picture from a color image to grayscale). In one embodiment, the spatial transformation is randomly or pseudo-randomly selected from a defined range of values. In one embodiment, as described in more detail below, the transformed image 210 is used as the input to the part segmentation network 206 to determine a set of segments or boundaries of the transformed image for calculating the equivariance loss.
[0026] In one embodiment, the input image 204 is used as the input to the part segmentation network 206 (e.g., as described above, the same part segmentation network is used to identify the part segmentation of the transformed image 210) to identify a set of segmentations 212. In one embodiment, the transformation 214 is an operation applied to the segments obtained from performing segmentation on the input image 204. In one embodiment, the transformations 208 and 214 are the same transformation operations applied in the same order. In one embodiment, the transformation 214 is a strict subset of the transformation 208 and includes only those operations that affect the geometric or spatial characteristics of the image or segments of the image. In one embodiment, the transformation 208 includes a spatial transformation operation for rotating the input image 204 and a color-changing operation for converting the input image 204 to grayscale (e.g., a type of appearance perturbation), and the transformation 214 applied to the segments 212 includes the spatial transformation operation but does not include color-changing, because color-changing does not affect the position or geometric structure of the determined segment boundaries. In one embodiment, the transformation 214 is applied to the segments 212 to generate transformed segments of the input image. In one embodiment, the equivariance loss function measures how closely the transformed segments detected from the input match the segments identified from the transformed image. In one embodiment, an equivariance loss of zero is ideal to reflect the concept that spatial transformations and appearance perturbations should not affect the segmentation.
[0027] In one embodiment, the system obtains an input image I from a set of images and generates a transformed image I′ = T s (T a (I)) through the segmentation network, and obtains corresponding response maps R and R′. Part centers and calculated given a component response mapping (e.g., according to the manner for calculating the component center described above Figure 1 ). In one embodiment, the isotropic variance loss 216 is defined as where D KL () is the Kullback-Leibler divergence distance, and is the loss balance coefficient. In one embodiment, the first term corresponds to component segmentation isotropic variance, and the second term indicates component center isotropic variance. In one embodiment, a spatial transformation is applied to the image by scaling, rotating, shifting, etc. In one embodiment, the spatial transformation is a random spatial transformation that is performed by randomly or pseudo-randomly selecting values for performing a specific transformation (e.g., randomly selecting a value between -180 and 180, which represents the angle for rotating the input image in degrees). In one embodiment, performing the spatial transformation includes one or more of the following operations: scaling; rotating; shifting; projective transformation; thin plate spline transformation; and more.
[0028] Figure 3 FIG. shows a system 300 according to one embodiment, where the computer system 302 uses semantic consistency constraints as part of component segmentation to encourage robustness to object variations. In one embodiment, the computer system implementing the semantic consistency constraints is a system that trains a neural network based on the input image set 304. In one embodiment, the input image set 304 includes one or more images belonging to a shared category. In one embodiment, the input image set 304 includes image sets from two or more image categories. In one embodiment, a category refers to the classification of an image such that the images share a common region (e.g., all pictures of the same type of animal or object). According to one embodiment, information related to the semantic meaning of objects and components is embedded in the intermediate convolutional neural network features of the classification network, and the semantic consistency loss function enters into the hidden layer information (e.g., the training features of the ImageNet). In one embodiment, the computer system 302 analyzes the image to find representative feature clusters corresponding to the classification features of different component segments.
[0029] In one embodiment, assuming C-dimensional classification features identifies K representative component features In one embodiment, the system simultaneously learns the component segmentation R, and these representative component features {w k} such that the classification feature V(u, v) of the pixel (u, v) belonging to the k-th component is close to w k (e.g., ‖V(u, v) - w k ‖ 2→0). In one embodiment, the number K of components is less than the feature dimension C, and the representative component features {w k} can be regarded as spanning a K-dimensional subspace in the C-dimensional space, and the representative component features can be referred to as component basis vectors. In one embodiment, the number of representative component features is a parameter specified by the user of the computer system 302 before training.
[0030] According to one embodiment, the semantic consistency loss is shown in Figure 3 where the image I of the input image set is obtained (e.g., retrieved from a data storage location or service), and the component segmentation network 308 is used to determine the component response map R of the image I. In one embodiment, the system passes I into the classification network 306 (e.g., a pre-trained classification network) and obtains the feature maps of one or more intermediate convolutional neural network layers. In one embodiment, the feature maps are bilinearly upsampled to have the same spatial resolution as I and R, thereby generating classification features In one embodiment, the computer system 302 uses the semantic consistency constraint 310 to learn a set of component basis vectors {w k} that are globally shared across different object instances (e.g., training images). In one embodiment, the semantic consistency constraint 310 is a loss function: where is the feature vector sampled at the spatial location (u, v). In one embodiment, the semantic consistency constraint 310 is calculated at least partially based on the classification feature 312 (e.g., the classification feature V shown in Figure 3 ) and the component response map 314 (the component response map R shown in Figure 3 ).
[0031] In one embodiment, the component segmentation vector R and the component basis vectors {w k}316 are learned simultaneously using appropriate backpropagation techniques. In one embodiment, and at least to ensure that different component basis vectors do not cancel each other out, the system enforces non-negativity on both the feature V and the basis vectors {w k}. In one embodiment, a rectified linear unit (ReLU) layer is used to enforce the non-negativity of the vector values. In one embodiment, the component segmentation R is the output of the softmax function and is naturally non-negative. In one embodiment, the learned component basis is improved during the training process. In one embodiment, the representative component features, as described below in conjunction with Figure 3 , are the component basis vectors learned using backpropagation.
[0032] In one embodiment, the semantic consistency loss is regarded as a linear subspace recovery problem with respect to the embedding space provided by a feature extractor (e.g., a classification network) on a set of input images. As training progresses, in one embodiment, in the embedding space provided by the pre-trained deep features, the part bases gradually converge to the most representative directions of each part, and the recovered subspace can be described as the span of the basis vectors <w k >. In one embodiment, non-negativity ensures that the weights R(k,u,v) are interpreted as part responses. With the proposed semantic consistency loss, in one embodiment, the system has similar semantic feature embeddings in the pre-trained feature space at least partially based on the same part responses, and explicitly enforces cross-instance semantic consistency through the learned part basis vectors <w k >.
[0033] In one embodiment, the orthogonality constraint 318 is imposed as an additional constraint on the part basis vectors w k to separate them. In one embodiment, the matrix W represents a set of basis vectors, where each row is the normalized part basis vector w k / ‖w k ‖, and the orthogonality constraint is formulated as a loss function on W: where is the Frobenius norm and I is the identity matrix of size K×K. In one embodiment, the orthogonality constraint is used to reduce (e.g., minimize) the correlation between different basis vectors to obtain a more concise basis set, resulting in better part responses.
[0034] In one embodiment, the saliency constraint 320 is imposed as an additional constraint. In one embodiment, an unsupervised saliency detection method is used to suppress the background features in V such that the learned part bases do not correspond to background regions. In one embodiment, for a given image I and an unsupervised saliency map D∈[0,1] H×W , the soft mask of the feature map V is represented as D°V, where ° is the Hadamard (entry-wise) product, and then it is passed into the semantic consistency loss function. In one embodiment, the semantic consistency loss is interpreted as solving R(k,u,v)w k = 0, which can be regarded as projecting the non-significant background regions into the null space of the learned subspace spanned by {w k}. In one embodiment, the saliency constraint encapsulates the prior knowledge that parts appear on objects (not the background) and the union of the parts forms an object. In one embodiment, the saliency constraint is imposed in the feature reconstruction loss.
[0035] Figure 4 FIG. Figure 4 illustrates an illustrative example of a process 400 for training a neural network in an unsupervised manner to determine one or more component segments of an image. In one embodiment, some or all of process 400 (or any other process described herein, or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with computer-executable instructions and can be implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors by hardware, software, or a combination thereof. In one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that includes multiple computer-readable instructions executable by one or more processors. In one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In one embodiment, at least some of the computer-readable instructions usable to perform process 400 are not stored using only transient signals (e.g., propagated transient electrical or electromagnetic transmissions). A non-transitory computer-readable medium need not include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of a transient signal.
[0036] In one embodiment, Figure 4 the process is implemented by any suitable computer system (e.g., an image processing system) that includes a memory, one or more hardware sensors that capture one or more images, one or more processors that detect one or more component segments of one or more objects in one or more images based at least in part on a neural network trained in an unsupervised manner to infer one or more component segments, and one or more memories that store parameters associated with one or more neural networks. In one embodiment, the parameters associated with one or more neural networks are the weights of the neural network, which are determined to train the neural network to segment a portion of an image. In one embodiment, the hardware sensors that capture one or more images include cameras, video cameras, and other such devices usable to capture or generate video and / or still images. In one embodiment, one or more hardware sensors are used to obtain at least a portion of an image set, where one or more hardware sensors include a video capture device and at least a portion of the image set is from video captured by the video capture device.
[0037] In one embodiment, the system is configured to obtain 402 a set of images. In one embodiment, the system obtains the set of images from a data store of images stored in a compressed format such as JPEG or GIF files. In one embodiment, the set of images includes video frames parsed from a video file (e.g., a video in MPEG or MP4 format). In one embodiment, the video is a live stream of video captured directly from a camera or a multimedia stream of another computer system over a network (e.g., captured from a webcam of a remote device or from a security camera accessible to one or more processors over a network). In one embodiment, the set of images includes images of one or more categories or classifications (e.g., a set of portraits). In one embodiment, the set of images includes images of the same category that differ in terms of pose, object appearance, camera viewpoint, presence of multiple objects, occlusion, and other variations. In one embodiment, the set of images includes a set of images of humans and facial features that can differ in terms of pose, angle, etc.
[0038] The system is configured to determine 404 one or more constraints on part segmentation. In one embodiment, the constraints are represented as a loss function that encodes rules for part segmentation. In one embodiment, the loss function encodes attributes of desired or undesired part segments. In one embodiment, the loss function encodes constraints in a geometric concentration that are constraints imposed during training that encourage geometric concentration of parts and penalize part responses based on distance from the part center. In one embodiment, the loss function for isovariance measures the degree of match between transformed segments detected from the input and segments identified from the transformed image. In one embodiment, a zero isovariance loss is desirable to reflect the concept that spatial transformations and appearance perturbations should not affect segmentation. In one embodiment, semantic consistency, orthogonality constraints, and saliency constraints are constraints for unsupervised training of a neural network.
[0039] In one embodiment, the system is configured to train a neural network 406 to determine one or more segments in an image of an image set. In one embodiment, the neural network is trained to detect one or more component segments without using or requiring reference to ground truth, at least in part based on optimization of a loss function that encodes characteristics for good segmentation. In one embodiment, a loss function for training the neural network for component segmentation is specified by a user, the loss function defining attributes of a desired loss segmentation and weighting different factors for good component segmentation or detecting good component segmentation. In one embodiment, the neural network is a fully convolutional neural network (FCN). In one embodiment, the network is trained to detect segments of an image without the aid of ground truth annotations of where the segments should be located. In one embodiment, the neural network is trained on one or more rules constraining component segmentation based on one or more attributes, the one or more attributes including at least one of the following: geometric density; invariance to spatial transformations; and semantic consistency. In one embodiment, a segment or component segmentation corresponds to pixels of an image representing an object or a part thereof (e.g., an eye of a face), and other pixels of another image represent another segment of another image corresponding to the same object or a part thereof. In one embodiment, the system generates boundaries from component segments in an image. In one embodiment, a boundary refers to a set of lines or edges representing an object boundary. In one embodiment, a segment refers to a line / edge representing an object and the portion of the segment within the edge. In one embodiment, the boundary is determined by tracing the contour of one or more component segments that limit or enclose all of the one or more component segments.
[0040] In one embodiment, one or more loss functions are used to train the neural network, the loss functions including a loss function that encodes constraints on component segmentation based on geometric density, as described in connection with Figure 1 that. In one embodiment, the loss function includes, for example, a loss function that encodes an isovariance constraint on component segmentation in the manner described in connection with Figure 2 described. In one embodiment, the isovariance constraint applies spatial transformations and appearance perturbations to an image during training to verify whether the isovariance constraint holds. In one embodiment, the loss function includes, for example, a loss function that encodes a constraint on semantic constitution in the manner described in connection with Figure 3 described.
[0041] In one embodiment, the neural network is trained by generating a component response map without relying on or requiring ground truth data for the image set as I: Where K represents the number of parts, and H×W is the image resolution. In one embodiment, the part segmentation network predicts K+1 channels including K foreground channels and one background channel. In one embodiment, the final part segmentation result is obtained by normalizing each part map with the maximum response value in the spatial dimension and setting the background map to have a value of T R of a constant value.
[0042] In one embodiment, the part segmentation network has one or more loss functions and ground truth. In one embodiment, the total loss function is summarized as a linear combination of each loss function multiplied by a corresponding weight. In one embodiment, the total loss function is normalized. In one embodiment, one or more loss functions are used to train a neural network (e.g., the part segmentation network), and different weight coefficients are adjusted based on the ground truth data to further improve the result of the neural network. In one embodiment, as part of training the neural network, the system obtains a set of images and calculates one or more weighted loss functions. If the ground truth data for image segmentation is available, the system can refer to the ground truth data as part of training and adjust the weights of each loss function coefficient so that the total loss function reflects the expected result based on the ground truth data. In one embodiment, some but not all of the images have available ground truth data.
[0043] Figure 5 An illustrative example of a process 500 according to one embodiment is shown. The process 500 is used to detect one or more part segments of one or more objects or object parts in one or more images based at least in part on a neural network trained in an unsupervised manner to infer one or more part segments. In one embodiment, some or all of the process 500 (or any other process described herein, or its variations and / or combinations) are executed under the control of one or more computer systems configured with computer-executable instructions and can be implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) jointly executed on one or more processors by hardware, software, or a combination thereof. In one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program, which includes a plurality of computer-readable instructions executable by one or more processors. In one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In one embodiment, at least some of the computer-readable instructions usable to execute the process 500 are not stored using only transient signals (e.g., propagated transient electrical or electromagnetic transmissions). A non-transitory computer-readable medium does not necessarily include non-transitory data storage circuits (e.g., buffers, caches, and queues) within a transceiver of a transient signal.
[0044] In one embodiment, Figure 5 the process of is implemented by any suitable computer system, such as an image processing system including a memory, one or more hardware sensors that capture one or more images, one or more processors that detect one or more component segments of one or more objects in one or more images based at least in part on a neural network trained in an unsupervised manner to infer one or more component segments, and one or more memories that store parameters associated with one or more neural networks. In one embodiment, the parameters associated with one or more neural networks are the weights of the neural network, which are determined as part of training the neural network to segment an image. In one embodiment, the hardware sensors that capture one or more images include cameras, cameras, and other such devices that can be used to capture or generate video and / or still images. In one embodiment, one or more hardware sensors are used to obtain at least a portion of an image set, where one or more hardware sensors include a video capture device, and at least a portion of the image set comes from video captured by the video capture device. In one embodiment, the system that executes process includes one or more processors and is integrated into a vehicle (e.g., a fully automated or semi-automated vehicle), an optical device (e.g., a smart phone, smart glasses, or other embedded device).
[0045] In one embodiment, the system is configured to obtain 502 an image. In one embodiment, the image is obtained from a data storage system or a data storage system (e.g., by invoking a remote server using a network service application programming interface). In one embodiment, the image is obtained by capturing an image (e.g., using a video or camera device), thereby obtaining the resulting image. In one embodiment, the image is a different type of image (e.g., does not match a classification or category) from one or more images used to train the neural network, as described in more detail below. In one embodiment, the system is according to the system described in combination with Figures 1 - 4 and Figures 6 - 10 described system.
[0046] In one embodiment, the system is configured to provide 504 the image as input to a neural network trained in an unsupervised manner. In one embodiment, the neural network is the neural network described in combination with Figure 4 described neural network. In one embodiment, the system is configured to obtain one or more segments of the image using the neural network. In one embodiment, the neural network generates an inference, which is one or more of the above-described component segments. In one embodiment, the neural network is trained based on one or more loss function encoding rules for component segmentation, as described in more detail elsewhere in this disclosure.
[0047] Figure 6 FIG. Figure 6 shows a parallel processing unit (“PPU”) 600 according to one embodiment. In one embodiment, the PPU 600 is configured with machine-readable code that, if executed by the PPU, causes the PPU to perform some or all of the processes and techniques described throughout this disclosure. In one embodiment, the PPU 600 is implemented on one or more integrated circuit devices and is a multi-threaded processor that utilizes multi-threading as a latency hiding technique designed to parallel process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) on multiple threads. In one embodiment, a thread refers to an execution thread and is an instance of a set of instructions configured to be executed by the PPU 600. In one embodiment, the PPU 600 is a graphics processing unit (“GPU”) configured to implement a graphics rendering pipeline for processing three-dimensional (“3D”) graphic data to generate two-dimensional (“2D”) image data for display on a display device such as a liquid crystal display (LCD) device. In one embodiment, the PPU 600 is used to perform computations such as linear algebra operations and machine learning operations. Figure 6 The example parallel processor is illustrated for illustrative purposes only and should be construed as a non-limiting example of a processor architecture contemplated within the scope of this disclosure, and any suitable processor may be employed to supplement and / or replace it.
[0048] In one embodiment, the PPU 600 includes one or more arithmetic logic units (ALUs) that, if the PPU 600 is executed, cause the one or more ALUs to assist in training one or more neural networks to detect one or more component segments of one or more objects in one or more images in an unsupervised manner. In one embodiment, the same or different ALUs are also configured to use one or more neural networks during inference to detect one or more component segments of one or more objects of one or more inferences. In one embodiment, various processes and techniques described throughout this disclosure are performed in parallel across multiple processing units of the PPU. In one embodiment, during backpropagation, two or more processing units across the PPU simultaneously learn the component segmentation vectors and the component basis vectors.
[0049] In one embodiment, one or more PPUs are configured to accelerate high performance computing (“HPC”), data center, and machine learning applications. In one embodiment, PPU 600 is configured to accelerate deep learning systems and applications, including the following non-limiting examples: autonomous vehicle platforms, deep learning, high-precision voice, image, and text recognition systems, intelligent video analysis, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, among others.
[0050] In one embodiment, PPU 600 includes an input / output (“I / O”) unit 606, a front-end unit 610, a scheduler unit 612, a work distribution unit 614, a hub 616, a crossbar (“Xbar”) 620, one or more general processing clusters (“GPCs”) 618, and one or more partition units 622. In one embodiment, PPU 600 is connected to a host processor or other PPU 600 via one or more high-speed GPU interconnects 608. In one embodiment, PPU 600 is connected to a host processor or other peripheral devices via an interconnect 602. In one embodiment, PPU 600 is connected to local memory including one or more memory devices 604. In one embodiment, the local memory includes one or more dynamic random access memory (“DRAM”) devices. In one embodiment, one or more DRAM devices are configured as and / or configurable as a high bandwidth memory (“HBM”) subsystem, and multiple DRAM dies are stacked within each device.
[0051] The high-speed GPU interconnect 608 may refer to a wired-based multi-channel communication link used by the system to scale and include one or more PPU 600s in combination with one or more CPUs, support cache coherence between the PPU 600 and the CPU, and CPU master control. In one embodiment, data and / or commands are sent by the high-speed GPU interconnect 608 to other units of PPU 600 via the hub 616, or received from other units of PPU 600, such as one or more copy engines, video encoders, video decoders, power management units, and other components not explicitly shown in Figure 6 the figure.
[0052] In one embodiment, the I / O unit 606 is configured to receive data from a host processor via the system bus 602 ( Figure 6send and receive communications (e.g., commands, data) (not shown in the figure). In one embodiment, I / O unit 606 communicates directly with the host processor via system bus 602 or through one or more intermediate devices (such as a memory bridge). In one embodiment, I / O unit 606 can communicate with one or more other processors (such as one or more PPU 600) via system bus 602. In one embodiment, I / O unit 606 implements a Peripheral Component Interconnect Express (“PCIe”) interface for communication via the PCIe bus. In one embodiment, I / O unit 606 implements an interface for communicating with external devices.
[0053] In one embodiment, I / O unit 606 decodes packets received via system bus 602. In one embodiment, at least some of the packets represent commands configured to cause PPU 600 to perform various operations. In one embodiment, I / O unit 606 sends the decoded commands to various other units of PPU 600 as specified by the commands. In one embodiment, the commands are sent to front-end unit 610 and / or sent to hub 616 or other units of PPU 600, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown in FIG. 6). In one embodiment, I / O unit 606 is configured to route communications between and among the various logical units of PPU 600.
[0054] In one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload to PPU 600 for processing. In one embodiment, the workload includes instructions and data to be processed by those instructions. In one embodiment, the buffer is an area in memory that is accessible (e.g., read / write) by both the host processor and PPU 600—the host interface unit can be configured to access the buffer in the system memory connected to system bus 602 via a memory request sent via I / O unit 606 through system bus 602. In one embodiment, the host processor writes the command stream to the buffer and then sends a pointer to the start of the command stream to PPU 600, such that front-end unit 610 receives the pointer to one or more command streams and manages one or more command streams, reads commands from these streams, and forwards the commands to the various units of PPU 600.
[0055] In one embodiment, the front-end unit 610 is coupled to a scheduler unit 612 that configures the respective GPCs 618 to process tasks defined by one or more streams. In one embodiment, the scheduler unit 612 is configured to track state information related to the various tasks managed by the scheduler unit 612, where the state information may indicate which GPC 618 a task is assigned to, whether the task is active or inactive, the priority associated with the task, and so on. In one embodiment, the scheduler unit 612 manages the execution of multiple tasks on one or more GPCs 618.
[0056] In one embodiment, the scheduler unit 612 is coupled to a work distribution unit 614 that is configured to distribute tasks for execution on the GPCs 618. In one embodiment, the work distribution unit 614 tracks the number of scheduled tasks received from the scheduler unit 612, and the work distribution unit 614 manages a pool of pending tasks and a pool of active tasks for each GPC 618. In one embodiment, the pool of pending tasks includes a plurality of time slots (e.g., 32 time slots) that contain certain tasks assigned to be processed by a particular GPC 618; the pool of active tasks may include a plurality of time slots (e.g., 4 time slots) for tasks actively being processed by the GPC 618, such that when the GPC 618 finishes executing a task, the task is evicted from the GPC 618's pool of active tasks, and one of the other tasks in the pool of pending tasks is selected and scheduled for execution on the GPC 618. In one embodiment, if an active task is idle on the GPC 618, e.g., while waiting to resolve data dependencies, the active task is evicted from the GPC 618 and returned to the pool of pending tasks, and another task in the pool of pending tasks is selected and scheduled for execution on the GPC 618.
[0057] In one embodiment, the work distribution unit 614 communicates with one or more GPCs 618 via an XBar 620. In one embodiment, the XBar 620 is an interconnect network that couples many of the units of the PPU 600 to other units of the PPU 600, and can be configured to couple the work distribution unit 614 to a particular GPC 618. Although not explicitly shown, one or more other units of the PPU 600 may also be connected to the XBar 620 via a hub 616.
[0058] Tasks are managed by the scheduler unit 612 and dispatched by the work distribution unit 614 to the GPC 618. The GPC 618 is configured to process the tasks and generate results. The results can be consumed by other tasks in the GPC 618, routed to a different GPC 618 via the XBar 620, or stored in the memory 604. The results can be written to the memory 604 by the partitioning unit 622. The partitioning unit implements a memory interface for reading data from and writing data to the memory 604. The results can be sent to another PPU 604 or CPU via the high-speed GPU interconnect 608. In one embodiment, the PPU 600 includes a number U of partitioning units 622, where U is equal to the number of independent and distinct memory devices 604 coupled to the PPU 600. The partitioning unit 622 will be described in more detail in conjunction with Figure 8 more detail.
[0059] In one embodiment, the host processor executes a driver kernel that implements an application programming interface ("API") that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 600. In one embodiment, multiple compute applications are executed simultaneously by the PPU 600, and the PPU 600 provides isolation, quality of service ("QoS"), and independent address spaces for the multiple compute applications. In one embodiment, an application generates instructions (e.g., in the form of API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 600, and the driver kernel outputs the tasks to one or more streams being processed by the PPU 600. In one embodiment, each task includes one or more related thread groups, which may be referred to as warps. In one embodiment, a warp includes multiple related threads (e.g., 32 threads) that can be executed in parallel. In one embodiment, cooperating threads can refer to multiple threads that include instructions for performing a task and exchanging data via shared memory. According to one embodiment, threads and cooperating threads are described in more detail in conjunction with Figure 9 more detail.
[0060] Figure 7 illustrates, according to one embodiment, such as Figure 6GPCs such as GPC 700 shown by the PPU 600. In one embodiment, each GPC 700 includes multiple hardware units for processing tasks, and each GPC 700 includes a pipeline manager 702, a pre-raster operation unit (“PROP”) 704, a raster engine 708, a work distribution crossbar (“WDX”) 716, a memory management unit (“MMU”) 718, one or more data processing clusters (“DPC”) 706, and any suitable combination of components. It will be understood that Figure 7 the GPC 700 may include other hardware units in place of or in addition to Figure 7 the units shown.
[0061] In one embodiment, GPC 700 includes or controls one or more processors for assisting in training one or more neural networks to detect one or more part segments of one or more objects within one or more images in an unsupervised manner, and one or more memories for storing parameters related to one or more neural networks. In one embodiment, the parameters related to one or more neural networks are weights for inferring part segments of an image. In one embodiment, GPC 700 is communicatively coupled (e.g., via an electronic circuit) to a hardware sensor device that obtains an image to generate part segment predictions for it. In one embodiment, the hardware sensor includes one or more of the following: a camera; a video camera; a webcam; a smartphone; smart glasses; an embedded sensor; an infrared sensor, etc.
[0062] In one embodiment, the operation of GPC 700 is controlled by the pipeline manager 702. The pipeline manager 702 manages the configuration of one or more DPC 706 to process tasks assigned to GPC 700. In one embodiment, the pipeline manager 702 configures at least one of one or more DPC 706 to implement at least a part of a graphics rendering pipeline. In one embodiment, DPC 706 is configured to execute a vertex shader program on a programmable streaming multiprocessor (“SM”) 714. In one embodiment, the pipeline manager 702 is configured to route packets received from work distribution to appropriate logic units within GPC 700, and some packets may be routed to fixed function hardware units in PROP 704 and / or the raster engine 708, while other packets may be routed to DPC 706 to be processed by the primitive engine 712 or SM 714. In one embodiment, the pipeline manager 702 configures at least one of one or more DPC 706 to implement a neural network model and / or a computational pipeline.
[0063] In one embodiment, the PROP unit 704 is configured to route data generated by the raster engine 708 and the DPC 706 to the raster operation ("ROP") unit in the memory partition unit, as described in more detail above. In one embodiment, the PROP unit 704 is configured to perform optimizations for color blending, organize pixel data, perform address translation, and so on. In one embodiment, the raster engine 708 includes a plurality of fixed-function hardware units configured to perform various raster operations, and the raster engine 708 includes a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile merge engine, and any suitable combination thereof. In one embodiment, the setup engine receives transformed vertices and generates plane equations associated with the geometric primitives defined by these vertices; the plane equations are sent to the coarse raster engine to generate coverage information for the primitives (e.g., x, y coverage masks for tiles); the output of the coarse raster engine is sent to the culling engine, where fragments associated with primitives that fail the z-test are culled and then sent to the clipping engine, where fragments located outside the view frustum are clipped off. In one embodiment, the fragments remaining after clipping and culling are passed to the fine raster engine to generate the attributes of the pixel fragments based on the plane equations generated by the setup engine. In one embodiment, the output of the raster engine 708 includes fragments to be processed by any suitable entity (e.g., a fragment shader implemented within the DPC 706).
[0064] In one embodiment, each DPC 706 included in the GPC 700 includes an M-pipeline controller ("MPC") 710; a primitive engine 712; one or more SMs 714, and any suitable combination thereof. In one embodiment, the MPC 710 controls the operation of the DPC 706 and routes packets received from the pipeline manager 702 to the appropriate units within the DPC 706. In one embodiment, packets associated with vertices are routed to the primitive engine 712, and the primitive engine 712 is configured to fetch vertex attributes associated with the vertices from memory; conversely, packets associated with shader programs can be sent to the SM 714.
[0065] In one embodiment, the SM 714 includes programmable streaming processors configured to process tasks represented by multiple threads. In one embodiment, the SM 714 is multi-threaded and configured to execute multiple threads (e.g., 32 threads) from a particular thread group simultaneously and implement a SIMD (Single Instruction, Multiple Data) architecture, where each thread in the thread group (e.g., a warp) is configured to process different data sets based on the same instruction set. In one embodiment, all threads in the thread group execute the same instruction. In one embodiment, the SM 714 implements a SIMT (Single Instruction, Multiple Threads) architecture, where each thread in the thread group is configured to process different data sets based on the same instruction set, but where individual threads in the thread group are allowed to diverge during execution. In one embodiment, a program counter, call stack, and execution state are maintained for each warp, enabling concurrency between the warp and serial execution within the warp when threads within the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, enabling equal concurrency between all threads within and between warps. In one embodiment, an execution state is maintained for each individual thread, and threads executing the same instruction can converge and execute in parallel for better efficiency. In one embodiment, the SM 714 will be described in more detail below.
[0066] In one embodiment, the MMU 718 provides an interface between the GPC 700 and the memory partition unit, and the MMU 718 provides virtual address to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, the MMU 718 provides one or more translation lookaside buffers ("TLBs") for performing virtual address to physical address translation in memory.
[0067] Figure 8 A memory partition unit of a PPU according to one embodiment is shown. In one embodiment, the memory partition unit 800 includes a raster operation ("ROP") unit 802; a secondary ("L2") cache 804; a memory interface 806; and any suitable combination thereof. The memory interface 806 is coupled to memory. The memory interface 806 can implement 32, 64, 128, 1024-bit data buses, etc., for high-speed data transfer. In one embodiment, the PPU includes U memory interfaces 806, one memory interface 806 per pair of partition units 800, where each pair of partition units 800 is connected to a corresponding memory device. For example, the PPU can be connected to up to Y memory devices, such as high-bandwidth memory stacks or graphics double data rate, version 5, synchronous dynamic random access memory ("GDDR5 SDRAM").
[0068] In one embodiment, the memory partitioning unit 800 includes or is coupled to a memory storing executable instructions that, if executed by one or more processors (e.g., of a PPU), cause the one or more processors to train one or more neural networks to detect one or more part segments of one or more objects in one or more images in an unsupervised manner and store parameters associated with the one or more neural networks in one or more memories. In one embodiment, the neural network is trained in an unsupervised manner at least in part based on a set of input images as training data when the neural network does not know or does not require ground truth data of part segments of the input images during training. In one embodiment, a set of loss functions is used to train the neural network, and the loss functions encode rules for determining a desired part segmentation.
[0069] In one embodiment, the memory interface 806 implements an HBM2 memory interface, and Y is equal to half of U. In one embodiment, the HBM2 memory stack is on the same physical package as the PPU, saving a significant amount of power and area compared to a conventional GDDR5 SDRAM system. In one embodiment, each HBM2 stack includes four memory dies, and Y is equal to 4, while the HBM2 stack includes two 128-bit channels per die, for a total of 8 channels and a data bus width of 1024 bits.
[0070] In one embodiment, the memory supports single error correction, double error detection (“SECDED”) error correction code (“ECC”) to protect data. ECC provides higher reliability for compute applications that are sensitive to data corruption. Reliability is particularly important in large-scale cluster computing environments where the PPU processes very large data sets and / or long-running applications.
[0071] In one embodiment, the PPU implements a multi-level memory hierarchy. In one embodiment, the memory partitioning unit 800 supports unified memory to provide a single unified virtual address space for CPU and PPU memories, enabling data sharing between virtual memory systems. In one embodiment, the access frequency of the PPU to memories located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the pages more frequently. In one embodiment, the high-speed GPU interconnect 608 supports address translation services, allowing the PPU to directly access the CPU's page table and providing the PPU with full access to the CPU memory.
[0072] In one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In one embodiment, the copy engine may generate a page fault for an address that is not mapped to a page table, and then the memory partition unit 800 services the page fault by mapping the address to the page table, after which the copy engine performs the transfer. In one embodiment, memory is fixed (i.e., non-pageable) for multiple copy engine operations between multiple processors, thereby significantly reducing the available memory. In one embodiment, in the case of a hardware page fault, the address can be passed to the copy engine regardless of whether the memory page is resident, and the copy process is transparent.
[0073] According to one embodiment, data from Figure 6 the memory or other system memory is fetched by the memory partition unit 800 and stored in the L2 cache 804, which is located on-chip and shared among the various GPCs. In one embodiment, each memory partition unit 800 includes at least a portion of the L2 cache 760 associated with the corresponding memory device. In one embodiment, lower-level caches are implemented in the various units within a GPC. In one embodiment, each SM 840 may implement a level one (“L1”) cache, where the L1 cache is private memory dedicated to a particular SM 840 and fetches data from the L2 cache 804 and stores it in each L1 cache for processing in the functional units of the SM 840. In one embodiment, the L2 cache 804 is coupled to the memory interface 806 and the XBar 620.
[0074] In one embodiment, the ROP unit 802 performs graphics raster operations related to pixel colors, such as color compression, pixel blending, etc. In one embodiment, the ROP unit $$50 performs a depth test together with the raster engine 825 and receives the depth of the sample position associated with the pixel fragment from the culling engine of the raster engine 825. In one embodiment, for the sample position associated with the fragment, the depth is tested against the corresponding depth in the depth buffer. In one embodiment, if the fragment passes the depth test for the sample position, the ROP unit 802 updates the depth buffer and sends the result of the depth test to the raster engine 825. It should be understood that the number of partition units 800 may be different from the number of GPCs, and thus, in one embodiment, each ROP unit 802 may be coupled to each GPC. In one embodiment, the ROP unit 802 tracks packets received from different GPCs and determines to which GPC to route the results generated by the ROP unit 802 via the Xbar.
[0075] Figure 9illustrates a streaming multiprocessor, such as a streaming multiprocessor like Figure 7 In one embodiment, the SM 900 includes: an instruction cache 902; one or more scheduler units 904; a register file 908; one or more processing cores 910; one or more special function units ("SFU") 912; one or more load / store units ("LSU") 914; an interconnect network 916; a shared memory / L1 cache 918; and any suitable combination thereof. In one embodiment, a work distribution unit dispatches tasks to be executed on the GPCs of the PPU, and each task is assigned to a specific DPC within a GPC, and if the task is associated with a shader program, the task is assigned to the SM 900. In one embodiment, the scheduler unit 904 receives tasks from the work distribution unit and manages the instruction scheduling of one or more thread blocks assigned to the SM 900. In one embodiment, the scheduler unit 904 schedules thread blocks to execute as warps of parallel threads, where each thread block is assigned at least one warp. In one embodiment, each warp executes threads. In one embodiment, the scheduler unit 904 manages multiple different thread blocks, assigns warps to different thread blocks, and then dispatches instructions from multiple different cooperative groups to respective functional units (e.g., cores 910, SFUs 912, and LSUs 914) in each clock cycle.
[0076] In one embodiment, the SM 900 includes one or more arithmetic logic units (ALUs) that, if executed, cause one or more ALUs to assist in training one or more neural networks to detect one or more component segments of one or more objects in one or more images in an unsupervised manner. In one embodiment, the same or different ALUs are also configured to use one or more neural networks at inference time to detect one or more component segments of one or more inferred objects. In one embodiment, the various processes and techniques described throughout this disclosure are performed in parallel across multiple processing units of the SM. In one embodiment, during backpropagation, two or more processing units of the SM simultaneously learn component segmentation vectors and component basis vectors.
[0077] A cooperative group can refer to a programming model for organizing communication thread groups, which allows developers to express the granularity of communication threads, enabling a richer and more efficient parallel decomposition to be expressed. In one embodiment, a cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. In one embodiment, the application of a conventional programming model provides a single, simple construct for synchronizing cooperative threads: a barrier for all threads across thread blocks (e.g., the syncthreads() function). However, programmers typically desire to define thread groups at a finer granularity than that of a thread block and synchronize within the defined groups to achieve higher performance, design flexibility, and software reuse in the form of collective group-wide functional interfaces. Cooperative groups enable programmers to explicitly define thread groups at sub-block (i.e., as small as a single thread) and multi-block granularities and perform collective operations (such as synchronization) on the threads within the cooperative group. The programming model supports clean composition across software boundaries so that libraries and utility functions can synchronize safely within their local contexts without having to make assumptions about convergence. Cooperative group primitives support new cooperative parallel patterns, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across an entire thread block grid.
[0078] In one embodiment, the dispatch unit 906 is configured to send instructions to one or more functional units, and the scheduler unit 904 includes two dispatch units 906, the two dispatch units 906 enabling two different instructions from the same warp to be dispatched in each clock cycle. In one embodiment, each scheduler unit 904 includes a single dispatch unit 906 or additional dispatch units 906.
[0079] In one embodiment, each SM 900 includes a register file 908 that provides a set of registers for the functional units of the SM 900. In one embodiment, the register file 908 is partitioned among each of the functional units such that each functional unit is assigned a dedicated portion of the register file 908. In one embodiment, the register file 908 is partitioned among different warps executed by the SM 900, and the register file 908 provides temporary storage for operands of data paths connected to the functional units. In one embodiment, each SM 900 includes a plurality (L) of processing cores 910. In one embodiment, the SM 900 includes a large number (e.g., 128 or more) of different processing cores 910. In one embodiment, each core 910 includes fully pipelined, single-precision, double-precision, and / or mixed-precision processing units that include a floating-point arithmetic logic unit (ALU) and an integer arithmetic logic unit. In one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In one embodiment, the core 910 includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.
[0080] The tensor cores are configured to perform matrix operations according to an embodiment. In one embodiment, one or more tensor cores are included in the core 910. In one embodiment, the tensor cores are configured to perform deep learning matrix algorithms, such as convolution operations for neural network training and inference. In one embodiment, each tensor core operates on a 4×4 matrix and performs matrix multiplication and accumulation operations D = A×B + C, where A, B, C, and D are 4×4 matrices.
[0081] In one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, and the accumulation matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In one embodiment, the tensor cores operate on 16-bit floating-point input data and perform 32-bit floating-point accumulation. In one embodiment, the 16-bit floating-point multiplication requires 64 operations and produces a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point addition for 4×4×4 matrix multiplication. In one embodiment, the tensor cores are used to perform larger two-dimensional or higher-dimensional matrix operations that are composed of these smaller elements. In one embodiment, APIs such as the CUDA 9 C++ API expose specialized matrix load, matrix multiplication and accumulation, and matrix store operations to effectively use the tensor cores from CUDA-C++ programs. In one embodiment, at the CUDA level, the warp-level interface assumes a 16×16 size matrix for all 32 threads across the warp.
[0082] In one embodiment, each SM 900 includes M SFUs 912 that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, the SFU 912 includes a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, the SFU 912 includes a texture unit configured to perform texture mapping filtering operations. In one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texture pixels) from memory and sample the texture map to generate a sampled texture value for use in a shader program executed by the SM 900. In one embodiment, the texture map is stored in shared memory / L1 cache. According to one embodiment, the texture unit implements texture operations, such as filtering operations using mip-maps (e.g., texture maps of different levels of detail). In one embodiment, each SM 900 includes two texture units.
[0083] In one embodiment, each SM 900 includes N LSUs 854 that implement load and store operations between the shared memory / L1 cache 806 and the register file 908. In one embodiment, each SM 900 includes an interconnect network 916 that connects each functional unit to the register file 908 and connects the LSU 914 to the register file 908 and the shared memory / L1 cache 918. In one embodiment, the interconnect network 916 is a crossbar switch that can be configured to connect any functional unit to any register in the register file 908 and connect the LSU 914 to memory locations in the register file and the shared memory / L1 cache 918.
[0084] In one embodiment, the shared memory / L1 cache 918 is an array of on-chip memories that allows data storage and communication between the SM900 and the primitive engine and between threads in the SM 900. In one embodiment, the shared memory / L1 cache 918 includes a storage capacity of 128 KB and is in the path from the SM 900 to the partition unit. In one embodiment, the shared memory / L1 cache 918 is used for caching reads and writes. One or more of the shared memory / L1 cache 918, the L2 cache, and the memory are backing memories.
[0085] In one embodiment, a data cache and a shared memory function are combined into a single memory block, providing improved performance for both types of memory access. In one embodiment, the capacity is used as or available as a cache by programs that do not use the shared memory. For example, if the shared memory is configured to use half of the capacity, textures and load / store operations can use the remaining capacity. According to one embodiment, the integration within the shared memory / L1 cache 918 enables the shared memory / L1 cache 918 to function as a high-throughput pipeline for streaming data, while providing high-bandwidth and low-latency access to frequently reused data. When configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In one embodiment, the fixed-function graphics processing unit is bypassed, creating a simpler programming model. In one embodiment, in a general-purpose parallel computing configuration, the work distribution unit directly assigns and distributes blocks of threads to the DPCs. According to one embodiment, the threads in a block execute the same program, using a unique thread ID in the computation to ensure that each thread generates a unique result, executing the program and performing the computation using the SM 900, communicating between threads using the shared memory / L1 cache 918, and reading and writing to global memory using the LSU 914 through the shared memory / L1 cache 918 and the memory partition unit. In one embodiment, when configured for general-purpose parallel computing, the SM 900 writes commands that the scheduler unit can use to start new work on the DPCs.
[0086] In one embodiment, the PPU is included in or coupled to a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smart phone (e.g., a wireless, handheld device), a personal digital assistant (“PDA”), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, etc. In one embodiment, the PPU is included on a single semiconductor substrate. In one embodiment, the PPU is included in a system-on-chip (“SoC”) together with one or more other devices (e.g., an additional PPU, a memory, a reduced instruction set computer (“RISC”) CPU), a memory management unit (“MMU”), a digital-to-analog converter (“DAC”), etc.
[0087] In one embodiment, the PPU can be included on a graphics card that includes one or more memory devices. The graphics card can be configured to interact with a PCIe slot on the motherboard of a desktop computer. In yet another embodiment, the PPU can be an integrated graphics processing unit (“iGPU”) included in the chipset of the motherboard.
[0088] Figure 10FIG. 1000 shows a computer system 1000 in which various architectures and / or functions may be implemented according to one embodiment. In one embodiment, the computer system 1000 is configured to implement the various processes and methods described throughout this disclosure.
[0089] In one embodiment, the computer system 1000 includes at least one central processing unit 1002 that is connected to a communication bus 1010 implemented using any suitable protocol (e.g., PCI (Peripheral Component Interconnect), PCI-Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol). In one embodiment, the computer system 1000 includes a main memory 1004 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data is stored in the main memory 1004, which may take the form of random access memory (“RAM”). In one embodiment, the network interface subsystem 1022 provides an interface to other computing devices and networks for receiving data from the computer system 1000 and sending data from the computer system 1000 to other systems. In one embodiment, the computer system 1000 includes a memory storing executable instructions that, if executed by one or more processors (e.g., the central processing unit 1002), cause the one or more processors to assist in training one or more neural networks to detect one or more component segments of one or more objects in one or more images in an unsupervised manner. In one embodiment, once trained, the neural network is stored in the memory or a non-volatile storage medium. In one embodiment, the computer system 1000 uses the trained neural network to determine the component segments of an image. In one embodiment, the computer system 1000 includes a hardware sensor device that obtains an image to generate a component segment prediction therefor. In one embodiment, the hardware sensor includes one or more of the following: a camera; a video camera; a webcam; a smartphone; smart glasses; an embedded sensor; an infrared sensor, etc.
[0090] In one embodiment, the computer system 1000 includes an input device 1008, a parallel processing system 1012, and a display device 1006 that may be implemented using a conventional CRT (Cathode Ray Tube), LCD (Liquid Crystal Display), LED (Light Emitting Diode), plasma display, or other suitable display technology. In one embodiment, user input is received from an input device 1008 such as a keyboard, mouse, touchpad, microphone, etc. In one embodiment, each of the foregoing modules may be located on a single semiconductor platform to form a processing system.
[0091] In this specification, a single semiconductor platform may refer to a unique single semiconductor-based integrated circuit or chip. It should be noted that the term "single semiconductor platform" may also refer to a multi-chip module with increased connectivity that emulates on-chip operations and has been substantially improved by using conventional central processing unit ("CPU") and bus implementation methods. Of course, according to the user's requirements, the individual modules can also be placed separately or in various combinations of semiconductor platforms.
[0092] In one embodiment, a computer program in the form of machine-readable executable code or a computer control logic algorithm is stored in the main memory 1004 and / or auxiliary memory. If executed by one or more processors, the computer program enables the system 1000 to perform various functions according to one embodiment. The memory 1004, storage, and / or any other storage are possible examples of computer-readable media. Auxiliary storage may refer to any suitable storage device or system, such as a hard disk drive and / or a removable storage drive, which represent a floppy disk drive, a tape drive, a compact disc drive, a digital versatile disc ("DVD") drive, a recording device, a universal serial bus ("USB") flash drive.
[0093] In one embodiment, the architectures and / or functions of the various previous figures can be implemented in the following contexts: a central processing unit 1002; a parallel processing system 1012; an integrated circuit having at least some of the capabilities of both the central processing unit 1002 and the parallel processing system 1012; a chipset (e.g., a set of integrated circuits designed to work and be sold as a unit for performing related functions, etc.), and any suitable combination of integrated circuits.
[0094] In one embodiment, the architectures and / or functions of the various previous figures are implemented in the context of a general-purpose computer system, a circuit board system, a game console system dedicated to entertainment purposes, an application-specific system, etc. In one embodiment, the computer system 1000 can take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smart phone (e.g., a wireless, handheld device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, a mobile phone device, a television, a workstation, a game console, an embedded system, and / or any other type of logic.
[0095] In one embodiment, parallel processing system 1012 includes a plurality of PPUs 1014 and associated memory 1016. In one embodiment, the PPUs are connected to a host processor or other peripheral devices via interconnect 1018 and switch 1020 or multiplexer. In one embodiment, the parallel processing system 1012 distributes computing tasks across the PPUs 1014, which may be parallel, for example, as part of the distribution of computing tasks across multiple GPU thread blocks. In one embodiment, the memory is shared and accessed (e.g., for read and / or write access) across some or all of the PPUs 1014, but such shared memory may incur a performance penalty as compared to using local memory and registers resident in the PPU. In one embodiment, the operations of the PPUs 1014 are synchronized by using a command such as _syncthreads(), which requires that all threads in a block (e.g., executing across multiple PPUs 1014) must reach a certain execution point in the code before continuing execution.
[0096] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it is obvious that various modifications and changes can be made thereto without departing from the broader spirit and scope of the invention as set forth in the claims.
[0097] Other variations are within the spirit of the disclosure. Thus, although the disclosed techniques are susceptible to various modifications and alternative constructions, certain illustrated embodiments have been shown in the drawings and have been described in detail above. However, it should be understood that it is not intended to limit the invention to the one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalent forms falling within the spirit and scope of the invention as defined by the appended claims.
[0098] In the context of describing the disclosed embodiments (especially in the context of the appended claims), the use of the terms "a", "an", and "the" and similar referents should be construed to cover the singular and plural, unless otherwise specified herein or clearly contradicted by the context. Unless otherwise indicated, the terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (i.e., meaning "including but not limited to"). The term "connected" (when unmodified and referring to a physical connection) should be understood to mean fully or partially contained in, attached to, or joined together, even if there is something intervening. References to numerical ranges herein are only intended to be used as a shorthand method, and unless otherwise specified herein, each individual value falling within the range is separately referred to as if individually recited herein and is incorporated into the specification. The use of the term "set" (e.g., "set of items" or "subset"), unless the context otherwise indicates or contradicts, should be construed as a non-empty set containing one or more members. Additionally, unless the context otherwise indicates or contradicts, the term "subset" of a corresponding set is not necessarily meant to denote a proper subset of the corresponding set, but the subset and the corresponding set may be equal.
[0099] Conjunctive language, such as phrases in the form of "at least one of A, B, and C" or "at least one of A, B, and C", unless otherwise explicitly stated or clearly contradicted by the context, can be understood in context as generally used to present items, clauses, etc., which can be A or B or C, or any non-empty subset of the set of A and B and C. For example, in an exemplary example of a set with three members, the conjunctive phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctive language generally is not intended to imply that certain embodiments require at least one of A, at least one of B, and at least one of C, each of which is used to present. Additionally, unless otherwise specified or contradicted by the context, the term "plural" indicates a plural state (e.g., "a plurality of items" means a plural number of items). The number of items in a "plurality" is at least two, but can be more when explicitly or implicitly indicated by the context. Further, unless otherwise specified or otherwise apparent from the context, the phrase "based on" means "at least partially based on" rather than "only based on".
[0100] The operations of the processes described herein may be performed in any suitable order, unless otherwise indicated herein or clearly contradicted by the context. In one embodiment, a process such as those described herein (or variations and / or combinations thereof) is performed under the control of one or more computer systems by hardware or a combination thereof, the one or more computer systems being configured with executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) to be executed jointly on one or more processors. In one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program, the computer program including a plurality of instructions executable by one or more processors. In one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that does not include transitory signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of a transitory signal. In one embodiment, the code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory storing executable instructions) on which the executable instructions, when executed by one or more processors of a computer system (e.g., as a result of being executed), cause the computer system to perform the operations described herein. In one embodiment, the set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more individual non-transitory storage media among the plurality of non-transitory computer-readable storage media lack all of the code, while the plurality of non-transitory computer-readable storage media together store all of the code. In one embodiment, the executable instructions are executed such that different instructions are executed by different processors - for example, the non-transitory computer-readable storage media stores instructions, and the main CPU executes some instructions while the graphics processing unit executes other instructions. In one embodiment, different components of the computer system have independent processors, and different processors execute different subsets of the instructions.
[0101] Thus, in one embodiment, a computer system is configured to implement one or more services that perform the operations of the processes described herein either individually or jointly, and such a computer system is configured with suitable hardware and / or software enabling the execution of the operations. Additionally, a computer system implementing an embodiment of the present disclosure is a single device, and in another embodiment, is a distributed computer system that includes a plurality of devices operating in different ways such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.
[0102] Unless otherwise required, the use of any and all examples or exemplary language (e.g., "such as") provided herein is for illustrative purposes only to better clarify embodiments of the present invention and does not limit the scope of the present invention. Language in this specification should not be construed as indicating that any non-claimed element is essential for practicing the present invention.
[0103] Embodiments of the present disclosure are described herein, including the best mode known to the inventors for practicing the invention. Variations of these embodiments will be apparent to those of ordinary skill in the art after reading the foregoing description. The inventors expect those skilled in the art to appropriately employ such variations, and the inventors intend for the embodiments of the present disclosure to be practiced in a manner different from that specifically described herein. Accordingly, the scope of the present disclosure includes all modifications and equivalents of the subject matter recited in the appended claims as permitted by applicable law. In addition, unless otherwise indicated herein or clearly contradicted by context, any combination of the above elements in all possible variations thereof is included within the scope of the present disclosure.
[0104] All references cited herein, including publications, patent applications, and patents, are incorporated herein by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0105] In the specification and claims, the terms "coupled" and "connected" and their derivatives may be used. It should be understood that these terms are not necessarily intended as synonyms for each other. Rather, in a particular example, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other.
[0106] Unless otherwise specified, it should be understood that throughout the specification, terms such as "processing," "computational processing," "computing," "determining," etc. refer to actions and / or processes of a computer or computing system, or similar electronic computing device, that manipulate and / or transform data represented as physical (e.g., electronic) quantities within the registers and / or memory of the computing system into other data similarly represented as physical quantities within the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0107] In a similar manner, the term "processor" can refer to any device or portion of a device that processes electronic data from registers and / or memory and converts the electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a central processing unit (CPU) or a graphics processing unit (GPU). A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process can refer to multiple processes that execute instructions sequentially or in parallel, continuously or intermittently. The terms "system" and "method" are used interchangeably herein to the extent that the system can embody one or more methods and the method can be considered a system.
[0108] In this document, reference may be made to obtaining, acquiring, receiving analog or digital data or inputting it into a subsystem, computer system, or computer-implemented machine. The process of obtaining, acquiring, receiving, or inputting analog and digital data can be accomplished in a variety of ways, such as by receiving the data as a parameter of a function call or a call to an application programming interface. In some embodiments, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting the data via a serial or parallel interface. In another embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transferring the data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be accomplished by transmitting the data as an input or output parameter of a function call, an application programming interface, or a parameter of an interprocess communication mechanism.
[0109] Although the above discussion sets forth example embodiments of the described techniques, other architectures can be used to implement the described functionality and are intended to be within the scope of the present disclosure. Additionally, although specific assignments of responsibilities were defined above for discussion purposes, the various functions and responsibilities may be assigned and divided in different ways depending on the circumstances.
[0110] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
[0111] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it is obvious that various modifications and changes can be made thereto without departing from the broader spirit and scope of the invention as set forth in the claims.
Claims
1. A processor, comprising: One or more circuits configured to generate one or more segmentations of one or more objects within one or more images based at least in part on one or more transformations of the one or more objects within the one or more images via one or more neural networks.
2. The processor according to claim 1, wherein the one or more neural networks are trained in an unsupervised manner based at least in part on one or more loss functions that encode rules for constraining how segments are determined.
3. The processor according to claim 2, wherein the one or more loss functions encode one or more constraints on how segments are generated, including at least one of the following: Geometric concentration constraint; Spatial invariance constraint; and Semantic consistency constraint.
4. The processor according to claim 1, wherein, The one or more neural networks are trained on a set of images without ground truth data indicating whether the segmentations have been correctly identified.
5. The processor according to claim 1, wherein, The one or more neural networks are trained in an unsupervised manner on images from any one of at least two classes of images.
6. The processor according to claim 1, wherein the number of segments to be detected in an image is a parameter specified by a user for training the one or more neural networks.
7. A system, comprising: One or more processors configured to generate one or more segmentations of one or more objects within one or more images based at least in part on one or more transformations of the one or more objects within the one or more images via one or more neural networks; And One or more memories for storing parameters associated with the one or more neural networks.
8. The system according to claim 7, wherein The one or more neural networks are trained based on a set of differentiable loss functions that encode constraints on how the one or more segments are generated.
9. The system according to claim 8, wherein The one or more neural networks are trained to detect a fixed number of segments in one or more images.
10. The system according to claim 7, wherein, The one or more neural networks are trained in a set of images sharing a class with the one or more images.
11. The system according to claim 10, wherein, The one or more images are a single image.
12. The system according to claim 7, wherein the one or more neural networks are trained without ground truth data indicating whether the segmentations have been correctly identified.
13. The system according to claim 7, wherein The one or more neural networks are fully convolutional neural networks.
14. An image recognition system, comprising: One or more hardware sensors for capturing one or more images; One or more processors configured to generate one or more segmentations of one or more objects within one or more images based at least in part on one or more transformations of the one or more objects within the one or more images via one or more neural networks; And One or more memories for storing parameters associated with the one or more neural networks.
15. The image recognition system according to claim 14, wherein, The one or more hardware sensors include video capture devices, and at least a portion of the one or more images is from a video captured by the video capture devices.
16. The image recognition system according to claim 14, wherein: The image recognition system further includes one or more data storage systems for storing identities and metadata related to segments of the identities; and The one or more processors are configured to determine whether an identity is recognized in one or more images based on the one or more segments that have been generated.
17. The image recognition system according to claim 16, wherein, The metadata includes at least one of the following: skin color; hair color; height; weight; and facial features.
18. The image recognition system according to claim 16, wherein, The one or more processors are used to determine whether the identity is a person.
19. The image recognition system according to claim 14, wherein, The image recognition system includes a first computer system that includes the one or more hardware sensors for communicating via a network with a second computer system that includes the one or more processors.
20. The image recognition system according to claim 14, wherein, The one or more segments include background segments.
21. A processor, comprising: One or more arithmetic logic units (ALUs) for assisting in training one or more neural networks to generate one or more segments of one or more objects based at least in part on one or more transformations of the one or more objects within one or more images.
22. The processor according to claim 21, wherein the one or more ALUs for assisting in training the one or more neural networks are the one or more ALUs for assisting in training the one or more neural networks to detect the one or more segments on a set of training images lacking ground truth data annotations.
23. The processor according to claim 21, wherein the one or more neural networks are neural networks that will be trained based at least in part on one or more rules that constrain segment generation according to one or more attributes, where the one or more attributes include at least one of the following: Geometric density; Invariance of spatial transformation; and Semantic consistency.
24. The processor according to claim 21, wherein, The one or more neural networks will be trained on a set of images of a class and predict segments in another image of the class.
25. The processor according to claim 21, wherein, The processor is communicatively coupled to an image processing system including one or more hardware sensors to generate the one or more images.
26. The processor according to claim 21, wherein, The first segment of the first image corresponds to the pixels of the first image associated with the object, and the second segment of the second image corresponds to the pixels of the second image associated with the object.
27. A system, comprising: One or more processors for assisting in training one or more neural networks to generate one or more segments of one or more objects in an unsupervised manner based at least in part on one or more transformations of the one or more objects within one or more images; and One or more memories for storing parameters associated with the one or more neural networks.
28. The system according to claim 27, wherein, The one or more processors are used to assist in training the one or more neural networks based at least in part on a loss function that encodes one or more constraints on segment generation.
29. The system according to claim 28, wherein The one or more constraints include a constraint on geometric concentration that regulates the neural network to minimize the variance of the spatial probability distribution.
30. The system according to claim 28, wherein, The one or more constraints include a constraint on equal variance, and the loss function is calculated based on instructions which, if executed by the one or more processors, cause the one or more processors to: Detect a first segmentation of an image; Apply one or more transformation operations to the image to generate a transformed image; Detect a second segmentation of the transformed image; Apply at least a portion of the one or more transformation operations to the image to generate a transformed first segmentation; and Compare the transformed first segmentation with the second segmentation.
31. The system according to claim 30, wherein: The one or more transformation operations include a spatial transformation and an appearance perturbation; and The portion of the one or more transformation operations includes the spatial transformation but lacks the appearance perturbation.
32. The system according to claim 30, wherein the transformed first segmentation and the second segmentation are compared by at least calculating the Kullback-Leibler divergence distance.
33. The system according to claim 27, wherein The one or more processors are further configured to determine one or more boundaries from the one or more segmentations.
34. A machine-readable medium storing an instruction set which, if executed by one or more processors, causes the one or more processors to at least: Train one or more neural networks to generate one or more segmentations of one or more objects based at least in part on one or more transformations of the one or more objects within one or more images in an unsupervised manner; and Store parameters associated with the one or more neural networks in one or more memories.
35. The machine-readable medium according to claim 34, wherein, The instructions for training the one or more neural networks in the unsupervised manner include instructions which, if executed by the one or more processors, cause the one or more processors to at least perform the following operations: Obtain an image from the one or more images; Obtain a feature map of the image based at least in part on a classification network; Obtain a component response map of the image based at least in part on a fully convolutional network; and Determine a loss function to backpropagate based at least in part on the feature map and the component response map.
36. The machine-readable medium according to claim 35, wherein, The feature map is bilinearly upsampled to have a spatial resolution matching that of the image.
37. The machine-readable medium according to claim 35, wherein, The instructions for training the one or more neural networks in an unsupervised manner include instructions which, if executed by the one or more processors, cause the one or more processors to at least learn a set of component basis vectors shared on at least a portion of one or more images, wherein the loss function is further determined at least in part based on the set of component basis vectors.
38. The machine-readable medium according to claim 37, wherein, The instructions for training the one or more neural networks in the unsupervised manner include instructions which, if executed by the one or more processors, cause the one or more processors to at least apply an orthogonality constraint to reduce the correlation between vectors in the set of component basis vectors.
39. The machine-readable medium according to claim 34, wherein, Instructions for training the one or more neural networks in the unsupervised manner include instructions that, if executed by the one or more processors, cause the one or more processors to at least apply a saliency constraint to reduce the correlation between the learned part basis and the image background.
40. The machine-readable medium according to claim 34, wherein, The parameters associated with the one or more neural networks include one or more weights determined as part of training the one or more neural networks, which are used to detect the one or more segmentations.