Image classifiers that include non-injective transformations

JP7926827B2Active Publication Date: 2026-09-30ROBERT BOSCH GMBH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021110560
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-03
Filing Date
2021-07-02
Publication Date
2026-09-30
Estimated Expiration
2041-07-02

Smart Images

  • Figure 0007926827000031
    Figure 0007926827000031
  • Figure 0007926827000032
    Figure 0007926827000032
  • Figure 0007926827000033
    Figure 0007926827000033
Patent Text Reader

Abstract

To provide a computer-implemented method of training an image classifier, a system, and a computer-readable medium.SOLUTION: Provided is a method of using any combination of labelled and / or unlabelled training images. An image classifier IC420 comprises a set of transformations between respective transformation inputs TI and transformation outputs TO. An inverse model is defined in which for a deterministic, non-injective transformation of the image classifier, its inverse is approximated by a stochastic inverse transformation. During training, for a given training image, a likelihood contribution for this transformation is determined based on a probability of its transformation inputs being generated by the stochastic inverse transformation given its transformation outputs. This likelihood contribution is used to determine a log-likelihood for the training image to be maximized or its label if the training image is labelled, based on the fact that the model parameters are optimized.SELECTED DRAWING: Figure 4e
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Field of Invention The present invention relates to a computer-implemented method for training an image classifier, and a corresponding system. The present invention also relates to a computer-implemented method for using a trained image classifier for image classification and / or image generation, and a corresponding system. The present invention further relates to a computer-readable medium containing instructions for performing one of the above methods and / or model data representing a trained image classifier.

[0002] Background of the Invention A key task in many computer-controlled systems is image classification. In image classification, an input image is classified into a class from a given set of classes. Image classification tasks arise, for example, in the control systems of (semi-)autonomous vehicles, where useful information about traffic conditions during vehicle operation is extracted. Image classification is also applied in manufacturing, healthcare, and other fields.

[0003] Machine learning techniques have proven to be very suitable for many practical problems in image classification. When using machine learning, classification is performed by applying a parameterized model to input images. The model is trained to learn the values ​​of a set of parameters to produce the best classification result. This training is typically supervised training, based on a training dataset of labeled training images labeled with each training class from a given set of classes.

[0004] In particular, machine learning techniques in the field of deep learning have been found to work well in practical applications. Generally, image classifiers can take the form of a categorical distribution p(y|x)=Cat(y|π(x)), where x is the input image and y is the label. When using deep learning, x is deterministically mapped to the class probability π, typically using a convolutional neural network consisting of convolutional layers, pooling layers, and / or tightly coupled layers.

[0005] Machine learning, particularly deep learning, can yield excellent results, but it also presents several practical challenges. Achieving satisfactory performance typically requires a large amount of training data. Furthermore, this training data often needs to be manually labeled. In many cases, obtaining sufficient training data is difficult, sometimes even dangerous (e.g., in autonomous driving applications), and sometimes simply impossible. Moreover, manually labeling all training data is often prohibitively expensive.

[0006] Summary of the Invention An image classifier is desired that can be trained in a semi-supervised manner using not only labeled training images with corresponding training classes, but also unlabeled training images with unknown classes. Furthermore, it is desired that the image classifier can generate additional images similar to those in the training dataset, for example, using a class-conditional approach. The ability to generate additional images with class conditions is particularly desirable in situations where labeled training data is scarce. This is because it allows for the generation of more samples, for example, in corner cases where data collection is difficult using other methods (e.g., dangerous traffic situations).

[0007] According to a first aspect of the present invention, a computer-implemented method and a corresponding system are provided for training an image classifier, as defined in claims 1 and 13, respectively. According to a further aspect of the present invention, a computer-implemented method and a corresponding system are provided for classifying and / or generating images using such an image classifier, as defined in claims 9 and 14, respectively. According to one aspect of the present invention, a computer-readable medium is described as defined in claim 15.

[0008] As the inventors have discovered, both training an image classifier using a semi-supervised method and using the image classifier to generate additional inputs can be achieved by defining an inverse model of the image classifier, in other words, by defining a model that maps the output class y of the image classifier to an input image x.

[0009] As will be explained in more detail below, given such an inverse model, we can define a joint probability distribution p(x,y) of the input images and the classes determined by the image classifier. Thus, we can train the image classifier by using labeled training images and maximizing the log-likelihood of the labeled training images that arise according to this joint probability distribution. Interestingly, however, in the inverse model, it is also possible to define a probability distribution p(x) of the images generated by the inverse model. Thus, we can train the inverse model using unlabeled training images by maximizing the log-likelihood of the unlabeled training images that arise according to the probability distribution p(x), and thereby indirectly train the image classifier itself. Thus, it becomes possible to train the image classifier based on any combination of labeled and / or unlabeled training images, and in particular, semi-supervised learning of the image classifier becomes possible.

[0010] Furthermore, as will be explained in more detail below, by defining an inverse model of the image classifier, the image classifier can be used as a class-conditional generative model, for example, according to a probability distribution p(x|y). It is also possible to generate class-independent images by sampling from the probability distribution p(x). The inverse model corresponds to the image classification model and therefore shares many of its trainable parameters with the forward mapping from image to class, resulting in the generation of highly representative images.

[0011] Separately, in the classification direction, robustness to perturbations is expected to improve by using a classifier that represents not only the conditional probability distribution p(y|x) but also the joint probability distribution p(x,y). Furthermore, it is possible to train the generative model more efficiently than training each model individually and reduce overfitting. Class-conditional generators are more efficient than, for example, rejection sampling type approaches of class-conditional generation because they can directly generate images from a given class. Images can also be generated according to a given vector of class probabilities, or more generally, according to values ​​specified in the inner layers of the model, rather than a given class, thereby allowing for customization of image generation, for example, generating images close to the decision boundary and / or images combining multiple characteristics.

[0012] In the context of generative modeling, it is known to use models that allow for the definition of inverse models, for example, using the so-called "normalizing flows" framework. In this framework, the probability density is represented by a differentiable bijection with a differentiable inverse transform. Such a bijection function contributes to the log-likelihood of the probability distribution in the form of its Jacobian determinant.

[0013] Unfortunately, the normalization flow framework is not applicable to image classifiers. The bijective nature of the transformations used in the normalization flow limits the capabilities of image classifiers, for example, their ability to change dimensions as needed for image classification. In particular, image classifiers are not bijective functions; that is, they can map multiple different input images to the same class. An image classifier can be thought of as consisting of multiple transformations between each transformation input and transformation output. Some of these transformations may be bijective functions. However, image classifiers typically include one or more nonjective transformations and therefore do not have a deterministic inverse transformation. This includes, for example, dimensionality reduction layers such as max pooling and tightly coupled layers commonly used in image classifiers. Therefore, these layers cannot be modeled with the normalization flow.

[0014] Interestingly, the present inventors have found that it is still possible to define an inverse model of an image classifier by approximating the inverse transformation of the non-injective transformation f:X→Z occurring in the image classifier by a probabilistic inverse transformation g:Z→X. Therefore, g does not define a function, but instead can define the conditional probability distribution p(x|z) of the transformation input x given the transformation output z. This inverse transformation can be an inverse transformation in the sense that, given the transformation output z, the support of the defined probability distribution is limited to the transformation inputs x that are mapped to the transformation output. For example, if p(x│z)≠0, then f(x)=z (this property can be slightly relaxed, for example, by considering only probabilities p(x│z) exceeding a specific threshold and / or requiring that f(x) approximates z within a predetermined tolerance). Therefore, the inverse transformation can be a right inverse in the sense that applying the inverse transformation first, then the forward transformation, corresponds to the identity mapping.

[0015] As further described below, various transformations used in image classification can be implemented using deterministic transformations involving such probabilistic inverse transformations. For example, an image classifier may include max pooling, ReLU, and / or dimension reduction fully-connected layers implemented using such transformations.

[0016] By combining the (approximated) inverse transformations of the respective transformations, a probabilistic inverse model of an image classifier can be obtained. The image classifier can be a deterministic function y=f(x), or more generally, can be a probability distribution p(y|x) over classes given an input image. The inverse model can be an inverse model in the sense that f(x)=y when p(x│y)≠0, or at least p(y│x)≠0. In this sense, it can be regarded as a right inverse of the image classifier, or at least an approximation of a right inverse.

[0017] The inventors have found a particularly attractive method for calculating the log-likelihoods of labeled and unlabeled training images using an inverse model. Specifically, they discovered that the difference in logarithmic marginal density between the transform output logp(z) and the transform input logp(x) of a deterministic non-injective transform f(x) can be approximated by the likelihood contribution based on the probability p(x|z) of the transform input x generated by the stochastic inverse transform, given the transform output z. Using this difference in logarithmic marginal density, both the combined log-likelihood logp(x,y) of the labeled training image and the log-likelihood p(x) of the unlabeled training image can then be calculated.

[0018] Therefore, given training images, the log-likelihood of those images can be evaluated in a Monte Carlo manner by applying a classifier to obtain the input and output of a transformation f(x), using the input and output to compute the likelihood contribution, and using the likelihood contribution to compute the log-likelihood of the training images (with and without associated labels). The image classifier can be trained by optimizing its parameters to maximize the log-likelihood of labeled and / or unlabeled training examples, thereby learning the distribution and available labels of the training input images. Thus, any combination of unsupervised, semi-supervised, or fully supervised training can be performed using the provided method.

[0019] In fact, the inventors have discovered that the likelihood contribution term can be defined not only for deterministic non-injective transformations but also for other types of transformations.

[0020] The inventors have identified four types of transformations that may be usefully used in image classifiers: 1) when both the transformation and its inverse are deterministic (referred to herein as “bijective transformations”), 2) when both the transformation and its inverse are stochastic (referred to herein as “stochastic transformations”), 3) when the transformation is deterministic and its inverse is stochastic (referred to herein as “inferential surjective transformations” or “inferential surjective”), or 4) when the transformation is stochastic and its inverse is deterministic (referred to herein as “generative surjective transformations” or “generative surjective”). As mentioned above, deterministic nonjective transformations are of the second type because they do not have a deterministic inverse. The term “surjective” is used herein to mean nonjective, in contrast to bijective. A single inferential / generative surjective may be referred to as a “slab”. A configuration of one or more bijective, surjective, and / or stochastic transformations may be referred to as a “flow”.

[0021] The inventors have discovered that for each of these four types of transformations, the difference in the log marginal probabilities of the transformation input and the transformation output can be approximated as the sum of likelihood contribution terms and marginal relaxation terms. For inferential surjectives and bijectives, the marginal relaxation term can be zero. For stochastic transformations and generative surjectives, a non-zero marginal relaxation term can be defined to represent the gap in the lower bound of the evidence. Thus, the log-likelihood of a training image can be obtained by combining the likelihood contributions of each transformation in the set of transformations, for example, by adding the likelihood contributions expressed as log differences (or by multiplying by likelihood contributions that do not use logarithms). The result can be an approximation of the degree of approximation given by the marginal relaxation term, which is not usually evaluated during training. Thus, various types of transformations can be arbitrarily combined while efficiently calculating the log-likelihood of a training image.

[0022] Apart from being able to handle dimensionality-reducing non-injective transformations, various non-bijective transformations can also be used to improve the modeling of distributions having discrete data and discrete structures or truncated components (e.g., structures in an image that contain truncated components). Examples are provided herein.

[0023] As explained, an image classifier typically includes at least one surjective inference. For example, at least one surjective inference may be in the inner layers of the image classifier, and may be preceded or followed by one or more other transformations. Surjective inferences can occur in multiple layers of the image classifier. For example, the input to a surjective inference may be obtained by successively applying several other surjective inferences and / or other transformations. For example, the output of a surjective inference may be used to successively apply several other surjective inferences and / or other transformations. At least one bijective transformation may precede and / or follow a surjective inference.

[0024] In one embodiment, the image classifier may consist only of bijective and inferential surjective transforms, with the optional use of a probabilistic output layer. This can produce a forward deterministic image classifier (excluding the output layer). This has the advantage of enabling efficient classification. Many conventional image classification models are of this type.

[0025] However, it is also possible to use generative surjective and / or stochastic transformations. Examples are provided herein. For example, various transformations particularly suitable for modeling the symmetry and segmented components of images are described herein.

[0026] Generally, transforms and / or their inverse transforms can be parameterized by parameters learned when training an image classifier. This applies not only to inferential surjective transforms but also to generative surjective, bijective, and stochastic transforms (although transforms without parameters are also possible). In many cases, the parameters of a transform and its inverse transform partially overlap or coincide. It is also possible that the set of parameters for a transform is a subset of the set of parameters for its inverse transform, or that the inverse transform has parameters but the transform itself does not. The latter case is particularly applicable to inferential surjective transforms, in which case the transform itself computes a deterministic function, and the inverse transform uses parameters to effectively infer the transform input given the transform output. Note that even if the parameters of the inverse transform are not used when applying the classifier, the parameters of the inverse transform are usually used (and therefore these parameters are trained) in the likelihood contribution of the inferential surjective transform.

[0027] The various building blocks known from image classification itself can be implemented using deterministic non-injective transformations. Several examples are shown below, which can be arbitrarily combined as needed for a particular application. As is common in image classification, for various transformations, the input and / or output can be represented as a three-dimensional volume containing one or more channels. Each channel can represent the input image two-dimensionally, and usually some spatial correspondence is maintained between the representation and the input image.

[0028] As an optional means, the image classifier may include a dimensionality-reducing, tightly coupled component implemented using a linear bijective transform and a slice transform. The linear bijective transform can be modified by applying a dimensionality-preserving linear transform. The slice transform can then select a subset of the output of the linear bijective transform. Interestingly, a dimensionality-reducing linear transform can be represented in this way. After (or before) applying the slice, an activation function, such as a nonlinear bijective function or an inferential surjective function like ReLU described herein, can be applied.

[0029] Since slice transformations perform dimensionality reduction, they are deterministic non-injective transformations. Their inverse transformation can be approximated by an inverse model using a stochastic inverse transformation. The inverse transformation can, for example, sample the non-selected output of a linear bijective transformation while preserving the non-selected output, given the selected output of a linear bijective transformation. As an arbitrary selection method, the inverse transformation can be parameterized, allowing for effective learning of substituting the non-selected output when a selected output is given.

[0030] As an optional means, the image classifier may include a convolutional combined transformation. As is known in itself, this is a combined layer, and the function applied is a convolution. In the combined layer, given first and second transformation outputs y1, y2 are obtained by combining the first transformation input with a first function of the second transformation input to obtain the first transformation output, e.g., y1 = x1 + F(x2), and by combining the second transformation input with a second function of the first transformation output, e.g., x2 + G(y1), to obtain the second transformation output.

[0031] Typically, the first and second transformation inputs are subsets of channels in an input activation volume, and similarly, the first and second transformation outputs each provide one or more channels in an output activation volume. The transformation is applied convolutionally to the input, but typically with a stride of 1 to provide reversibility.

[0032] Convolution is well known to be useful for image classification. The use of a combined layer is particularly beneficial in the current setup because it allows for efficient inversion without the need for a stochastic inverse transformation. After the convolutional combined transformation, a slice transformation can be optionally performed to select a subset of the output channels.

[0033] As an optional means, an image classifier may include a maximum transformation. A maximum transformation can compute the transformed output as the maximum value of multiple transformed inputs. This is another type of transformation commonly used in image classification. For example, it may be implemented by a maximum pooling layer convolutionally applying the maximum transformation across the entire input volume. The maximum transformation is a deterministic non-injective transformation. In its inverse model, its inverse transformation can be approximated by sampling the index of the maximum transformed input given the transformed output, and sampling the values ​​of the non-maximized transformed input given the transformed output and the index of the maximum transformed input. Again, the inverse model can be parameterized to learn to make optimal predictions of the transformed inputs, but this is not mandatory.

[0034] As an optional means, the image classifier may include, for example, a ReLU transform, in which the transformed output is calculated by mapping the transformed input from a predetermined interval to a predetermined constant, such that z=max(x,0) maps the input in the interval [-∞,0] to zero. Furthermore, this transform is deterministic and non-injective. In the inverse model, its inverse transform can be approximated by an inverse transform that samples the transformed input from a predetermined interval, given that the transformed output is equal to a predetermined constant.

[0035] As an optional method, an image classifier can be configured to classify an input image into a class by obtaining a vector of class probabilities for each class, and then determine the class from this vector in the output layer. This decision can be deterministic, such as selecting the most likely class, or probabilistic, such as sampling based on probability. In the inverse model, the inverse transform of the output layer can generally be approximated using trainable parameters based on a conditional probability distribution of the vector for class probabilities given the determined class. Furthermore, this conditional probability distribution can be used to determine a class based on class probabilities, for example, according to Bayes' theorem. This provides, in principle, a way to define the output and its inverse transform using a relatively small number of parameters.

[0036] As an optional means, the image classifier may include a probabilistic transformation with a deterministic inverse transformation, in other words, a generative surjective transformation. In this case, the likelihood contribution can also be calculated, but conversely, it can be calculated based on the probability that the transformation output is produced given a transformation input. Various types of image data can be modeled more accurately using generative surjective transformations. For example, a generative rounding surjective, such as the one discussed herein, may be used as an initial layer to effectively dequantize discrete image data into continuous values.

[0037] Interestingly, an image classifier trained according to the techniques presented herein can be used not only to classify images but also to generate additional images by using the inverse model. As mentioned above, the parameters of the transformations that constitute an image classifier and the parameters of their inverse transformations are not necessarily identical. Therefore, when using a classifier solely for image classification or solely for image generation, it may be necessary to access only the respective subsets of the trained classifier's parameters.

[0038] Specifically, the inverse model can be used as a class-conditional generative model by obtaining a target class and applying the inverse model to generate images representing the target class. For example, as described elsewhere, the class probability vector may be obtained based on the target class, following the inverse of the output layer. The class probability vector can also be arbitrarily set to generate images with specified correspondences to multiple classes. Alternatively, the target class can be sampled first, and then images from that class can be sampled to obtain images representing the entire training dataset. Images generated by applying the inverse model can be used, for example, as training and / or test data for training further machine learning models.

[0039] As an optional mechanism, the image classifier can be configured to calculate a confidence score for the requested classification. Because the training is based on log-likelihood, the confidence score can accurately represent the probability that the input image actually belongs to the requested class.

[0040] It will be understood by those skilled in the art that two or more of the above embodiments, implementations, and / or any aspects of the present invention can be combined in any way that is deemed useful.

[0041] Modifications and variations for any system and / or any computer-readable medium corresponding to the described modifications and variations for the methods implemented in the corresponding computer can be performed by those skilled in the art based on this description.

[0042] These and other aspects of the present invention will become apparent and further clarified by reference to embodiments described as examples in the following description and by reference to the accompanying drawings. [Brief explanation of the drawing]

[0043] [Figure 1] This is a diagram showing a system for training a model. [Figure 2] This figure shows a system for using a pre-trained model. [Figure 3] This figure shows a (semi)autonomous vehicle using an image classifier. [Figure 4a] This figure shows a detailed example of the conversion. [Figure 4b] This figure shows a detailed example of the conversion. [Figure 4c] This figure shows a detailed example of the conversion. [Figure 4d] This figure shows a detailed example of the conversion. [Figure 4e] This figure shows a detailed example of a trained model. [Figure 5a] This figure shows a detailed example of the output layer. [Figure 5b] This figure shows a detailed example of slice transformation. [Figure 5c] This figure shows a detailed example of maximum value transformation. [Figure 5d] This figure shows a detailed example of an image classifier. [Figure 6] This figure shows the method implemented in the computer used to train the model. [Figure 7] This diagram shows how a computer implements the use of a trained model. [Figure 8] This is a diagram showing a computer-readable medium containing data.

[0044] Please note that these diagrams are purely illustrative and not drawn to scale. In the diagrams, elements corresponding to elements already described may share the same reference number.

[0045] Detailed description of the embodiment Figure 1 shows a system 100 for training a model. The model can be configured to produce a model output given an input instance. In one embodiment, the input instance may be an image. In one embodiment, the model may be a classification model configured to classify the input instance into a class from a set of classes.

[0046] System 100 may include a data interface 120 for accessing the training dataset 030. The training dataset 030 may include at least one labeled training instance (e.g., an image) labeled with the output of the training model (e.g., a class from a set of classes). Alternatively or in addition to this, the training dataset 030 may include at least one unlabeled training instance (e.g., an image). For example, the training dataset may contain at least 1000, at least 100000, or at least 10000000 training instances. For example, up to or at least 1%, up to or at least 5%, or up to or at least 10% of the training instances may be labeled.

[0047] As shown in the figure, the data interface 120 may also be for accessing model data 040 representing the model being trained. In particular, the model data may include a set of parameters for the model being trained. For example, the model may include at least 1000, at least 10000, or at least 100000 trainable parameters. The model data may define a forward model for obtaining the model output given an input instance, and an inverse model for obtaining the input instance from the model output. Typically, the parameter sets for the forward and inverse models overlap. They may coincide, but not necessarily, and for example, the inverse model may include additional parameters. Using the trained model, the model can be applied and / or input instances can be generated by, for example, system 200 in Figure 2, according to the method described herein. Systems 100 and 200 can also be combined into a single system.

[0048] For example, as shown in Figure 1, the data interface 120 can be configured by a data storage interface 120, which can access data 030,040 from data storage 021. For example, the data storage interface 120 may be a memory interface or persistent storage interface, such as a hard disk or SSD interface, or a personal, local, or wide area network interface such as Bluetooth, Zigbee, Wi-Fi interface, Ethernet, or fiber optic interface. The data storage 021 may be internal data storage of the system 100, such as a hard drive or SSD, or it may be external data storage, such as network-accessible data storage. In some embodiments, data 030,040 may each be accessed from different data storage, for example, via different subsystems of the data storage interface 120. Each subsystem may be of the type described above for the data storage interface 120.

[0049] System 100 may further include a processor subsystem 140 which can be configured to define an inverse model of a model during the operation of System 100, where the model includes a set of transformations. The set of transformations may include at least one deterministic non-injective transformation whose inverse transformation can be approximated by a stochastic inverse transformation in the inverse model. Alternatively, or in addition to the above, the set of transformations may include at least one stochastic transformation with a deterministic inverse transformation.

[0050] The processor subsystem 140 can be further configured to train a model using log-likelihood optimization while the system 100 is running. During optimization, training instances can be selected from the training dataset 030. Model 040 can be applied to the training instances. This may include determining the transformation inputs to each transformation based on the training images, and applying the transformations to obtain the respective transformation outputs.

[0051] In the case of a deterministic non-injective transformation, the processor subsystem 140 can be configured to calculate the likelihood contribution based on the probability of the transformation input of the transformation generated by the inverse stochastic transformation, given the transformation output of the transformation. In the case of a stochastic transformation accompanied by a deterministic inverse transformation, the processor subsystem 140 can be configured to calculate the likelihood contribution based on the probability of the transformation output of the stochastic transformation generated by the transformation, given the transformation input of the stochastic transformation.

[0052] System 100 can be configured to use labeled training instances. In this case, if the selected training instances are labeled, the processor subsystem 140 can be configured to use the calculated likelihood contributions to calculate the log-likelihood of the labeled training instances and their labels according to the joint probability distribution of the input instances and outputs determined by the model.

[0053] System 100 may be configured to use unlabeled training instances instead of, or in addition to, this. In this case, if the selected training instances are unlabeled, the processor subsystem 140 may be configured to use the calculated likelihood contribution to calculate the log-likelihood of the unlabeled training instances according to the probability distribution of the input instances generated by the inverse model.

[0054] The system 100 may further include an output interface for outputting model data 040 representing a learned (or "trained") model. For example, as shown in Figure 1, the output interface may be configured by a data interface 120, which in these embodiments is an input / output ("IO") interface through which the trained model data 040 can be stored in data storage 021. For example, the model data defining an "untrained" model can be replaced, at least partially, by the model data 040 of the trained model during or after training, and the model parameters, such as neural network weights and other types of parameters, may be adapted so as to reflect training on the training data 030. This is also shown in the model data 040 in Figure 1. In other embodiments, the trained model data 040 may be stored separately from the model data defining an "untrained" model. In some embodiments, the output interface may be separated from the data storage interface 120, but the output interface may generally be the type of interface described above with respect to the data storage interface 120.

[0055] Figure 2 shows a system 200 for using a trained model. The model can be configured to produce a model output given an input instance. In one embodiment, the input instance may be an image. In one embodiment, the model may be a classification model configured to classify the input instance into a class from a set of classes.

[0056] System 200 may include a data interface 220 for accessing model data 040 representing a trained model, as defined by System 100 in Figure 1 or as described elsewhere. As discussed with respect to Figure 1, the model data may include parameters for the forward transformation to obtain the model output from the input instances and / or parameters for the inverse transformation to obtain the input instances from the model output. These parameters may partially overlap. System 200 can be configured to obtain the model output using the trained model, in which case the model data 040 may include at least the forward transformation parameters, but not necessarily the inverse transformation parameters. Alternatively or in addition to this, System 200 can be configured to obtain the input instances using the trained model, in which case the model data 040 may include at least the inverse transformation parameters.

[0057] For example, as shown in Figure 2, the data interface can be configured by a data storage interface 220, which can access data 040 from data storage 022. Generally, the data interface 220 and data storage 022 may be of the same type as those described with reference to Figure 1 for data interface 120 and data storage 021. The data storage may optionally include, for example, input instances to which the model is applied, which include sensor data. The input instances may also be received directly from sensor 072 via another type of interface or via sensor interface 260, instead of being accessed from data storage 022 via data storage interface 220.

[0058] System 200 may further include a processor subsystem 240, which can be configured to use a trained model during the operation of System 200. Using a trained model may include taking input instances and applying the model to obtain model inputs, for example, classifying input images into classes from a set of classes. Alternatively or in addition to this, using a model may include applying the inverse model to generate synthetic instances. This may include, for example, sampling the transform input of a deterministic non-injective transform based on the transform output of the said transform by a probabilistic inverse transform. The obtained model outputs or instances may be output using a data / output interface as described elsewhere.

[0059] It will be understood that the same considerations and implementation options as for processor subsystem 140 in Figure 1 apply to processor subsystem 240. Unless otherwise specified, it will be understood that the same considerations and implementation options as for system 100 in Figure 1 generally apply to system 200.

[0060] Figure 2 further illustrates various components of the system 200 as desired. For example, in some embodiments, the system 200 may include a sensor interface 260 for direct access to sensor data 224 acquired by a sensor 072 in the environment 082. Input instances to which the model applies may be based on or include the sensor data 224. The sensor may be located in the environment 082, but may also be located remotely from the environment 082, for example, if the quantity may be measured remotely. Sensor 072 may or may not be part of the system 200.

[0061] Sensor 072 may have any suitable form, such as an image sensor, LiDAR sensor, radar sensor, pressure sensor, or storage temperature sensor. This figure shows sensors for providing image data, such as a video sensor, radar sensor, LiDAR sensor, ultrasonic sensor, motion sensor, or thermal imaging sensor.

[0062] In some embodiments, the sensor data 072 can sense measurements of different physical quantities, in that it may be obtained from two or more different sensors sensing different physical quantities. The sensor data interface 260 may have any suitable form whose type corresponds to the type of sensor. This includes, but is not limited to, a low-level communication interface based on I2C or SPI data communication, or a data storage interface of the type described above for the data interface 220.

[0063] In some embodiments, the system 200 may include an actuator interface 280 for providing control data 226 to an actuator (not shown) in the environment 082. Such control data 226 is generated by a processor subsystem 240 and can control the actuator based on the output of a model applied and / or based on the generated input instance. The actuator may be part of the system 200. For example, the actuator may be an electric, hydraulic, pneumatic, thermal, magnetic, and / or mechanical actuator. Non-limiting examples include electric motors, electroactive polymers, hydraulic cylinders, piezoelectric actuators, pneumatic actuators, servo mechanisms, solenoids, and stepping motors. Such a type of control is illustrated with reference to Figure 3 for a (semi)autonomous vehicle.

[0064] In other embodiments (not shown in Figure 2), the system 200 may include an output interface to rendering devices such as a display, light source, speaker, or vibration motor. This can be used to generate sensory-perceptible output signals that may be generated based on a requested model output or a generated input instance. For example, the signals may be for use in guidance, navigation, or other types of control in a computer-controlled system.

[0065] In further embodiments (not shown in Figure 2), the system 200 may include an output interface that outputs multiple input instances generated by applying the inverse model, as described with respect to Figure 1, for example, and is an interface for training further machine learning models using them as training and / or test data. For example, the input instances may be used as labeled data, or they may be labeled (e.g., manually) as an optional means and used as labeled data. The input instances can also be used to refine model 040. For example, the system 200 may provide the generated instances (and optionally their labels) to the system 100 to refine model 040.

[0066] In general, each system described herein includes, but is not limited to, System 100 in Figure 1 and System 200 in Figure 2. Such systems can be embodied as a single device or apparatus, such as a workstation or server, or within such a single device or apparatus. The device may be an embedded device. The device or apparatus may include one or more microprocessors running appropriate software. For example, the processor subsystem of each system may be embodied by a single central processing unit (CPU), or by a combination or system of such CPU and / or other types of processing units. The software may be downloaded and / or stored in corresponding memory, such as volatile memory like RAM or non-volatile memory like flash. Alternatively, the processor subsystem of each system may be implemented in the form of programmable logic in a device or apparatus, for example, as a field-programmable gate array (FPGA). In general, each functional unit of each system may be implemented in the form of circuitry. Each system can also be implemented in a distributed manner, including different devices or apparatus, such as distributed local or cloud-based servers. In some embodiments, system 200 may be part of a vehicle, robot, or similar physical entity, and / or may represent a control system configured to control a physical entity.

[0067] Various specific applications of System 200 are envisioned. In one embodiment, System 200 can be used to detect glaucoma in images of (e.g., human) eyes. In one embodiment, System 200 can be used for fault detection in the manufacturing process based on images of manufactured products. In one embodiment, System 200 can be used to classify plants or weeds to determine the need for fertilizers and pesticides.

[0068] Figure 3 shows, as an example, that system 200 is a control system for a (semi)autonomous vehicle 62 operating in environment 50. The autonomous vehicle 62 may be autonomous in that it may include an automatic driving system or a driver assistance system, the latter also referred to as a semi-autonomous system. In this case, the trained model may be an image classifier applied to image data acquired from a video camera 22 integrated into the vehicle 62.

[0069] The autonomous vehicle 62 may incorporate, for example, a system 200 for controlling the steering and / or brakes of the autonomous vehicle based on image data. For example, the system 200 may control the electric motor 42 to perform (regenerative) braking when the autonomous vehicle 62 is in a dangerous traffic situation, for example, when a collision with a traffic participant is expected. The system 200 may control the steering and / or brakes to avoid a collision with a traffic participant, for example. For this purpose, the system 200 may classify images representing the environment of the vehicle 62 as dangerous or non-dangerous based on image data acquired from a video camera.

[0070] As another example, system 200 may classify input images obtained from, for example, the ultrasonic sensors of vehicle 62 to perform near-range object height classification. System 200 can also be used to perform free-space detection in video data from camera 22, for example, a trained model which in this case may be a semantic segmentation model. More generally, the detection of various types of objects in the vehicle's environment, such as traffic signs, road surfaces, pedestrians, and / or other vehicles, by an image classifier can be used in various upstream tasks when controlling and / or monitoring (semi-)autonomous vehicles, such as in driver assistance systems.

[0071] In (semi-)autonomous driving, collecting and labeling training data can be costly and sometimes dangerous. Therefore, the ability to train image classifiers using unlabeled training data and / or generate additional synthetic training data is particularly advantageous.

[0072] The following describes various methods that use image classification as their primary application. However, as those skilled in the art will understand, the methods provided are generalizable to other types of data, in which case the benefits provided (including improvements in semi-supervised learning, improvements in training generative models, and improvements in the ability to model specific types of datasets, such as discrete or symmetric data) will also apply.

[0073] Various embodiments relate to trainable models (e.g., image classifiers) and their inverse models. Such trainable models can be constructed by configuring one or more transformations between a transformation input and a transformation output.

[0074] Mathematically,

number

[0075] In the context of generative modeling, the paper “Variational Inference with Normalizing Flows” by D. Rezende et al. (incorporated herein by reference and available at https: / / arxiv.org / abs / 1505.05770) describes normalizing flows. These utilize a bijective transformation f to transform a simple fundamental density p(z) into a more expressive density p(x), with the variable transformation formula p(x)=p(z)|det∇ x f -1 (x)| is being used.

[0076] Furthermore, in the context of generative modeling, the paper “Auto-Encoding Variational Bayes” by D. Kingma et al. (incorporated herein by reference and available at https: / / arxiv.org / abs / 1312.6114) discusses variational autoencoders (VAEs). A VAE defines a probabilistic graphical model in which each observed variable x is associated with a latent variable z, and the generative process is z~p(z) and x~p(x|z), where p(x|z) can be considered a probabilistic transformation. VAEs use variational inference with an amortized variational distribution q(z|x) to approximate the usually cumbersome posterior p(z|x). This facilitates the calculation of a lower bound on p(x), known as the evidence lower bound (ELBO), for example,

number

[0077] The inventors have found it desirable to train and use models that include both bijective and stochastic transformations. Furthermore, they have found that it is often desirable to have additional types of transformations, particularly in image classification. Indeed, bijective transformations are deterministic and allow for accurate likelihood calculations, but require the preservation of dimensionality. Stochastic transformations, on the other hand, can change the dimensionality of the random variable, but only provide a stochastic lower bound estimate of the likelihood. Interestingly, the inventors have devised a technique for training and using models that include transformations that can change dimensionality while still allowing for accurate likelihood estimation.

[0078] To facilitate any combination of different types of transformations, the inventors have provided a forward transformation f: Z → X with a related conditional probability p(x|z), and an inverse transformation f with a related distribution q(z|x). -1 :X→Z, and (iii) the likelihood contribution term used to approximate the difference between the marginal probability distributions of the transformed input and output, which is used in the calculation of the log-likelihood. Specifically, the inventors have found that under any transformation, the density p(x) is

number

[0079] Figures 4a and 4b show detailed and non-limiting examples of transformations that may be used in the models described herein.

[0080] Figure 4a shows the bijective transformation BT, 441. From the figure, it can be seen that each of the transformation inputs 431-433 is mapped to each of the transformation outputs 451-453 by the bijective transformation. A unique inverse transformation is defined that maps each of the transformation outputs 451-453 to each of the transformation inputs 431-433. Mapping the inputs to the outputs and then mapping the outputs back to the inputs returns the original inputs. Similarly, mapping the outputs back to the inputs returns the original outputs.

[0081] Figure 4b shows a stochastic transformation ST, 442, e.g., p(x|z). At least some of the transformation inputs 431-433 do not deterministically map to unique transformation outputs 451-453 (for example, transformation input 431 shown in the figure maps to transformation output 451 with a non-zero probability and to transformation output 452 with a non-zero probability). The inverse transformation is usually shown as a variational distribution q(z|x) that approximates the posterior p(z|x). Also, in the case of the inverse transformation, at least some of the transformation outputs 451-453 do not deterministically map to unique transformation inputs 431-433 (for example, transformation output 451 shown in the figure maps to transformation input 431 with a non-zero probability and to transformation input 432 with a non-zero probability). Therefore, applying the inverse transformation after applying the transformation to the transformation input does not necessarily yield the original input. Similarly, applying the transformation after applying the inverse transformation to the transformation output does not necessarily yield the original output.

[0082] The bijective transformation BT and the stochastic transformation ST can be described in terms of forward transformation, inverse transformation, and likelihood contribution as follows:

[0083] Forward transformations: In the case of a stochastic transformation ST, the forward transformation can be defined by a conditional distribution p(x|z). In the case of a bijective transformation BT, the forward transformation can be a deterministic function, such as p(x|z)=δ(xf(z)) or x=f(z).

[0084] Inverse transform: In the case of a bijective transform BT, the inverse transform is also a deterministic function, for example, z=f -1(x). In the case of stochastic transformation ST, the inverse transformation is also stochastic. According to Bayes' theorem, the inverse transformation can be defined, for example, as p(z|x)=p(x|z)p(z) / p(x). In many cases, p(z|x) cannot be calculated or is too computationally expensive, so variational approximation q(z|x) can be used.

[0085] Likelihood contribution: In the case of bijective transformation BT, the density p(x) is obtained from p(z) and the mapping f using the variable transformation formula, logp(x)=logp(z)+log|det∇ x f -1 (x)|, z=f -1 (x) can be calculated as shown above. In the formula, |det∇ x f -1 (x)| is the Jacobian matrix of f

Formula

[0086] In the case of stochastic transformation ST, the marginal density p(x) is

Formula

Formula

[0087] For example, each bijective transformation and stochastic transformation

number

number

[0088] As those skilled in the art will understand, this code can be adapted to also cover generative surjective and inferential surjective methods, as discussed herein.

[0089] Figure 4c shows a detailed and non-restrictive example of a transformation that may be used in the model described herein. The transformation shown in this figure is the generative surjective transformation GST, 443. Here,

number

[0090] Interestingly, the inventors have discovered that generative surjective GSTs can also be expressed as forward transforms, inverse transforms, and likelihood contributions.

[0091] Forward transformations: Similar to the bijection in Figure 4a, a surjective GST can be expressed as a deterministic forward transformation p(x|z)=δ(xf(z)) or x=f(z).

[0092] Inverse transform: However, unlike the bijection in Figure 4a, the surjective f:Z→X is not invertible because multiple inputs can be mapped to the same output. However, a right inverse can also be defined for surjectives in the inference direction (e.g., moving from image to class), for example,

number

number

[0093] Likelihood contribution: The likelihood contribution is,

number

[0094] Figure 4d shows a detailed and non-restrictive example of a transformation that may be used in the models described herein. This figure shows the inference surjective transformation IST, 444.

[0095] Forward Transform: In contrast to the transformation in Figure 4c, which is surjective in the generation direction Z→X (e.g., transformation from class to image), this transformation is surjective in the inference direction X→Z (e.g., transformation from image to class). Therefore, in the generation direction, the forward transform p(x│z) can be stochastic, for example, at least one transformation input 431~432 is not uniquely mapped to transformation outputs 451~453. For example, as shown in the figure, transformation input 432 may map to transformation output 452 with a non-zero probability and to transformation output 453 with a non-zero probability. According to Bayes' theorem, the forward transform can also be derived from the inverse transform.

[0096] Inverse Transform: Inverse transforms can be deterministic, for example, each transform output maps to a unique transform input. However, they are not injective; for example, there may be two transform outputs 452 and 453 that map to the same transform input 432. Preferably, a probabilistic forward transform for a given transform input has support only for the set of transform outputs that map to that transform input.

[0097] Likelihood contribution: The likelihood contribution is,

number

[0098] The following table summarizes the transformations, inverse transformations, likelihood contributions, and marginal relaxation terms explained in relation to Figures 4a to 4d, i.e. [Table 1] That is the case.

[0099] Figure 4e shows a detailed and non-restrictive example of a trained model, specifically an image classifier. This figure shows the image classifier IC420 applied to classify an input image II, 410 into class CL, 460 from a set of classes. The input image II can be represented, for example, as a 2D volume or as a 3D volume containing one or more channels (e.g., 1 for grayscale, 3 for color). The set of classes can be, for example, a finite predefined set of two or more classes, a maximum or minimum of five classes, or a maximum or minimum of ten classes. The classes do not need to be mutually exclusive, and for example, classes may be considered as attributes (e.g., in traffic situations, "motorcycle rider", "go straight", "overtake"). The image classifier is configured to classify classes as having a set of one or more attributes.

[0100] Interestingly, the inventors have discovered that an image classifier IC can be implemented using a combination of one or more transformations, as described with respect to Figures 4a to 4d. Specifically, if we consider the transformations in Figures 4a to 4d as a transformation in the generation direction Z→X, then the image classifier IC can be considered as a transformation p(y│x) composed of inverse transformations in the inference direction X→Z, as described with respect to Figures 4a to 4d. Next, the inverse model of the image classifier can be defined by inverting each of the transformations that constitute the image classifier IC, for example, by using the forward transformations described with respect to Figures 4a to 4d. Thus, the inverse model of the image classifier IC can define a class-conditional generative model p(x│y) from class CL to input image II.

[0101] The provided technology allows, advantageously, the image classifier IC to be a deterministic function y=f(x) with a probabilistic inverse model p(x│y). This is not possible when using only bijective transformations, for example, because the function f representing the image classifier is not bijective (multiple input images II can be mapped to the same class). It is also not possible when using only probabilistic transformations, because a probabilistic mapping from image to class is created.

[0102] Alternatively, this can be achieved by including at least one inference probabilistic transformation IST, 444 in the image classifier IC, as explained with respect to Figure 4d, for example. This transformation is deterministic and non-injective in the inference direction.

[0103] By applying the image classifier IC to the input image II and applying each transformation of the image classifier in the inference direction, the input image can be classified into class CL. Thus, the transformation input TI, 430 of transformation IST can be obtained from the input image II, and then this transformation can be applied to the transformation input TI as a deterministic function z=g(x) to obtain the transformation output TO, 450. Next, the output classification CL can be determined based on the transformation output TO. The function g is parameterizable by a set of parameters for the image classifier IC.

[0104] The image classifier IC can also be used to generate a composite input image II by applying the inverse transforms of each of the image classifier's transformations in the generation direction. For example, the image classifier can generate an input image given a class CL (or multiple non-mutually exclusive attributes), or more generally, given values ​​in the inner or output layers of the model. As also explained in Figure 4d, the inverse transform of the inferential surjective IST can be approximated by the inverse model using a stochastic inverse transform (referred to as the forward transform). The stochastic inverse transform can define a probability distribution p(x|z) of the transformed input TI given a transformed output TO. In other words, the stochastic inverse transform can be configured to probabilistically generate the transformed input TI given a transformed output TO. The definition of the probability distribution p(x|z) can be parameterized by a set of parameters of the image classifier CL and may overlap with the parameters of the function g. Thus, when generating image II, the transformed input TI can be sampled according to a probability distribution given a transformed output TO.

[0105] (Note that Figure 4d illustrates the transformation of the generation direction. Therefore, the transformation input TI in Figure 4e corresponds to the transformation outputs 451-453 in Figure 4d, and the transformation output TO in Figure 4e corresponds to the transformation inputs 431-432 in Figure 4d.)

[0106] The image classifier IC can be trained using maximum likelihood estimation; that is, optimization is feasible, and the parameters of the image classifier IC, including the parameters of the inferential surjective IST (and its forward and / or inferential transform), can be optimized with respect to the objective function. The objective function can include the log-likelihood of images from the training dataset and is maximizable. The objective function may include additional terms such as normalizers.

[0107] Interestingly, the image classifier IC can be trained on a training dataset containing any combination of labeled training images for which the relevant training classes are available and unlabeled training images for which the relevant training classes may not be available. For labeled training images, the maximized log-likelihood may be the log-likelihood of the labeled training images and their labels logp(x,y), according to the joint probability distribution of the input images and classes determined by the image classifier IC. For unlabeled training images, the maximized log-likelihood may be the log-likelihood of the unlabeled training images logp(x), according to the probability distribution of the input images generated by the inverse model.

[0108] Interestingly, in each case, the log-likelihood can be efficiently calculated based on the likelihood contributions of the various transformations, as explained with respect to Figures 4a to 4d. Since the likelihood contributions allow us to estimate the difference in log-likelihoods between transformed outputs, adding the likelihood contributions of each transformation allows us to estimate the difference in log-likelihoods between the input and output of the combined transformation. This difference can then be used to calculate the log-likelihood of the training images.

[0109] Typically, training is performed using stochastic optimization, such as stochastic gradient descent. For example, the Adam optimizer can be used, as disclosed in Kingma and Ba's "Adam: A Method for Stochastic Optimization" (incorporated herein by reference and available at https: / / arxiv.org / abs / 1412.6980). As is known, such optimization methods are heuristic and / or capable of reaching local optima. Training can be performed per instance or in batches, for example, up to or at least 64 instances, or up to or at least 256 instances.

[0110] It should be noted that an image classifier does not have to consist solely of transformations, as explained with respect to Figures 4a to 4d. Specifically, the output of the classifier can be obtained using an output layer, for example, as explained with respect to Figure 5a. The output layer is not considered to belong to the set of transformations. Apart from the use of an output layer as an optional means, in some embodiments, the image classifier IC can be composed entirely of the four types of transformations described. Apart from the use of an output layer as an optional means, in some embodiments, the image classifier IC can be composed entirely of bijective and inferential surjective transformations. However, it is also possible to use generative surjective and / or stochastic transformations in the image classifier.

[0111] Furthermore, as those skilled in the art will understand, by appropriately adapting the image classifier IC, it can also be used for non-image input instances (e.g., other types of sensor data) and / or non-classification tasks (e.g., regression, data generation, etc.).

[0112] The following describes various advantageous model components that can be used in image classifier ICs or other pre-trained models. Examples of various model architectures for image classifier ICs based on these components are also described below.

[0113] Figure 5a shows a detailed and non-restrictive example of an output layer for use in, for example, an image classifier IC or another classifier. The image classifier can be configured to obtain a vector of class probabilities π1,...,πk, 531 for each class of input images. This is usually done using a set of transformations of the type shown in Figures 4a-4d, as also explained with respect to Figure 4e. The output layer of the image classifier can then obtain the class y, 551 for the input images from the class probabilities πi.

[0114] The output layer described in this diagram is not considered part of the set of transformations. This is similar to a stochastic transformation in that p(y│π) and p(π|y) are defined stochastically, but interestingly, it can be evaluated analytically, thus avoiding variational inference.

[0115] As shown in the figure, the output layer can be defined by a conditional probability distribution p(π|y) of the vector of class probabilities given the required class. This conditional probability distribution is, for example, a normal distribution p(π│y)=N(π|μ) as shown in the figure. y ,σ y ) can be defined by the respective probability distributions of each class.

[0116] During training, the log-likelihood of the training images can be calculated based on this conditional probability distribution. If a predetermined prior distribution on the class labels, for example p(y)=1 / K, is used, then K is the number of classes, for example,

number

[0117] For unlabeled training images, the log-likelihood of the training images can be calculated by combining the marginal probabilities of the class vector with the likelihood contribution of the transformation set, for example.

number

number

[0118] Specifically, during training, images can be selected, an image classifier can be applied, and the class probability π can be obtained. For unlabeled training images, logp(π) can be calculated based on the class probability. The likelihood contribution of each transformation can be calculated based on the input and output of each transformation, and logp(x) can be obtained by adding the likelihood contributions as described above. For labeled training images, similarly, logp(y) and logp(π│y) can be combined with the likelihood contributions to obtain logp(x,y). The log-likelihood of labeled and / or unlabeled training examples can be maximized, for example, by evaluating the gradient of the log-likelihood with respect to the parameters of the image classifier and using gradient descent.

[0119] Using a trained image classifier, for example, we can classify an input image into a class by calculating the class probability π of the input image and then using Bayes' theorem to calculate the probability that the image belongs to a class. logp(y|x)=logp(y,x)-logp(x)=logp(y)+logp(π│y)-p(π)=logp(y│π) The final output label y can be selected deterministically, for example, as the most probable class, or probabilistically, for example, by sampling y according to the class probabilities. For example, the calculated probability p(y|x) can also be output for some or all classes. In particular, the most probable class can be returned with a given confidence score p(y|x).

[0120] A trained image classifier can also be used as a class-conditional generative model p(x|y) by sampling images according to a probability distribution of a given class vector and a given inverse transform, for example, logp(x|y)=logp(x|π)-logp(π|y)-logp(π) This is the result. In the formula, p(x|π) is the inverse probabilistic transformation of the combination of the set of transformations. A trained image classifier can also be used as a generative model based on logp(x|π), for example, by starting with a set of class probabilities and generating input images from them. In this way, a given combination of output classes can be realized.

[0121] Figure 5b shows a detailed and non-restrictive example of an inference surjective slice transform for use in an image classifier or other types of pre-trained models.

[0122] A slice transformation allows us to obtain the transformed output z, 552 from the transformed inputs x1, x2, 532 by obtaining a (strict, non-empty) subset of the elements of the transformed input. For example, the transformed input

number

[0123] As shown in the figure, the inverse stochastic transform of a slice transform can set elements selected from the transform input, similar to the transform output, for example, x1 = z. Unselected elements can be approximated by sampling from the transform output, for example, x2 ~ p(x2|z).

[0124] By satisfying the general formula for the likelihood contribution of a surjective inference transformation, the likelihood contribution of this transformation can be calculated as logp(x2|z), for example, as the entropy of the probability distribution used to infer the sliced ​​element x2.

[0125] As explained elsewhere, generative surjective slice transformations can be defined similarly.

[0126] Figure 5c shows a detailed and non-restrictive example of an inference surjective maximum transformation for use in an image classifier or other types of pre-trained models.

[0127] The maximum value transformation allows us to obtain the transformed output z, 553, as the maximum value of multiple transformed inputs x1, ..., xk, ..., xK, 533.

[0128] As shown in the figure, the inverse stochastic transform is performed by (i) taking the index k of the maximum transform input, for example, x k (ii) Sample such that = z, and set z to x k (iii) Map deterministically to z, and the remaining non-maximal transformed input value x of x -k All of these are x k , x -k ~p(x -k This can be done by sampling in such a way that the result is smaller than |z,k). Here, k is the index of x, K is the number of elements in x, x -k x is x excluding element k.

[0129] The probability distribution p(k|z) can be used to train a classifier, for example, or it can be fixed, for example, p(k|z) = 1 / K. In order for the inverse transformation to be the right inverse of the forward transformation, p(x -k |z,k) is (-∞,z) K-1 It is preferable that it be defined so that it is only supported by x. k This will definitely be the maximum value.

[0130] For example, logp(k│z) can be defined such that the output is likely to be copied equally to one of its inputs. The remaining inputs can be sampled such that the copied value remains at its maximum, and this maximum value can be set to be equal to the noise of a standard semi-normal distribution, such as a Gaussian distribution with only positive values.

[0131] The likelihood contribution of this transformation is given by the general formula V = log p(k│z) + log[(x -k It can be found as shown in the equation (|z,k). In the equation, z=x k = maxx, k = argmaxx.

[0132] For example, an image classifier may include a max pooling layer for use in downsampling, which is implemented as a set of surjective maximum values ​​operating on each subset of the input volume. The maximum value transformation can also be fitted to compute the minimum value.

[0133] As explained elsewhere, a generative surjective maximal transformation can be defined similarly.

[0134] Another advantageous transformation for use in this specification is the rounding surjective, which takes a transformation input and rounds it, for example, by calculating its floor. The forward transformation can be a discrete, deterministic, non-injective function, for example

number

number

number

[0135] Another favorable transformation for use in this specification is the absolute value surjective, which returns the magnitude of its input, z=|x|. As an inference surjective, its forward and inverse transforms are: p(x|z)=Σs∈{-1,1} δ(x-sz)P(s|z),q(z|x)=Σ s∈{-1,1} δ(z-sx)δ s,sign(x) It can be expressed as follows: Here, q(z|x) is deterministic and corresponds to z=|x|. The inverse transform p(x|z) involves the following steps: (i) sample the sign s of the transform input, given the transform output z, and (ii) apply the sign to the transform output z to obtain the transform input x=sz. The absolute value surjective is useful for modeling data with symmetry.

[0136] The probability distribution p(s|z) used to sample the sign can also be trained as a classifier, and can be fixed, for example, to p(s|z) = 1 / 2. Fixing the sign can be particularly useful in ensuring exact symmetry around the origin.

[0137] The likelihood contribution of a surjective inference can be calculated as V = log p (s|z), where z = sx = |x| and s = sign(x).

[0138] As a generator surjective, the forward and inverse transforms are: p(x|z)=Σ s∈{-1,1} p(x|z,s)p(s|z)=Σ s∈{-1,1} δ(x-sz)δ s,sign(z) q(z│x)=Σ s∈{-1,1} q(z│x,s)q(s│x)=Σ s∈{-1,1} δ(z-sx)q(s│x) It can be defined as follows: In the formula, the forward transformation p(x|z) is completely deterministic and corresponds to x=|z|. The inference direction involves two steps: 1) sampling the sign s of the transformed input z given the transformed output x, and 2) deterministically mapping the transformed output x to z=sx. Here, the probability distribution of the sign q(s|x) can also be trained as a classifier, and can be fixed to, for example, q(s|x)=1 / 2. The last choice is particularly useful when p(z) is symmetric.

[0139] In this case, the likelihood contribution is:

number

[0140] Absolute value surjectives can be usefully used, for example, to model antisymmetry using a trainable classifier P(s|z) for learning the expansion.

[0141] Another advantageous transformation for use in this specification is the sort surjective. A sort surjective can be used as a generative surjective x = sortz or an inference surjective z = sortx.

[0142] As a generative surjective, a sort surjective is,

number

[0143] The forward transformation p(x|z) is completely deterministic and corresponds to the input sort x=sortz. The inference directions are 1) sampling of permutation indices I conditional on the sorted transformed output x, and 2) inverse permutation I. -1 The converted output x is deterministically rearranged according to the converted input

number

[0144] The likelihood contribution of this transformation is,

number

[0145] As a surjective inference, a surjective sort is,

number

[0146] The transformation q(z|x) is completely deterministic and corresponds to the input sort z=sortx. The inverse transformation involves 1) sampling of permutation indices I conditioned on the transformed output z, and 2) the inverse permutation I -1 The converted output z is deterministically rearranged according to the converted input

number

[0147] The likelihood contribution of this transformation can be calculated as V = log p (I | z), where z = x I =sortx, I=argsortx.

[0148] Sort surjectives are particularly useful for modeling naturally sorted data, learning order statistics, and training interchangeable models using flows. In particular, interchangeable data can be modeled by constructing any number of transformations along with sort surjectives.

[0149] Another favorable transformation is a probabilistic permutation. This is a probabilistic transformation that randomly rearranges the inputs. An inverse path can be defined that reflects the forward path.

[0150] Forward and reverse conversions are performed as follows:

number

[0151] The transformation is probabilistic and may involve the same steps in both directions: 1) sampling permutation indices I, for example, randomly and uniformly, and 2) deterministically rearranging the inputs according to the sampled indices I.

[0152] When using uniform random sampling, the likelihood contribution can be shown to be zero. Probabilistic permutations are useful for modeling interchangeable data by constructing any number of transformations using a probabilistic permutation layer and applying permutation invariance.

[0153] For example, interchangeable data can be modeled using one or more connected flows parameterized by a Transformer network (available from A. Vaswani et al.'s "Attention Is All You Need," https: / / arxiv.org / abs / 1706.03762, incorporated herein by reference) without using positional encoding. For example, probabilistic permutations may be inserted between each connected layer, or fixed permutations may be used after inducing ordering using an initial sort surjective.

[0154] The following table summarizes several advantageous surjective inference layers described herein, namely, [Table 2] That is the case.

[0155] The following table summarizes several advantageous generative surjective layers described herein, namely, [Table 3] That is the case.

[0156] Various favorable combinations of the above transformations can be defined. The image classifier may be a neural network, and for example, the class probability may be determined for the input image by a function representable by the neural network, at least in the inference direction. This is the case, for example, when the trainable part uses a bijective and inferential surjective function given by the neural network.

[0157] Generally, an image classifier may include one or more convolutional layers in which an input volume (e.g., size m × n × c) is transformed by the layer into an output volume (e.g., size m' × n' × c'), and the spatial correspondence between the input and output volumes is preserved. Such layers may be implemented by one or more transformations as described herein. An image classifier including such layers may be referred to as a convolutional model. For example, an image classifier may be a convolutional neural network. An image classifier may include, for example, up to or at least 5, up to or at least 10, or up to or at least 50 convolutional layers.

[0158] For example, the image classifier may include a convolutional join transform, as described elsewhere. In one embodiment, the image classifier includes a ReLU layer that applies a ReLU transform to each portion of its input vector. In one embodiment, the image classifier includes a max pooling layer that convolutionally applies a max transform to its input volume to reduce the spatial dimension of the input volume. In one embodiment, the image classifier includes a slice transform that selects a subset of channels, thereby reducing the number of channels.

[0159] A convolutional layer may be followed by one or more non-convolutional layers, such as one or more tightly coupled layers. Such tightly coupled layers may be implemented, for example, by combining linear bijective transformations and slice transformations. For example, the number of non-convolutional layers may be one, two, a maximum or at least five, or a maximum or at least ten.

[0160] Figure 5d shows a detailed and non-limiting example of an image classifier based on the image classifier in Figure 4e, for example.

[0161] Specifically, the diagram shows an input image x, 510, which is transformed through a series of transformations into a set of class probabilities π, 560 for each class. For example, as discussed with respect to Figure 5a, the class can be determined based on the set of class probabilities. Since this example image classifier uses only bijective and inferential surjective transformations, π = g(x) is given by a deterministic function.

[0162] Specifically, the example shown is a convolutional combined transform CC, 541 applied to the input image x. This is a bijective transform. As described elsewhere, such a layer can compute first and second transformed outputs based on first and second transformed inputs by applying two transforms, for example, as disclosed in A. Gomez et al.'s "The Reversible Residual Network: Backpropagation Without Storing Activations" (available at https: / / arxiv.org / abs / 1707.04585, incorporated herein by reference). Both transformeds applied are convolutions applied to their respective input volumes.

[0163] After the convolutional join transformation, in this example, a ReLU layer 542 is applied, as described elsewhere. This is an inference surjective layer. Next, a max pooling layer MP, 543, is applied. This layer performs downscaling by convolving the max transformation across the entire input.

[0164] Layers 541-543 are convolutional layers that determine the output volume from each input volume. Layers 541-543 can be repeated multiple times, individually or in combination.

[0165] A flattening layer F, 544 is then applied to convert the output of the last convolutional layer into a one-dimensional feature vector. Subsequently, a tensor slice layer TS, 545 is applied to select a subset of features.

[0166] In this particular example, the training log-likelihood is calculated by summing the likelihood contributions of each transformation, for example. -V1=log|detJ|(in the case of convolutional join transformation CC) -V2=I(z=0)logp(x)(in the case of a ReLU layer) -V3=logp(k|z)+logp(x -k |z,k) (in the case of maximum pooling layer MP) -V4=0 (in the case of flattening layer F) -V5=logp(x2|z)(For a tensor slice layer TS, x=(x1,x2) and z=x1) It can be calculated like this.

[0167] Many modifications will be conceivable to those skilled in the art. In particular, the ReLU layer can be replaced with the "Sneaky ReLU" activation function, as described in M. Finzi et al.'s "Invertible Convolutional Networks," from the proceedings of the first workshop on Invertible Neural Networks and Normalizing Flows at ICML 2019. Interestingly, this activation function is reversible and has closed-form inverse and logarithmic determinants. As an optional measure, the model may include an initial generative rounding surjective to accommodate the discrete nature of the input image data and to avoid causing divergence during training. It has also been found that the performance of the model can be improved to include a generative slice surjective as described herein, and the number of input channels can be increased from, for example, one layer (grayscale) or three layers (color images) to a larger number N, e.g., N≧5. The model may also include a tightly coupled part after flattening, for example, as described herein.

[0168] Figure 6 shows a block diagram of Method 600 implemented on a computer for training a model, such as an image classifier. The model can be configured to derive a model output from an input instance, for example, to classify an input image into a class from a set of classes. Method 600 may correspond to the operation of System 100 in Figure 1. However, this is not a limitation, and Method 600 may be performed using a different system, apparatus, or device.

[0169] Method 600 may include accessing a training dataset (610) in an operation titled “Accessing training data.” The training dataset may include at least one labeled training instance and / or at least one unlabeled training instance labeled with a model output (e.g., a class).

[0170] Method 600 may include defining an inverse model of the model (620) in an operation titled “Define an Inverse Model”. The model may include a set of transformations. The set of transformations includes at least one deterministic non-injective transformation. The inverse transformation of this transformation can be approximated by an inverse model through a stochastic inverse transformation. Alternatively, or in addition to this, the set of transformations may include at least one stochastic transformation with a deterministic inverse transformation.

[0171] Method 600 may include training the model using log-likelihood optimization in an operation titled “Train the model” (630).

[0172] As part of the training operation 630, method 600 may include selecting training instances from a training dataset (632) in an operation titled “Selection of training instances”.

[0173] As part of the training operation 630, method 600 may further include applying the model to a training instance (634) in an operation titled “Apply Model to Instance.” This may include applying the transformation to the transformation input of the transformation to obtain the transformation output of the transformation.

[0174] As part of training operation 630, method 600 may further include calculating the likelihood contribution of a transformation (636) in an operation titled “Calculate Likelihood Contribution.” For a deterministic non-injective transformation, this contribution may be based on the probability that a stochastic inverse transformation produces a transformation input given a transformation output of the transformation. For a stochastic transformation using a deterministic inverse transformation, this may be based on the probability that a stochastic transformation produces a transformation output given a transformation input.

[0175] If the training instances are labeled as part of the training operation 630, method 600 may include, in an operation titled “Calculate Joint Log-Likelihood,” using the obtained likelihood contributions to calculate the log-likelihood of the labeled training instances and their labels according to the joint probability distribution of the input instances and classes obtained by the image classifier (638).

[0176] As part of the training operation 630, if the training instance is unlabeled, method 600 may include, in an operation titled “Calculate Log-Likely,” using the obtained likelihood contribution to calculate the log-likely value of the unlabeled training instance according to the probability distribution of the input instance generated by the inverse model (639).

[0177] The various steps 632-639 of the training operation 630 may be performed once or multiple times, for example, for a certain number of iterations and / or until convergence occurs, in order to train the model.

[0178] Figure 7 shows a block diagram of Method 700 implemented on a computer using a trained model, such as an image classifier. The model can be configured to derive a model output from an input instance, for example, it may be configured to classify an input image into a class from a set of classes. Method 700 may correspond to the operation of System 200 in Figure 2. However, this is not a limitation, and Method 700 may be performed using a different system, apparatus, or device.

[0179] Method 700 may include accessing model data representing a trained model (710) in an operation titled “Accessing Model Data.” The model may be trained according to the methods described herein. When using the model to obtain model outputs, the model data includes at least the parameters of the forward transformation of the trained model, but does not need to include the parameters of the inverse transformation. When using the model to generate model inputs, the model data includes at least the parameters of the inverse transformation, but does not need to include the parameters of the forward transformation.

[0180] Method 700 may further include the use of a pre-trained model.

[0181] The trained model can be used by obtaining an input instance in an operation titled “Get Input Instances” (720), and then by applying the model to the input instance in an operation titled “Apply Model to Instances” (722), for example, to classify an input image into a class from a set of classes.

[0182] Instead of, or in addition to, operation 720, the trained model may be used in an operation titled “Apply Inverse Model” by applying an inverse model of the trained model (730) to generate a synthetic input instance, such as a synthetic image. This may include, for example, sampling the transform input of a deterministic non-injective transform of the trained model based on the transform output of the transform by a stochastic inverse transform.

[0183] In general, it will be understood that the operations of Method 600 in Figure 6 and Method 700 in Figure 7 can be performed in any suitable order, for example, sequentially, simultaneously, or in combination thereof, where applicable, according to any specific order required by input / output relationships, etc. Some or all of the methods can also be combined; for example, Method 700 using a model under training may be applied after the model has been trained according to Method 600.

[0184] This method can be implemented on a computer as a computer-implemented method, as dedicated hardware, or as a combination of both. As shown in Figure 8, computer instructions, e.g., executable code, can be stored on a computer-readable medium 800, e.g., in the form of a series of machine-readable physical marks 810 and / or as a set of elements having different electrical, e.g., magnetic or optical properties or values. Executable code can be stored in a temporary or non-temporary manner. Examples of computer-readable mediums include memory devices, optical storage devices, integrated circuits, servers, and online software. Figure 8 shows an optical disk 800. Alternatively, the computer-readable medium 800 may include temporary or non-temporary data 810 representing model data representing a model trained according to the method described herein, in particular model data that provides parameters for a forward transformation to apply the model to an input instance to obtain a model output, and / or model data that provides parameters for an inverse transformation to generate an input instance using the inverse model of the model. Model data can include parameters for each transformation of the model, and may indicate the type of transformation, such as an inferential surjective transformation, a generative surjective transformation, a bijective transformation, or a stochastic transformation.

[0185] Examples, embodiments, or any features, whether indicated as non-limiting, should not be understood as limiting the invention as described in the claims.

[0186] The embodiments described above are illustrative, not limiting, of the invention, and it should be noted that those skilled in the art can design many alternative embodiments without departing from the scope of the appended claims. In the claims, reference symbols placed in parentheses should not be construed as limiting the claims. The use of the verb “comprise” and its conjugations does not preclude the existence of elements or stages other than those described in the claims. The article “a” or “an” preceding an element does not preclude the existence of multiple such elements. Expressions such as “at least one” when preceding a list or group of elements represent a selection of all or any subset of elements from the list or group. For example, the expression “at least one of A, B, and C” should be understood to include A only, B only, C only, both A and B, both A and C, both B and C, or all of A, B, and C. The invention can be implemented by hardware comprising multiple distinct elements and by a appropriately programmed computer. In claims for a device enumerating multiple means, multiple of these means may be embodied by a single identical hardware item. The mere fact that certain means are described in different dependent claims does not mean that combinations of these means cannot be used advantageously.

Claims

1. A computer-implemented method (600) for training an image classifier, wherein the image classifier is configured to classify an input image into a class from a set of classes, and the method is The steps (610) include accessing a training dataset which includes at least one labeled training image and / or at least one unlabeled training image labeled with a training class from the set of classes, A step (620) of defining an inverse model of the image classifier, wherein the image classifier includes a set of transformations, the set of transformations includes at least one deterministic non-injective transformation, and the inverse transformation of the transformation is approximated in the inverse model by a stochastic inverse transformation. The process (630) involves training the image classifier using log-likelihood optimization, The training process includes, The process of selecting training images from the aforementioned training dataset (632), A step (634) of applying the image classifier to the training image, wherein the application step (634) includes a step of applying the transformation to the transformation input of the transformation and obtaining the transformation output of the transformation, Given the conversion output, the process (636) involves determining the likelihood contribution of the conversion based on the probability that the inverse stochastic conversion generates the conversion input, If the training image is a labeled training image, the process (638) involves using the likelihood contribution to determine the log-likelihood of the labeled training image and its label according to the joint probability distribution of the input image and the class determined by the image classifier, If the training image is an unlabeled training image, the steps (639) are to use the obtained likelihood contribution to determine the log-likelihood of the unlabeled training image according to the probability distribution of the input image generated by the inverse model, Method (600), including.

2. The method (600) according to claim 1, wherein the step of obtaining the log-likelihood of the training image includes the step of obtaining the sum of the likelihood contributions to each transformation in the set of transformations.

3. The method (600) of claim 2, wherein the image classifier includes tightly coupled components given by a linear bijective transform and a slice transform, the slice transform is configured to select a subset of the outputs of the linear bijective transform, and the inverse transform of the slice transform is approximated by the inverse model by a stochastic inverse transform configured to sample the unselected outputs of the linear bijective transform based on the selected outputs of the linear bijective transform.

4. The method (600) according to claim 2 or 3, wherein the image classifier includes a combined transform configured to obtain a first transform output by combining the first transform input with a first function of the second transform input, and obtain the first and second transform outputs by combining the second transform input with a second function of the first transform output, wherein the first and second functions are convolutions.

5. The method (600) according to any one of claims 1 to 4, wherein the image classifier includes a maximum pooling transform that calculates the transform output as the maximum value of a plurality of transform inputs, and the inverse transform of the maximum pooling transform is approximated by the inverse model by an inverse transform configured to sample the index of the maximum transform input and the values ​​of the non-maximum transform inputs.

6. The method (600) according to any one of claims 1 to 5, wherein the image classifier includes a ReLU transform configured to calculate a transform output by mapping a transform input from intervals to constants, the inverse transform of the ReLU transform being approximated by the inverse model by an inverse transform configured to sample the transform input from predetermined intervals when given a transform output equal to a predetermined constant.

7. The method (600) according to any one of claims 1 to 6, wherein the image classifier is configured to classify the input image into the class by obtaining a vector of class probabilities for each class, and in the output layer, by obtaining the class from the vector, and the inverse transform of the output layer is approximated by the inverse model based on the conditional probability distribution of the vector of class probabilities given the obtained class.

8. The method (600) according to any one of claims 1 to 7, wherein the image classifier further includes a stochastic transform with a deterministic inverse transform, the method comprising the step of calculating the likelihood contribution of the stochastic transform based on the probability that the transform produces the transform output of the stochastic transform given the transform input of the stochastic transform.

9. A system (100) for training an image classifier, wherein the image classifier is configured to classify an input image into a class from a set of classes, and the system is A data interface (120) for accessing a training dataset (030), wherein the training dataset includes at least one labeled training image and / or at least one unlabeled training image labeled with a training class from the set of classes, A processor subsystem (140), A step of defining an inverse model of the image classifier, wherein the image classifier includes a set of transformations, the set of transformations includes at least one deterministic non-injective transformation, and the inverse transformation of the transformation is approximated in the inverse model by a stochastic inverse transformation. A step of training the image classifier using log-likelihood optimization, A processor subsystem (140) configured to perform the following: The training process includes, The steps include selecting training images from the aforementioned training dataset, A step of applying the image classifier to the training images, wherein the application step includes a step of applying the transformation to the transformation input of the transformation and obtaining the transformation output of the transformation, Given the conversion output, the process involves determining the likelihood contribution of the conversion based on the probability that the inverse stochastic conversion generates the conversion input, If the training image is a labeled training image, the process involves using the obtained likelihood contribution to calculate the log-likelihood of the labeled training image and its label according to the joint probability distribution of the input image and the class determined by the image classifier. If the training image is an unlabeled training image, the process involves using the obtained likelihood contribution to calculate the log-likelihood of the unlabeled training image according to the probability distribution of the input image generated by the inverse model. A system (100) including this.

Citation Information

Patent Citations

  • Apparatus, method and program for automatically classifying content

    JP2011145951A

  • Relevance Score Assignment for Artificial Neural Networks

    JP2018513507A

  • Learning program, learning method and learning device

    JP2019159576A

  • Systems and methods for generating data interpretations in neural networks and related systems

    JP2019520655A

  • Method of creating training data, computer and program

    JP2020052936A