Training the generative model and the discriminative model

By using the discriminant model in GAN training to provide detailed feedback on some discriminator scores of the input instances, and updating the gradient of the generated model, it solves the problem that the quality of the generated model is difficult to evaluate and improve during the GAN training process, and achieves faster and more robust training effects and visual feedback.

CN112241784BActive Publication Date: 2025-07-22ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010690706.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-19
Filing Date
2020-07-17
Publication Date
2025-07-22
Estimated Expiration
2040-07-17

AI Technical Summary

Technical Problem

The existing generative adversarial network (GAN) training methods are difficult to effectively train generative models and discriminative models, which makes it difficult to evaluate and improve the quality of synthetic instances generated by generative models, and the training process is easily stuck, and there is a lack of an effective feedback mechanism.

Method used

By using the discriminator scores on multiple parts of the input instance during the training of the generative model, providing detailed-level feedback, updating the gradient of the loss function to improve the training of the generative model, the specific methods include backpropagation and partial derivative adjustment based on the discriminator score.

Benefits of technology

Faster, more robust and interpretable training of generative models is achieved, the quality of generating synthetic instances is improved, and visual feedback on the training process is provided to help optimize training progress.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112241784B_ABST
    Figure CN112241784B_ABST
Patent Text Reader

Abstract

A system (100) for training a generative model and a discriminative model is disclosed. The generative model generates synthetic instances from latent feature vectors by generating intermediate representations from the latent feature vectors and generating synthetic instances from the intermediate representations. The discriminative model determines a plurality of discriminator scores for a plurality of parts of an input instance, which indicate whether the parts are from synthetic instances or actual instances. The generative model is trained by backpropagation. During backpropagation, the partial derivatives of the loss with respect to the entries of the intermediate representation are updated based on the discriminator scores for the parts of the synthetic instances, wherein the parts of the synthetic instances are generated based at least in part on the entries of the intermediate representation, and wherein the value of the partial derivative is decreased if the discriminator score indicates an actual instance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to systems for training generative models and discriminative models and to corresponding computer-implemented methods. The present invention also relates to a computer-readable medium comprising instructions for performing methods and / or parameters of generative models and / or discriminative models. Background Art

[0002] When applying machine learning models in the field of motor vehicle perception systems, the main challenge is the availability of training and test data. For example, such data can include image data and various other types of sensor data, which can be fused to build a 360-degree view around the vehicle. The more such data is available, the better the training and testing that can be performed. Unfortunately, it is difficult to obtain such data. In fact, obtaining real data requires testing motor vehicle perception systems in actual traffic situations. Not only is it expensive to drive around test vehicles equipped with such systems, but it is also dangerous if the decisions are based on motor vehicle perception systems that have not been fully trained. In addition, collecting real data would require collecting data for many different combinations of various parameters, such as amount of daylight, weather, traffic volume, etc. In particular, it is difficult to collect real-world data to perform proper training and testing of extreme cases of such models as near collisions. More generally, in various application domains, and especially when applying machine learning to sensor data, there is a need to efficiently obtain real test and training data.

[0003] The so-called generative adversarial networks (GANs) were proposed in "Generative Adversarial Networks" by I. Goodfellow et al. (available at https: / / arxiv.org / abs / 1406.2661 and incorporated herein by reference). Such GANs include a generative model for generating synthetic data, which can be used, for example, to train or test another machine learning model. The generative model is trained simultaneously with a discriminative model, which estimates the probability that a sample is from the training data rather than the generator. The generative model is trained to maximize the probability that the discriminative model makes a mistake, while the discriminative model is trained to maximize the accuracy of the estimated probability that a sample is from the training data. Promising results have been achieved with generative adversarial networks (GANs). For example, GANs have been shown to be able to reproduce natural-looking images with high resolution and sufficient quality to even deceive human observers.

[0004] Unfortunately, existing methods for training GANs are not satisfactory. Training the discriminative model together with the generative model is generally more difficult than training only the discriminative model. For example, it involves not one but two main components that work adversarially in a zero-sum game and are trained to find a Nash equilibrium. For example, in various cases, early in training, the discriminative model may be very good at distinguishing real instances from synthetic instances, such that the generative model may be too difficult to generate convincing synthetic images. For example, in various cases, it has been observed that the training process of GANs gets stuck reasonably early during training, where the generative model focuses on specific aspects where it fails to make progress. However, it is difficult to detect or correct this: the loss curve of the GAN training process does not provide a good measure of the quality of the generated synthetic instances, and due to the dependence of the generator and discriminator on each other, it is difficult to interpret any apparent patterns in the learning history of the components. This also reduces the confidence that meaningful training of the generator will occur. Summary of the Invention

[0005] According to a first aspect of the present invention, a system for training a generative model and a discriminative model is proposed. According to another aspect of the present invention, a computer-implemented method for training a generative model and a discriminative model is proposed. According to one aspect of the present invention, a computer-readable medium is provided.

[0006] The above aspects of the present invention relate to training a generative model to generate synthetic instances from latent feature vectors, and training a discriminative model to distinguish between synthetic instances and actual instances (such as images, audio waveforms, etc.). The training process may involve repeatedly training the discriminative model to reduce a first loss in distinguishing between actual instances and synthetic instances generated by the generative model, and training the generative model to reduce a second loss in generating synthetic instances that the discriminative model indicates as actual instances.

[0007] Interestingly, the discriminative model can be configured to determine multiple discriminator scores for multiple parts of an input instance. Such discriminator scores for parts of the input instance can indicate whether the part is from a synthetic instance or an actual instance. For example, a common discriminator can be applied to corresponding parts of the input instance, where a final discriminator score for the overall input instance is determined based on the discriminator scores for the corresponding parts, e.g., by averaging. The parts of the input image that affect the discriminator score can be referred to as its receptive field. Preferably, the receptive fields of the respective discriminator scores cover the overall input instance. In the case of an image, the respective discriminator scores can be obtained by applying the discriminator to corresponding regions of the image in a convolutional manner. Typically, the parts of the input instance overlap, e.g., the discriminator scores can be obtained by applying the discriminator to the input instance in a convolutional manner with a stride smaller than the size of the parts. In the context of an image, obtaining the discriminator score itself by averaging the output of the discriminator over a smaller region of the image (referred to as a "patch") is considered a way to force the discriminator to model only the local structure of the image, see, e.g., P. Isola et al., "Image-to-Image Translation with Conditional Adversarial Networks" (available at https: / / arxiv.org / abs / 1611.07004 and incorporated herein by reference).

[0008] However, as recognized by the present inventors, such discriminator scores for parts of the input instance can be used not only to base the final discriminator score, but also to obtain useful feedback regarding the training process. Since traditional training only uses the final discriminator score, the generative model can potentially update all parts of its current synthetic instance in a largely unconstrained manner to convince the discriminative model that it has produced a real sample. Thus, with traditional methods, it may take a long time before the correct modifications are made that cause the discriminative model to start believing in the authenticity of the generated samples.

[0009] However, interestingly, the discriminator scores provide access to an evaluation of the discriminative model at a much higher level of detail than the final discriminator scores: the discriminator scores can be considered as indicating the underlying reasons for classification by the discriminative model. The present inventors have recognized that this detailed information can be utilized to improve the training of the generative model. In fact, if the discriminator scores for parts of the synthesized instance indicate the actual instance, this can indicate that the generator model has performed well in generating that part of the synthesized instance. Thus, it may be beneficial to ensure that the parameters of the generative model affecting that part of the synthesized instance are updated less during training than the parameters of the generative model affecting other parts of the synthesized instance that the discriminator is less confident are real.

[0010] Accordingly, the present inventors have devised a way to incorporate such feedback into the training process. In various cases, the generative model generates a synthesized instance from a latent feature vector by generating an intermediate representation from the latent feature vector and generating the synthesized instance from the intermediate representation. For example, this is the case if the generative model includes a network having one or more convolutional layers for successively refining the synthesized instance. Such a generative model can be trained by backpropagation, where the gradient of the loss with respect to the intermediate representation is computed and then the loss is further backpropagated based on the computed gradient. In traditional methods, such a loss would be based generally on the overall discriminator score of the overall synthesized instance rather than on any individual discriminator scores of its parts.

[0011] However, interestingly, the present inventors have designed to update the gradient of the loss with respect to the intermediate representation of the generative model based on the discriminator scores for the parts of the synthesized instance. Specifically, the partial derivative of the loss with respect to the entries of the intermediate representation can be updated based on the discriminator scores for parts of the synthesized instance that are generated at least in part based on the entries of the intermediate representation. For example, if the discriminator score indicates the actual instance, the value of the partial derivative can be decreased, e.g., decreased compared to its previous value or compared to the value obtained in cases where the discriminator score does not indicate the actual instance. In this way, compared to the parts of the generative model responsible for generating less realistic parts and thus potentially having more room for improvement, the update of the parts of the generative model responsible for generating the realistic parts of the synthesized image can be inhibited. It should be understood that by "decrease", it means a decrease in the absolute value, rather than, for example, making a negative value even more negative.

[0012] Using these measures, faster, more robust, and / or more interpretable training of the generative model can be obtained. Compared to traditional training, using discriminator scores for the parts of the synthetic instance can effectively provide explicit cues as to which parts of the synthetic instance need more modification and which parts need less modification. By using these discriminator scores, the risk of rewriting good parts of the synthetic instance in subsequent training steps can be reduced. This can help the optimization reach its lowest cost point more quickly.

[0013] In addition, the discriminator scores for the parts of the synthetic instance can indicate the progress of the training of the generative model by indicating which parts of the synthetic instance are considered good by the discriminative model and thus indicating which parts of the synthetic instance the generative model is currently focusing on improving. For example, the discriminator scores can be provided to the user of the training system, or to another component that controls the training process, as an explanation of what is happening throughout the training process. Thus, the discriminator scores can provide evidence for learning a meaningful model beyond known metrics (such as mean squared error in the context of cross-entropy or accuracy). For example, the discriminator scores can be used to determine whether the generative model has absorbed specific important parts of the training data, rather than learning a simple proxy for a complex problem in a way that may not be immediately clear to the developer.

[0014] Note that with the proposed measures, various advantages can be obtained in a relatively efficient manner. For example, at a high level, few changes to the existing GAN training process may be required. For example, there may be no need for internal training loops, external saliency methods, explicit attention maps, additional inputs to the generator, etc., thus providing a relatively fast and natural solution.

[0015] Optionally, the model can work on images. The generative model can be configured to generate synthetic images. The discriminator model can be configured to determine discriminator scores indicating the parts of the input image. For example, the parts of the input image can include overlapping rectangular regions of the image. Images are of particular interest not only because they are used in a wide range of applications such as motor vehicle control systems, medical image analysis, etc.; but also because they can use well-performing generative models where intermediate representation entries affect well-defined parts of the synthetic image.

[0016] In the case of images, it is interesting that there may be a spatial correspondence between regions of the image and the discriminator scores they influence. In this sense, the discriminator scores themselves can be regarded as images. For example, if the discriminator scores are determined by moving a fixed-size patch horizontally and vertically over the image, the discriminator scores can be transformed into an image by using the same horizontal and vertical spatial relationships for the discriminator scores. Note that for other modalities that can be modeled as a spatial grid (such as audio), there can also be the same kind of spatial correspondence.

[0017] Optionally, the generative model includes a neural network. Synthetic instances can be computed from an intermediate representation by one or more convolutional layers of the neural network. For example, such convolutional layers can be combined with pooling layers, ReLU layers, and other types of layers known for convolutional neural networks. It is interesting that when convolutional layers are used, there is generally a well-defined relationship between regions of subsequent convolutional layers in terms of how they influence each other. For example, the dimension of the output of a convolutional layer can correspond to one or more instances of the synthetic instance to be generated. For example, a convolutional layer can output one or more images that are the same size or a scaled version of the synthetic image they generate. Since the relationship between subsequent layers can be preserved between layers (such as a spatial relationship in the case of images), the use of convolutional layers allows for a good determination of which part of the synthetic instance is generated by which part of the neural network, which can allow for a particularly efficient use of the feedback in the form of discriminator scores. For example, the intermediate representation can be the output of a convolutional layer. In particular, if the discriminator scores have a similar correspondence to parts of the synthetic image that are the intermediate representation of the generative model, for example, if the discriminative model also includes a convolutional neural network, the feedback mechanism as discussed above can be particularly direct in the sense that after possible scaling, a particular discriminator score corresponds to a particular value of the convolutional layer output.

[0018] Optionally, the intermediate representation of the generative model can be computed in a layer of a neural network (such as a convolutional neural network or other type of neural network). It may be beneficial for this layer to occur relatively far into the neural network. For example, the number of layers before this layer can be greater than or equal to the number of layers after said layer. At least, this layer can be in the last 75% of the layers of the neural network, or even in the last 25% of the layers of the neural network. Using a later layer may be beneficial because it allows a larger part of the neural network to utilize the feedback signal.

[0019] Optionally, the discriminative model includes a neural network, where the corresponding discriminator scores can be calculated by the convolutional layer of the neural network based on the corresponding receptive fields in the synthetic instances. For example, in the case of an image, each spatial unit in the output of the discriminative model can have a well-defined receptive field in the input instance. In this way, as also described above, a well-defined correspondence (e.g., spatial correspondence in the case of an image) between parts of the synthetic instance and the corresponding discriminator scores can be achieved. In fact, in the case of an image, the convolutional layer output can be regarded as the image itself, which can be used as, for example, the input to another neural network, the output presented to the user, etc.

[0020] Optionally, the discriminator scores determined by the discriminative model form a first data volume, and the intermediate representation generated by the generative model forms a corresponding second data volume. For example, the output of the discriminative model can form a first volume having the same spatial dimensions as one or more of the volumes of the intermediate representation. The first and second volumes can be related by scaling. For example, a sub-volume of the first volume of a particular part of the input image having a specific receptive field can correspond by scaling to a sub-volume of the second volume that largely affects the part of the synthetic image corresponding to that receptive field. In this way, a particularly direct feedback mechanism can be achieved, where the use of the discriminator scores leads to particularly relevant updates to the intermediate representation. The gradient of the loss can be updated by scaling the discriminator scores from the first volume to the corresponding second volume and updating the corresponding partial derivatives based on the corresponding scaled discriminator scores. For example, the volume representing the discriminator scores can be scaled up or down to the size of the second volume, where the corresponding scaled-up discriminator scores are used one-to-one to update the corresponding entries for each second volume.

[0021] Optionally, the corresponding partial derivatives are updated based on the corresponding optionally scaled discriminator scores by calculating the Hadamard product (e.g., entry-wise) of the original partial derivatives and the scaled discriminator scores. For example, high values of the discriminator scores can indicate synthetic instances, while low values can indicate real instances. Thereby, a continuous mechanism can be provided, where as the discriminator scores of parts of the generative model more strongly indicate real instances, the updates to the parts of the generative model are more gradually inhibited. Optionally, various functions such as thresholding, non-linear functions, etc. can be applied to the discriminator output before or after calculating the Hadamard product to control the way the training of the generator is affected, e.g., to make the training more or less eager to adapt to the generative model.

[0022] Optionally, the loss function includes a binary cross entropy. Binary cross entropy is often used to train GANs, and is particularly convenient in this case because it provides a metric between 0 and 1 that combines well with various ways of updating partial derivatives (especially Hadamard products). As is known in the art, the binary cross entropy metric can include a regularizer, etc.

[0023] Optionally, the partial derivatives of the loss are updated as a function of the partial derivatives and the discriminator scores. It may be convenient to apply the same function to update the corresponding partial derivatives based on the corresponding discriminator scores, not only because of its simplicity, but also because it may allow parallelization of updates to the corresponding partial derivatives.

[0024] Optionally, multiple partial discriminator scores for input instances together form a discriminator instance. One or more discriminator instances can be output to the user in a sensory perceptible manner. For example, as described above, there may be a systematic relationship between the discriminator scores and the parts of the input instances for which they are determined, such as a spatial relationship in the case of an image. In this case, for example, the discriminator instance can be output, for example, shown on the screen, next to the input instance for which it is determined, for example, both are shown together, or one can be superimposed on the other. The discriminator instance can be scaled so that its size corresponds to the size of the input instance. Therefore, a visualization of the GAN training process (or other types of sensory perceptible output) can be effectively provided. As training proceeds, such visualization can show how the discriminant model distinguishes between actual instances and synthetic instances, and then shows how the generated model tries to deceive the discriminant model. Therefore, providing such a visualization over time provides the model developer with insight into the internals of the training mechanism in terms of the dynamic interaction between the generator and the discriminator, thereby allowing the training process to be corrected by changing the architecture of the generation model and / or the discriminant model, obtaining additional test data, etc. Discriminator instances have the additional advantage that they are searchable, e.g., it is possible to browse through the discriminator instances to find parts where the outputs generally indicate synthetic instances or generally indicate real instances.

[0025] Optionally, partial discriminator scores for portions of an image determined during training may be aggregated, for example to determine the amount of attention the training process has spent on that portion of the generated image. For example, the partial discriminator scores may be averaged to provide such a metric. As an example, in various settings, the generated synthetic images may often be very similar to each other. If the discriminator score indicates that a portion of the image is an actual instance, this may indicate that the training process was not a primary causal factor in how that portion appeared in the synthetic instance. This may provide a particularly good indication of whether meaningful training has occurred that may not be well visible from the synthetic instance itself.

[0026] Optionally, discriminator instances can be output by mapping discriminator scores to the parts of the input instances from which they were determined and outputting the mapped discriminator scores. For example, an image with the size of the input image can be shown, where the pixels that result in high (or low) discriminator scores are highlighted to highlight the parts of the image that the generator was good (or bad) at generating. Alternatively, the input instance can be shown where the pixels of the original instance that result in high (or low) discriminator scores are removed (e.g., blacked out).

[0027] Optionally, user input is obtained and, based on the user input, training is reset or continued. During the training of a GAN, it is not uncommon to become stuck in cases where the generative model stops improving before it can generate convincing synthetic instances, e.g., because the generative model focuses on parts of the synthetic instances that are too difficult for it to fool the discriminative model. In practice, developers typically use manual inspection of a batch of generated images at different stages of the training pipeline to determine whether it makes sense to continue training. By showing the discriminator instances to the user, the user obtains better feedback on the internal state of the training process to use as a basis for that decision, thus allowing the user to better control the training process being performed.

[0028] Optionally, the generative model is used to generate one or more synthetic instances. One possible use of such generated synthetic instances can be as training and / or test data for additional machine learning models such as neural networks, etc. Then, one or more synthetic instances can be used as test and / or training data to train an additional machine learning model. As an example, the additional machine learning model can be a classification model or a regression model. In such cases, for example, labels can be obtained for one or more synthetic instances, and the additional machine learning model can be trained based on the one or more synthetic instances and the labels. Thus, the generative model improved with the various features described herein can be used to improve the training of an additional machine learning model such as an image classifier and ultimately obtain a more accurate model output for classification, for example. The additional machine learning model does not need to be a classification or regression model: for example, the training dataset can also be used for unsupervised learning, in which case labels may not be required.

[0029] It is also conceivable to have various other applications of generative and / or discriminative models. In some embodiments, generative and discriminative models can be used to determine anomaly scores for detecting anomalies, for example, in multivariate time series of networks of sensors and actuators, as in "MAD-GAN: Multivariate Anomaly Detection for Time Series Data with Generative Adversarial Networks" by D. Li et al. (available at https: / / arxiv.org / abs / 1901.04997 and incorporated herein by reference). In other embodiments, generative models can be used for data completion. For example, in autonomous driving, various types of sensor data can be fused to build a 360-degree view around the vehicle. In such cases where the fields of view of different types of sensors are different, a generative model can be trained and used to synthesize missing sensor data outside the field of view of a sensor based on the sensor data of another sensor.

[0030] It should be noted that gradient updates can be more generally used to control the training process of the generator. For example, in some or all iterations of training the generative model, the partial derivatives can be updated based not on the discriminator scores of the discriminative model but on scores obtained in other ways (e.g., provided by a user) to control which specific or semantic parts of the synthetic instances the training of the generative model is to address at that point in time.

[0031] Those skilled in the art should understand that two or more of the above-described embodiments, implementations, and / or alternative aspects of the present invention can be combined in any manner deemed useful.

[0032] Those skilled in the art can perform any computer-implemented method and / or modification of any computer-readable medium corresponding to the described modifications and variations of the corresponding system based on this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] These and other aspects of the present invention will become apparent and be further elucidated from the following embodiments described by way of example and with reference to the accompanying drawings, in which:

[0034] Figure 1 A system for training generative and discriminative models is shown;

[0035] Figure 2 A system for using generative and / or discriminative models is shown;

[0036] Figure 3 A detailed example of how generative and discriminative models can be trained and used is shown;

[0037] Figure 4 A computer-implemented method for training a generative model and a discriminative model is shown;

[0038] Figure 5 A computer-readable medium including data is shown.

[0039] It should be noted that the drawings are purely schematic and not drawn to scale. In the drawings, elements corresponding to elements already described may have the same reference numerals. Detailed Description

[0040] Figure 1 A system 100 for training a generative model and a discriminative model is shown. The generative model can be configured to generate synthetic instances from latent feature vectors, while the discriminative model can be configured to determine a plurality of discriminator scores for a plurality of parts of an input instance. The discriminator scores for the parts of the input instance can indicate whether the part is from a synthetic instance or an actual instance. The system 100 can include a data interface 120 and a processor subsystem 140, which can communicate internally via a data communication 124. The data interface 120 can be used to access a set of actual instances 030 and parameters 041 of the generative model and parameters 042 of the discriminative model.

[0041] The processor subsystem 140 can be configured to access data 030, 041, 042 during operation of the system 100 and using the data interface 120. For example, as Figure 1 shown, the data interface 120 can provide access 122 to an external data storage device 020 that can include the data 030, 041, 042. Alternatively, the data 030, 041, 042 can be accessed from an internal data storage device that is part of the system 100. Alternatively, the data 030, 041, 042 can be received from another entity via a network. Generally, the data interface 120 can take various forms, such as a network interface to a local area network or a wide area network (e.g., the Internet), a storage interface to an internal or external data storage device, etc. The data storage device 020 can take any known and appropriate form.

[0042] The processor subsystem 140 can be configured to learn the parameters 041 of the generative model and the parameters 042 of the discriminative model during operation of the system 100 and using the data interface 120 by repeatedly training the discriminative model to reduce a first loss in distinguishing between actual instances 030 and synthetic instances generated by the generative model and training the generative model to reduce a second loss of generating synthetic instances that the discriminative model indicates as actual instances. The generative model can be configured to generate synthetic instances from latent feature vectors by generating an intermediate representation from the latent feature vectors and generating synthetic instances from the intermediate representation.

[0043] The processor subsystem 140 can be configured to train a generative model by a second loss of synthetic instances generated from latent feature vectors through backpropagation. The backpropagation can include using a discriminative model to determine a plurality of discriminator scores for a plurality of parts of the synthetic instance. The backpropagation can also include calculating the gradient of the loss with respect to an intermediate representation. The gradient can include partial derivatives of the loss with respect to the entries of the intermediate representation. The backpropagation can also include updating the partial derivatives of the loss with respect to the entries of the intermediate representation based on the discriminator scores for the parts of the synthetic instance. The parts of the synthetic instance can be generated at least in part based on the entries of the intermediate representation. If the discriminator scores indicate an actual instance, the value of the partial derivative can be decreased. The backpropagation can also include further backpropagating the loss based on the updated gradient.

[0044] As an optional component, the system 100 can include an image input interface 160 or any other type of input interface for obtaining sensor data from a sensor such as a camera 170. The sensor data can be part of a set of actual instances 030. For example, the camera can be configured to capture image data 162, and the processor subsystem 140 is configured to obtain the image data 162 obtained via the input interface 160 and store it as part of a set of actual instances 030. The input interface can be configured for various types of sensor signals, e.g., video signals, radar / LiDAR signals, ultrasonic signals, etc.

[0045] As an optional component, the system 100 can include a display output interface 180 or any other type of output interface for outputting one or more discriminator instances to a rendering device such as a display 190. For example, the display output interface 180 can generate display data 182 for the display 190, which causes the display 190 to present one or more discriminator instances in a sensorily perceivable manner, e.g., as a screen visualization 192.

[0046] Reference will be made to Figure 3 Further elaborate on the details and aspects of the operation of the system 100, including its optional aspects.

[0047] Typically, system 100 can be implemented as a single device or apparatus, or be implemented within a single device or apparatus, such as, for example, a laptop- or desktop-based workstation, or a server. The device or apparatus may include one or more microprocessors that execute appropriate software. For example, the processor subsystem may be implemented by a single central processing unit (CPU), but may also be implemented by a combination or system of such CPUs and / or other types of processing units. The software may have been downloaded and / or stored in a corresponding memory, such as, for example, volatile memory such as RAM or non-volatile memory such as flash memory. Alternatively, the functional units of the system (such as the data interface and the processor subsystem) may be implemented in the device or apparatus in the form of programmable logic, such as, for example, implemented as a field-programmable gate array (FPGA) and / or a graphics processing unit (GPU). Typically, each functional unit of the system may be implemented in the form of a circuit. Note that system 100 may also be implemented in a distributed manner, such as, for example, involving different devices or apparatuses, such as distributed servers, for example, in the form of cloud computing.

[0048] Figure 2 System 200 for applying a generative model and / or a discriminative model trained, for example, by system 100 is shown. System 200 may include a data interface 220 and a processor subsystem 240, which may communicate internally via data communication 224. The data interface 220 may be used to access the parameters 041 of the generative model and / or the parameters 042 of the discriminative model. Similar to system 100, the processor subsystem 240 of system 200 may be configured to access the parameters 041 and / or 042 during operation of system 200 and using the data interface 220, such as, for example, via access 222 to an external data storage device 022 or otherwise. Similar to system 100, system 200 may be implemented as a single device or apparatus, or may be implemented in a distributed manner.

[0049] The processor subsystem 240 may be configured to apply the generative model and / or the discriminative model during operation of the system. For example, the processor subsystem 240 may use the generative model to generate one or more synthetic instances; obtain labels for the one or more synthetic instances; and train a classification model based on the one or more synthetic instances and the labels.

[0050] As a specific example, system 200 can be included in a motor vehicle perception system, such as a sub-component of an autonomous vehicle. For example, an autonomous vehicle can be controlled at least in part based on data generated by a generative model. For example, the generative model can be used to synthesize missing sensor data outside the field of view of an image sensor of the system. Based on the sensor data and the synthesized data, an electric motor can be controlled. Generally, such a data synthesis system can be part of a physical entity such as a vehicle, a robot, etc., or a connected or distributed system of physical entities, such as a lighting system or any other type of physical system, such as a building.

[0051] Reference will be made to Figure 3 further elaborate on the details and aspects of the operation of system 200, including its optional aspects.

[0052] Figure 3 A detailed but non-limiting example of training and using a generative model and a discriminative model is shown.

[0053] In this example, a generative model GM 350 parameterized by a set of parameters GPAR 341 is shown. The generative model GM can be configured to generate a synthetic instance SI 353 from a latent feature vector LFV 351. In this particular example, the synthetic instance SI can be a synthetic image, e.g., an image of size M×N represented by M×N×c real numbers, where c is the number of channels of the image, e.g., 1 for grayscale, and 3 for RGB, etc. Note that various other types of synthetic instances SI are also possible, such as audio data like a spectral range.

[0054] The generative model GM can generate the synthetic instance SI from the latent feature vector LFV by generating an intermediate representation IR 352 from the latent feature vector LFV and generating the synthetic instance SI from the intermediate representation IR. As also discussed elsewhere, various types of generative models can use an intermediate representation, e.g., a Gaussian process, a Markov model, etc. In this case, the generative model GM can be a neural network, and more specifically a convolutional neural network. The parameters GPAR include the weights of the neural network. As is known per se in the art, a convolutional neural network can include one or more convolutional layers. Such a convolutional layer typically provides an output volume of size M i ×N i ×d i where di is the number of filters in the layer, and M i ×N i is the size of the output of the filter, which is typically applied convolutionally to the corresponding output of the previous convolutional layer, optionally with various other layers between them, such as a pooling layer or a ReLU layer. The size M i ×N iTypically has the same scale as the synthetic image SI to be generated, although this is not required.

[0055] In this case, the intermediate representation IR is the output of one of the convolutional layers. As such, it is also denoted by M i ×N i ×d i sized volume in this figure. Due to the convolutional nature of the convolutional neural network GM, in this example, it can be seen that the intermediate representation IR includes d i multiplied by M i ×N i ×1 sized data volume. Given coordinates (i, j), a set of entries (i, j, k) of the intermediate representation IR typically all affect the same sub-region of the synthetic instance SI. Due to the convolutional nature, this sub-region can approximately correspond to the region of the synthetic instance SI obtained by scaling the point (i, j) to the region of the synthetic instance.

[0056] Typically, there can be any number L b of convolutional layers between the latent feature vector LFV and the intermediate representation IR, and there can be L a convolutional layers between the intermediate representation IR and the synthetic instance SI. For example, the total number of layers L b +L a can be at most or at least 10, at most or at least 20, or at most or at least 100. The total number of parameters of the generative model can be at most or at least 10000, at most or at least 1000000, or at most or at least 10000000, etc. Preferably, L a ≥L b , for example, L a =L b =5, for example, to ensure that a large enough part of the generative model GM benefits from the feedback provided by the discriminative model, as described below.

[0057] More generally, the generative model GM can sample from a k-dimensional latent space to a tensor that matches the size of the image to be generated. Along this way, an intermediate representation IR can be generated, whose spatial size matches the output of the discriminative model, as described below:

[0058]

[0059] where, typically, in the above convolutional neural network example, m < M and n < N, for example, (m, n, j) = (M i , N i , d i ).

[0060] In addition to the generative model GM that generates the synthetic instance SI, the actual instances AI1, …, Ain 330 and the discriminative model DM 360 parameterized by a set of parameters DPAR 342 are also shown. The discriminative model DM can be used to determine whether a given input instance II 361 is a synthetic instance SI or an actual instance AIi. Specifically, the discriminative model DM can be configured to determine a plurality of discriminator scores DS 363 for a plurality of parts of the input instance II. For example, the discriminator score 364 for the part 365 of the input instance shown in the figure. Such discriminator scores (e.g., discriminator score 364) can indicate whether such a part (e.g., part 365) is from the synthetic instance SI or from an actual instance. For example, where 0 indicates an actual instance and 1 indicates a synthetic instance. As shown by the score 364 and the part 365, there can be a spatial relationship or other type of systematic relationship between the score and the part for which it is determined. For example, here the discriminator score 364 is shown in a grid, where the upper left score 364 is determined for the upper left part 365 of the input image II.

[0061] Similar to the generative model GM, the discriminative model DM in this example can also be a neural network, specifically a convolutional neural network including one or more convolutional layers. For an illustrative example, the output 362 of an intermediate convolutional layer of the neural network is shown. In addition, the discriminator scores DS can be calculated by the convolutional layers of the neural network DM based on the corresponding receptive fields in the synthetic instance. This can result in the spatial relationship between the scores and the parts discussed above. To make the corresponding discriminator scores DS affected by reasonable parts of the input instance, the number of layers of the discriminative model DM can be reasonably large, such as at least 4 layers or at least 8 layers. For example, for a 128×128 input image, the discriminative model DM can include 5 or 6 layers; generally, the number of layers can be at most or at least 10, at most or at least 100, etc. The total number of parameters of the discriminative model can be at most or at least 10000, at most or at least 100000, at most or at least 10000000, etc.

[0062] As a specific example, the discriminative model DM can be designed according to the patch discriminator design disclosed in, for example, "Image-to-Image Translation with Conditional Adversarial Networks" by P. Isola et al. In this design and other designs, the discriminative model DM can take the input image II and output a smaller flattened image, for example:

[0063]

[0064] Similarly, typically, m' < M and n' < N. The output dimensions (m', n') can be equal to the dimensions (m, n) of the intermediate representation IR of the generative model GM, but this is not necessary. For example, the dimensions (m', n') and (m, n) can correspond to each other by scaling, as discussed in more detail below.

[0065] So far, how the discriminative model DM and the generative model GM can be applied has been discussed. The results output of the models can be used for various purposes. For example, considering the discriminator scores 364 together as discriminator instances, one or more such discriminator instances can be output to the user in a perceptible manner. For example, when the convolutional neural network DM is applied to the input image II, the output DS itself can be regarded as a discriminator image that can be output, for example, shown on the screen. Such an output can be used as a debugging output to allow the user to control the training of the generative model GM, and more generally, for example, for anomaly detection as discussed in "MAD-GAN: Multivariate Anomaly Detection for Time Series Data with Generative Adversarial Networks" by D. Li et al. Specifically, the discriminator scores can be output by mapping them to the parts of the input instances from which they are determined and outputting the mapped discriminator scores, and / or can be overlaid on the original input instance II. For example, starting from the input image, pixels that affect the discriminator scores can be changed based on the scores, for example, made darker or lighter.

[0066] The generative model GM can also be used to generate one or more synthetic instances SI, for which labels can be obtained, for example, by manual annotation. After that, a classification model (not shown) can be trained based on the one or more synthetic instances and the labels. The synthetic instances can also be used as training inputs, for example, without such labels, to perform unsupervised learning.

[0067] Proceed to the training of the discriminative model DM and the generative model GM. The parameters GPAR of the generative model and the parameters DPAR of the discriminative model can be learned by repeatedly training the discriminative model DM to reduce a first loss in distinguishing between actual instances AIi and synthetic instances SI generated by the generative model GM, and training the generative model GM to reduce a second loss in generating synthetic instances SI that the discriminative model DM indicates as actual instances. For example, a cycle of the training process can include one or more training iterations of the discriminative model DM, followed by one or more training iterations of the generative model GM. For example, an advanced training method of "Generative Adversarial Networks" proposed by I. Goodfellow et al. can be used. Generally, the training is performed using a stochastic method such as stochastic gradient descent, for example, using the Adam optimizer. The training can be performed on an instance-by-instance basis or in batches (e.g., at most or at least 64 or at most or at least 256 instances). As is known, such optimization methods can be heuristic and / or reach a local optimum.

[0068] The training of the discriminative model DM is shown as operation Td 366 in the figure. In this example, the discriminative model DM can be trained by independently calculating the binary cross-entropy loss to obtain an m×n discriminator output and averaging the results to determine the total loss to be minimized. As a ground truth, for example, a batch of m×n matrices of all ones can be used in the case of synthetic instances of the current training batch, or a batch of m×n matrices of all zeros can be used in the case of actual instances AIi of the current batch. Thus, each of the m×n discriminator outputs can intuitively be responsible for identifying whether the part of the input instance II in its receptive field in the input is an actual instance or a synthetic instance. However, various other ways of training the discriminative model DM with respect to the discriminator output and thereby iteratively determining the parameter set DPAR will be obvious. Compared with a normal GAN, generally m > 1 and / or n > 1, such that multiple discriminator outputs are determined.

[0069] As shown in the figure, the training of the generative model GM can be performed by backpropagating, for example, the second loss of generating synthetic instances that the discriminative model indicates as actual instances. In other words, the loss (e.g., binary cross-entropy loss) can penalize the generation of instances that are easily recognized as synthetic by the discriminative model DM. For one or more synthetic instances SI generated from the latent feature vector LFV, the loss can be backpropagated. Thus, the parameter set GPAR can be iteratively updated to reduce this loss. The backpropagation can include calculating the gradient of the loss with respect to the intermediate representation IR in the backpropagation operation B1354. This backpropagation can be performed using conventional techniques also described above. The gradient includes the partial derivatives of the loss with respect to the entries of the intermediate representation IR.

[0070] Interestingly, the backpropagation shown in the figure may include an update operation UPD 355. The update UPD may be performed on the computed gradient of the synthetic instance SI with respect to the intermediate representation IR based on the discriminator score DS determined by the discriminative model DM for the synthetic instance SI. Thus, the discriminator score DS may be used as a feedback signal to update the gradient. In particular, the partial derivative of the loss with respect to the entry of the intermediate representation IR may be updated based on the discriminator score for the portion of the synthetic instance SI. The portion of the synthetic instance SI may be selected such that it is generated at least in part based on the entry of the intermediate representation IR. In other words, the partial derivative of the loss may be updated based on the discriminator score for the portion of the synthetic instance it affects. For example, the discriminator score may in effect provide feedback on the quality of the portion of the synthetic instance generated at least in part based on the entry of the intermediate representation. As discussed in more detail elsewhere, if the discriminator score indicates an actual instance, the value of the partial derivative may be decreased; in other words, in the case where the discriminator score indicates an actual instance, less backpropagation may be performed than in the case where the discriminator score indicates a synthetic instance.

[0071] The selection of which partial derivative to update based on which discriminator score DS may be performed in various ways. For example, for some or all discriminator scores DS, one or more corresponding entries of the intermediate representation IR may be updated. It may be beneficial to update all entries of the intermediate representation IR that affect the discriminator score DS, but it is generally beneficial to select only those entries that most strongly (e.g., most directly) affect the discriminator score DS. For example, in the case of a convolutional neural network, the entries of the intermediate representation IR may correspond to the discriminator score DS as corresponding sub-volumes (e.g., corresponding image regions). In such a convolutional neural network, due to their convolutional nature, sub-volumes, for example in the case of an image, most strongly affect the corresponding sub-volumes of the next layer of the convolutional network spatially, but may also less strongly affect neighboring sub-volumes. In such a case, to improve the quality of the feedback signal, it may be beneficial to update only those portions of the intermediate representation IR that more strongly affect the portion (e.g., corresponding sub-volume) of the synthetic instance for which the discriminator score DS is determined.

[0072] In some embodiments, for each discriminator score, exactly one corresponding entry of the intermediate representation may be updated. However, this is not necessary. For example, the discriminator score DS may be used to update multiple entries of the intermediate representation IR and / or the entries of the intermediate representation IR may be updated based on multiple discriminator scores DS in, for example, multiple update operations. For example, the receptive fields of multiple discriminator scores DS in the synthetic instance SI may overlap, which may result in different entries of the intermediate representation IR affecting the same discriminator score DS.

[0073] As a specific example, the case where the discriminative model DM and / or the generative model GM includes a convolutional neural network is now discussed. As described above, such a network may include multiple intermediate representations representing a data volume of a corresponding size M i ×N i ×d i where d i is the number of filters applied, and M i ×N i is the size of the output of the filter. Starting from an input image of size M×N×c, the discriminative model DM can, for example, use one or more convolutional layers with corresponding intermediate representations to end at a discriminator instance, e.g., a set of discriminator scores DS forming a first data volume, e.g., a layer of size m'×n' representing a data volume of size m'×n'. Similarly, the generative model GM can start from a latent feature vector of size k, use multiple intermediate representations to end at a synthetic image of size M×N×c. In particular, the intermediate representation 362 can be computed in the layers of the network GM as one or more corresponding second data volumes, e.g., including M i of number d i ×N i second data volumes of size M i ×N i ×d i size layers. Such data volumes can also appear in other types of models besides convolutional neural networks, but for ease of explanation, convolutional neural networks are used as an example here.

[0074] In the case where the intermediate representation IR and the discriminator scores DS form corresponding data volumes, the entries of the intermediate representation IR can be updated based on the corresponding discriminator scores DS by scaling the discriminator scores DS from the first volume to a corresponding second volume and updating the corresponding partial derivatives based on the corresponding scaled discriminator scores. Depending on the situation, the scaling can be upscaling or downscaling. In many cases, the corresponding second volumes may each have the same size, in which case a single scaling may be sufficient, but corresponding scalings relative to the corresponding second volumes can also be performed. Of course, if the first and second volumes have the same size, the scaling can be skipped.

[0075] Now, the case where the sizes of the first and second volumes do not exactly match will be described. For example, the discriminative model DM can start with an input instance of size 100×100×3, such as an RGB image representing a 100×100 size. After one or more operations, the intermediate representation can be determined to have a size of 10×10×20, which, for example, corresponds to 20 convolutional neural network filters, each with a size of 10×10. To determine the discriminator output DS, the discriminative model DM can apply an operation such as a 2D convolution with a filter window size of 1 to produce a volume of discriminator scores DS of size 10×10×1, for example. Thus, each of these discriminator scores can have a corresponding receptive field of a specific size in the input. The intermediate representation IR of the generative mode GM can, for example, have a volume of 10×10×N. In this case, scaling can be skipped, and each entry (i, j, k) of the intermediate representation IR can be updated based on the corresponding entry (i, j) of the discriminator scores DS. However, it is also possible for the generative model GM to generate an intermediate representation IR of a different size (such as size 5×5×N). In this case, the update of the gradient of the loss can include scaling the first volume (in this case, size 10×10×1) to the second volume (in this case, 5×5×1), and updating each entry (i, j, k) of the intermediate representation IR based on the corresponding entry (i, j) of the rescaled volume.

[0076] Regarding how the partial derivatives are updated, as described above, if the discriminator scores DS indicate an actual instance, the value of the partial derivative can be decreased, thus achieving stronger feedback for seemingly synthetic instances than for seemingly actual instances. This can be done by possibly updating the partial derivative of the loss based on the partial derivative itself and the discriminator scores after rescaling.

[0077] As an example, the discriminative model DM can be trained to provide discriminator scores ranging from 0 to 1 indicating whether a part of instance 365 looks like an actual instance or a synthetic instance, for example, by training the discriminative model DM using a binary cross-entropy loss function, etc. For example, a value of 0 can indicate an actual instance, while a value of 1 can indicate a synthetic instance. In such a case, it can be considered that the discriminator scores DS represent the degree to which the gradient of the generative model GM can be modulated at that specific location. For example, the generative model GM can be prevented from changing parts of the discriminative model DM that are, for example, uncertain that the input is a fake synthetic instance SI, and focus on minimizing the GAN loss by working on the locations where the discriminative model is more certain that they are synthetic.

[0078] In some particularly appealing embodiments, the corresponding partial derivatives are updated based on the respective optionally scaled discriminator scores by computing the Hadamard (e.g., element-wise) product of the raw partial derivatives and the (scaled) discriminator scores. For example, the Hadamard (element-wise) product can be computed between discriminator scores DS of size m×n×1 and the gradient volume of size m×n×j produced by the intermediate representation IR. Intuitively, if the discriminator score is 1 at a given location, e.g., which indicates a synthetic instance, the gradient can flow backward to reconfigure the parameters GPAR, such as neural network weights, which can change the output of the generative model GM at that location. If the discriminator score is 0, e.g., the discriminative model believes the sample is real, the parameters GPAR can remain unchanged.

[0079] Instead of or in addition to the Hadamard product, various other functions can be used to update the partial derivatives. Varying the function can allow control over how eager one is to make the generative model GM based on the output predictions of the discriminative model DM. For example, a threshold can be applied, e.g., given the output of the discriminative model represented as a floating point number between 0 and 1, a threshold at 0.5 etc. can be applied to ensure that only values below the threshold result in updated partial derivatives. Note that controlling eagerness can be particularly relevant given the difference between modulating based on discriminator scores indicating the discriminative model's certainty that it sees an actual instance versus discriminator scores indicating the model's uncertainty that it sees a synthetic instance. Specifically, note that in this context, the theoretical optimum of the discriminator in a GAN may not be that it can perfectly discriminate real images from fake images, but rather, e.g., that it is 50% certain whether the input image is real or fake.

[0080] After the update UPD, in operation B2356, based on the updated gradients, the loss can be further backpropagated. For example, this can be performed according to the advanced training methods of I. Goodfellow et al. "Generative Adversarial Networks". Regularization and / or custom GAN loss functions can be applied as is known in the art itself.

[0081] Interestingly, in some embodiments, during training, one or more discriminator instances are presented to a user, e.g., shown on a screen. User input can be obtained and, based on the user input, training can be reset or continued. Specifically, the discriminator instance can be shown beside or superimposed on the instance for which it is determined. Interestingly, discriminator instance DS can not only highlight which parts of the generated synthetic instance SI are considered convincing by the discriminative model DM, but also, due to the feedback mechanism of the updated partial derivative, which parts of the synthetic instance SI the training process is currently focusing on. Thus, particularly effective feedback on the internal workings of the training process can be provided to the user, who can then adjust the training process accordingly, e.g., by stopping training if insufficient progress is being made. For example, multiple training processes can be active simultaneously, where the user selects which training processes to continue. However, the feedback can also be used for various other purposes, e.g., to inform the generation of additional training data, etc. Thereby, a generally more efficient training process can be achieved.

[0082] Figure 4 FIG. 4 shows a block diagram of a computer-implemented method 400 for training a generative model and a discriminative model. The generative model can be configured to generate synthetic instances from latent feature vectors. The discriminative model can be configured to determine a plurality of discriminator scores for a plurality of parts of an input instance. The discriminator score for a part of the input instance can indicate whether that part is from a synthetic instance or an actual instance. Method 400 can correspond to Figure 1 the operation of system 100. However, this is not limiting, as method 400 can also be performed using another system, device, or apparatus.

[0083] Method 400 can include: in an operation entitled "ACCESSING ACTUAL INSTANCES, PARAMETERS", accessing 410 a set of actual instances, as well as the parameters of the generative model and the discriminative model. Method 400 can also include learning the parameters of the generative model and the discriminative model. The learning can be performed by repeatedly training discriminative model 420 in an operation entitled "TRAINING DISCRIMINATIVE MODEL" to reduce a first loss in distinguishing between actual instances and synthetic instances generated by the generative model; and training generative model 430 in an operation entitled "TRAINING GENERATIVE MODEL" to reduce a second loss of generating synthetic instances that the discriminative model indicates as actual instances. The generative model can be configured to generate synthetic instances from latent feature vectors by generating an intermediate representation from the latent feature vectors and generating synthetic instances from the intermediate representation.

[0084] Training of the generative model 430 can include backpropagating a second loss of synthetic instances generated from latent feature vectors. The backpropagation can include computing 432 the gradient of the loss with respect to an intermediate representation in an operation entitled "COMPUTING GRADIENT W.R.T. INTERMEDIATE REPRESENTATION". The gradient can include partial derivatives of the loss with respect to entries of the intermediate representation. The backpropagation can also include using a discriminative model to determine 434 a plurality of discriminator scores for a plurality of parts of a synthetic instance in an operation entitled "DETERMINING DISCRIMINATOR SCORES". The backpropagation can also include updating 436 the partial derivatives of the loss with respect to entries of the intermediate representation based on the discriminator scores for parts of the synthetic instance. The parts of the synthetic instance can be generated at least in part based on entries of the intermediate representation. If the discriminator scores indicate an actual instance, the value of the partial derivative can be decreased. The backpropagation can also include further backpropagating 438 the loss based on updated gradients that adapt a base classifier to one or more new classifications in an operation entitled "FUTHER BACKPROPAGATING".

[0085] It should be understood that, generally, Figure 4 the operations of method 400 can be performed in any suitable order, e.g., sequentially, simultaneously, or a combination thereof, subject in appropriate cases to a particular order required, e.g., by input / output relationships. The (one or more) methods can be implemented on a computer as a computer-implemented method, as dedicated hardware, or as a combination of both. Also as Figure 5 shown, instructions (e.g., executable code) for a computer can be stored on a computer-readable medium 500, e.g., in the form of a sequence 510 of machine-readable physical markings and / or as a sequence of elements having different electrical, e.g., magnetic or optical, properties or values. The executable code can be stored in a transient or non-transient manner. Examples of computer-readable media include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Figure 5 A compact disc 500 is shown. As an alternative, the computer-readable medium 500 can include transient or non-transient data 510 representing parameters of a generative model and / or a discriminative model as described elsewhere in this specification.

[0086] Whether or not indicated as non-limiting, examples, embodiments, or alternative features should not be construed as limiting the claimed invention.

[0087] It should be noted that the above embodiments are illustrative and not restrictive of the present invention, and those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In a claim, any reference signs placed between parentheses shall not be construed as limiting the claim. The use of the verb "comprise" and its conjugations does not exclude the presence of elements or steps other than those recited in a claim. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. When used before a list or group of elements, expressions such as "at least one" represent any subset of all elements or elements selected from the list or group. For example, the expression "at least one of A, B, and C" should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The present invention can be implemented by hardware comprising several discrete elements and by a suitably programmed computer. In a device claim enumerating several components, several of these components may be implemented by the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously.

Claims

1. A system (100) for training a generative model and a discriminative model, wherein, The generative model is configured to generate synthetic instances from latent feature vectors, and the discriminative model is configured to determine a plurality of discriminator scores for a plurality of parts of an input instance, the discriminator scores for the parts of the input instance indicating whether the parts are from a synthetic instance or an actual instance. The system includes: - A data interface (120) for accessing a set of actual instances (030) and the parameters (041) of the generative model and the parameters (042) of the discriminative model; and - A processor (140) configured to learn the parameters of the generative model and the discriminative model by repeatedly training the discriminative model to reduce a first loss in distinguishing between the actual instances and the synthetic instances generated by the generative model and training the generative model to reduce a second loss of generating synthetic instances that the discriminative model indicates as actual instances, where: - The generative model is configured to generate synthetic instances from the latent feature vectors by generating an intermediate representation from the latent feature vectors and generating synthetic instances from the intermediate representation; and - The processor (140) is configured to train the generative model by backpropagating the second loss of the synthetic instances generated from the latent feature vectors, by: - Using the discriminative model to determine a plurality of discriminator scores for the plurality of parts of the synthetic instance; - Calculating the gradient of the loss with respect to the intermediate representation, the gradient including the partial derivatives of the loss with respect to the entries of the intermediate representation; - Updating the partial derivatives of the loss with respect to the entries of the intermediate representation based on the discriminator scores for the parts of the synthetic instance, where the parts of the synthetic instance are generated at least in part based on the entries of the intermediate representation, and where, if the discriminator scores indicate an actual instance, the value of the partial derivative is decreased; and - Further backpropagating the loss based on the updated gradient, The generative model is configured to generate synthetic images, and the discriminative model is configured to determine discriminator scores for the parts of an input image.

2. The system (100) according to claim 1, wherein, The generative model includes a neural network, and the processor (140) is configured to compute the synthetic instance from the intermediate representation through one or more convolutional layers of the neural network.

3. The system (100) according to claim 2, wherein, The processor (140) is configured to compute the intermediate representation in a layer of the neural network, and the number of layers after the layer is greater than or equal to the number of layers before the layer.

4. The system (100) according to any one of claims 1 to 3, wherein, The discriminative model includes a neural network, and the processor (140) is configured to compute corresponding discriminator scores based on corresponding receptive fields in the synthetic instance through the convolutional layers of the neural network.

5. The system (100) according to any one of claims 1 to 3, wherein, The discriminator scores determined by the discriminative model form a first data volume, and the intermediate representation generated by the generative model forms a corresponding second data volume. The processor (140) is configured to update the gradient of the loss by scaling the discriminator scores from the first volume to the corresponding second volume and updating the corresponding partial derivatives based on the corresponding scaled discriminator scores.

6. The system (100) according to claim 5, wherein, The processor (140) is configured to update the respective partial derivatives based on the respective scaled discriminator scores by computing the Hadamard product of the raw partial derivatives and the scaled discriminator scores.

7. The system (100) according to any one of claims 1 to 3, wherein, The loss function includes binary cross entropy.

8. The system (100) according to any one of claims 1 to 3, wherein, The processor (140) is configured to update the partial derivatives of the loss according to the partial derivatives and the discriminator scores.

9. The system (100) according to any one of claims 1 to 3, wherein, The plurality of discriminator scores for the input instance together form a discriminator instance, and the system further includes an output interface (180) configured to output one or more discriminator instances to a user in a perceptible manner.

10. The system (100) according to claim 9, wherein, The processor (140) is configured to output the discriminator instance by mapping the discriminator scores to the portions of the input instance from which they were determined and outputting the mapped discriminator scores.

11. The system (100) according to claim 9, wherein, The processor (140) is further configured to obtain an input from the user and reset or continue the training according to the input of the user.

12. A computer-implemented method (400) for training a generative model and a discriminative model, wherein, The generative model is configured to generate synthetic instances from latent feature vectors, and the discriminative model is configured to determine a plurality of discriminator scores for a plurality of portions of an input instance, the discriminator scores for the portions of the input instance indicating whether the portions are from a synthetic instance or an actual instance. The method includes: - accessing (410) a set of actual instances and the parameters of the generative model and the discriminative model; and - learning the parameters of the generative model and the discriminative model by repeatedly training the discriminative model (420) to reduce a first loss in distinguishing between the actual instances and synthetic instances generated by the generative model and training the generative model (430) to reduce a second loss of generating synthetic instances that the discriminative model indicates are actual instances, where: - the generative model is configured to generate synthetic instances from the latent feature vectors by generating an intermediate representation from the latent feature vectors and generating synthetic instances from the intermediate representation; and - training the generative model (430) includes backpropagating the second loss of the synthetic instances generated from the latent feature vectors by: - computing (432) the gradient of the loss with respect to the intermediate representation, the gradient including the partial derivatives of the loss with respect to the entries of the intermediate representation; - using the discriminative model to determine (434) a plurality of discriminator scores for the plurality of portions of the synthetic instance; - updating (436) the partial derivatives of the loss with respect to the entries of the intermediate representation based on the discriminator scores for the portions of the synthetic instance, where the portions of the synthetic instance are generated at least in part based on the entries of the intermediate representation, and where, if the discriminator score indicates an actual instance, the value of the partial derivative is decreased; and - further backpropagating (438) the loss based on the updated gradient, The generative model is configured to generate synthetic images, and the discriminative model is configured to determine discriminator scores for the respective portions of an input image.

13. The method (400) according to claim 12, further comprising using the generative model to generate one or more synthetic instances.

14. The method (400) according to claim 13, further comprising using the one or more synthetic instances as test and / or training data to train an additional machine learning model.

15. A computer-readable medium (500), comprising transient or non-transient data (510) representing: - instructions which, when executed by a processor system, cause the processor system to perform the computer-implemented method according to claim 12; and / or - parameters of a generative model and / or a discriminative model trained according to the computer-implemented method according to claim 12.

Citation Information

Patent Citations

  • A method and an apparatus for general image classification base on a semi-supervised generative adversarial network

    CN109190665A

  • Image instance segmentation method and device

    CN109635812A