Large-Scale Generative Neural Network Models with Inference for Representation Learning Using Adversarial Training

By combining the joint discriminator loss term and the single discriminator loss term in the loss function, the generator neural network and the encoder neural network are jointly trained, and the efficiency and stability problems of training large-scale neural networks in the existing technology are solved, achieving more efficient training and more realistic data generation.

CN113795851BActive Publication Date: 2025-05-06GDM HOLDING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080033423.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-23
Filing Date
2020-05-22
Publication Date
2025-05-06
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively train large-scale generator neural networks and encoder neural networks, especially in processing large-scale data and generating real data.

Method used

The generator neural network and the encoder neural network are used to train the generator neural network together. The loss function includes the joint discriminator loss term and the single discriminator loss term. The samples generated by the generator network are distinguished from the real samples through the discriminator neural network.

Benefits of technology

Effective training of large-scale generator and encoder neural networks is realized, the ability to generate real data and the stability of the training process is improved, and the demand for computing resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113795851B_ABST
    Figure CN113795851B_ABST
Patent Text Reader

Abstract

The present invention provides methods, systems and apparatus for training a generator neural network and an encoder neural network, including a computer program encoded on a computer storage medium. The generator neural network generates data items based on a set of potential values, which are samples of a distribution. The encoder neural network generates a set of potential values ​​for the corresponding data items. The method includes jointly training the generator neural network, the encoder neural network and the discriminator neural network, the discriminator neural network being configured to distinguish between samples generated by the generator network and samples of the distribution that are not generated by the generator network. The discriminator neural network is configured to distinguish input pairs including sample parts and potential parts by processing through the discriminator neural network. The training is based on a loss function, the loss function including a joint discriminator loss term based on the sample part and the potential part of the input pair processed by the discriminator neural network and at least one single discriminator loss term based on only one of the sample part or the potential part of the input pair.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a non-provisional application of and claims priority to U.S. Provisional Patent Application No. 62 / 852,250 filed on May 23, 2019. Background Art

[0002] The present specification relates to methods and systems for training large-scale generative neural networks and encoding neural networks for performing inference.

[0003] A neural network is a machine learning model that uses one or more layers of nonlinear units to predict outputs from received inputs. In addition to the output layer, some neural networks include one or more hidden layers. The output of each hidden layer is used as the input to the next layer in the network, i.e., the next hidden layer or output layer. Each layer of the network generates an output from the received inputs based on the current values ​​of the corresponding set of parameters. Summary of the invention

[0004] This specification generally describes how a system implemented as a computer program in one or more computers at one or more locations can perform a method to train (i.e., adjust parameters of) an adaptive system that is a generative adversarial network (GAN) that includes an inference model, which includes a generator neural network, an encoder neural network, and a discriminator neural network. The neural networks are trained based on a training set of data items selected from a distribution. The generator neural network, once trained, can be used to generate samples from the distribution based on potential values ​​(or just "potential") selected from a potential value distribution (or "latent distribution"). The encoder neural network, once trained, can be used to generate potential values ​​from the potential value distribution based on data items selected from the distribution. That is, the encoder neural network can be considered to implement the inverse function of the generator neural network.

[0005] More specifically, the present specification relates to a computer-implemented method for training a generator neural network and an encoder neural network. The generator neural network may be configured to generate data items based on a set of potential values, which are samples of a distribution representing a set of training data items. The encoder neural network may be configured to generate a set of potential values ​​for the corresponding data items. The training method may include jointly training the generator neural network, the encoder neural network, and a discriminator neural network, the discriminator neural network being configured to distinguish between samples generated by the generator network and samples of the distribution that are not generated by the generator network. The discriminator neural network may be configured to distinguish, by processing, input pairs including sample portions and potential portions by the discriminator neural network. The sample portions and potential portions of the input pairs may include samples of the distribution generated by the generator neural network and training data items for generating corresponding sets of potential values ​​or sets of training data items for the samples, respectively, and sets of potential values ​​generated by the encoder neural network based on the training data items. The training may be based on a loss function that includes a joint discriminator loss term based on the sample portions and potential portions of the input pairs processed by the discriminator neural network and at least one single discriminator loss term based only on one of the sample portions or potential portions of the input pairs.

[0006] In an embodiment, it has been found that large-scale generator neural networks and encoder neural networks can be effectively trained by training the generator neural network and the encoder neural network using a loss function, the loss function including a joint discriminator loss term based on the input pairs processed by the discriminator neural network and a single discriminator loss term based on only one of the sample portions or potential portions of the input pairs. This may allow more efficient processing of large-scale data compared to known methods. In particular, experiments have found that the encoder's examples are better than known methods in extracting salient information from data items, such as for use by a classification system or for other purposes, such as controlling an agent. It has been found that the generator is better than known techniques in generating data items that users believe to be real. It has also been found that the use of a loss function provides a more stable training process than some known techniques, which allows large-scale generation and inference neural networks to be effectively trained, the loss function including a joint discriminator loss term based on the input pairs processed by the discriminator neural network and a single discriminator loss term based on only one of the sample portions or potential portions of the input pairs.

[0007] The training method may also include the following optional features.

[0008] The single discriminator loss term may be based on the sample portion of the input pair. The single discriminator loss term may include a sample discrimination score generated based on processing the sample portion of the input pair using the sample discriminator subnetwork. The sample discrimination score may indicate the likelihood that the sample portion of the input pair is a sample generated by the generator neural network or a true training data item of the set of training data items. In this regard, the sample discrimination score may be a probability.

[0009] The sample discriminator subnetwork can be based on a convolutional neural network. For example, the sample discriminator subnetwork can be based on a discriminator network from the "BigGAN" framework ("Large scale GAN training for high fidelity natural image synthesis" by Andrew Brock, Jeff Donahue, and Karen Simonyan, submitted in arXiv 1809:11096 at ICLR 2019, the disclosure of which is incorporated herein by reference).

[0010] The sample identification score may be further generated based on applying the projection to the sample discriminator sub-network. For example, the projection may be implemented as a further linear neural network layer, which may have trainable parameters to be trained using the described training method.

[0011] The single discriminator loss term may be based on a latent portion of the input pair. The single discriminator loss term includes a latent discriminator score generated based on processing the latent portion of the input pair using the latent discriminator subnetwork. The latent discriminator score may indicate a likelihood that the latent portion of the input pair is a set of latent values ​​generated based on the training data item using the encoding neural network or a set of latent values ​​corresponding to a sample generated using the generator neural network. In this regard, the latent discriminator score may be a probability.

[0012] The latent discriminator score may be further generated based on applying the projection to the latent discriminator sub-network. For example, the projection may be implemented as a further linear neural network layer, which may have trainable parameters to be trained using the described training method.

[0013] The latent discriminator subnetwork may be based on a multi-layer perceptron. For example, the latent discriminator subnetwork may be a "ResNet" type neural network including residual blocks and skip connections.

[0014] The loss function may include multiple single discriminator loss terms. For example, the loss function may include a sample discriminator score and a latent discriminator score.

[0015] The joint discriminator loss term may include a joint discriminator score generated using the joint discriminator subnetwork. The joint discriminator score may indicate the likelihood that an input pair includes a sample from a distribution generated by a generator neural network and a training data item of a corresponding set of potential values ​​or a set of training data items used to generate the sample, respectively, and a set of potential values ​​generated by the encoder neural network based on the training data items.

[0016] The joint discriminator subnetwork may be configured to process an input pair. Alternatively, the joint discriminator subnetwork may be configured to process an output of a sample discriminator subnetwork and an output of a potential discriminator subnetwork, wherein the sample discriminator subnetwork is configured to process a sample portion of an input pair and the potential discriminator subnetwork is configured to process a potential portion of an input pair. The sample discriminator subnetwork and the potential discriminator subnetwork may be the same as described above.

[0017] The joint discriminator score may be further generated based on applying the projection to the joint discriminator sub-network. For example, the projection may be implemented as a further linear neural network layer, which may have trainable parameters to be trained using the described training method.

[0018] The joint discriminator sub-network may be based on a multi-layer perceptron. For example, the joint discriminator sub-network may be a "ResNet" type neural network including residual blocks and skip connections.

[0019] The loss function may be based on the sum of the joint discriminator loss term and the single discriminator loss term. It will be appreciated that where there are multiple single discriminator loss terms, the sum may include all single discriminator loss terms or a subset of the single discriminator loss terms.

[0020] The loss function may include a hinge function applied to a component of the loss function (e.g., a fraction or the negative of a fraction). The hinge function may be defined as h(t) = max(0,1-t). The hinge function may be applied to a component of the loss function individually or to the sum of components of the loss function or to any combination of individual and aggregated loss function components.

[0021] The encoder neural network can represent a probability distribution, and generating a potential value set can include sampling from the probability distribution. In this way, the encoder neural network is non-deterministic. The output of the encoder neural network can include a mean and a standard deviation for defining a normal probability distribution, and the potential value set can be sampled from the normal probability distribution. The encoder neural network can include a final neural network layer that implements a non-negative "soft-plus" nonlinearity for generating a standard deviation. The "soft-plus" nonlinearity can be defined as log(1+exp(x)). The potential value set can be generated based on the reparameterized sampling. For example, a potential value z can be generated as z=mean+ε*standard deviation, where ε is sampled from a unit Gaussian with zero mean. Alternatively, the potential value can be based on a discrete probability distribution.

[0022] The encoder neural network may be based on a convolutional neural network. For example, the encoder neural network may be a "ResNet" or "RevNet" type neural network having standard residual blocks or reversible residual blocks having further fully connected layers having skip connections.

[0023] The generator neural network can be a large-scale deep neural network, and for example, can be based on the "BigGAN" framework. The generator neural network can generate samples unconditionally or conditionally.

[0024] The training may further include alternating updates of the discriminator neural network parameters and updates of the encoder neural network parameters and the generator neural network parameters, wherein the updates are generated based on the loss function.

[0025] In general, training follows the GAN framework. Training is an iterative process where each iteration is based on a mini-batch of samples that are used to determine the value of the loss function from which the neural network parameter updates are determined using gradient descent and backpropagation.

[0026] Training can further include jointly updating the encoder neural network parameters and the generator neural network parameters.

[0027] Alternating the updating of the discriminator neural network parameters and the updating of the encoder neural network parameters and the generator neural network parameters may include performing multiple updates of the discriminator neural network parameters followed by updates of the encoder neural network parameters and the generator neural network parameters. For example, training may include two updates of the discriminator neural network parameters followed by a joint update of the encoder and generator neural network parameters.

[0028] In some implementations, the latent values ​​can include categorical variables, for example, by concatenating the latent values ​​with categorical variables, for example, 1024-way categorical variables. In this way, the generator neural network can learn clusters of data items, and the encoder neural network can learn to classify the data items (making predictions in the embedding space rather than the latent variable space itself).

[0029] A method for performing inference using an encoder neural network is also described, the method comprising: processing an input data item using the encoder neural network to generate a potential value set representing the input data item. The encoder neural network is jointly trained with a generator neural network and a discriminator neural network, the generator neural network is configured to generate data items based on the potential value set, the data items are samples of a distribution representing a set of training data items, and the discriminator neural network is configured as samples generated by the generator network and samples of a distribution not generated by the generator network. The discriminator neural network is configured to distinguish input pairs including sample portions and potential portions by the discriminator neural network through processing. The sample portion and the potential portion of the input pair include samples of a distribution generated by the generator neural network and training data items for generating the corresponding potential value set or training data item set of the sample, respectively, and the potential value set generated by the encoder neural network based on the training data item. The training is based on a loss function, the loss function including a joint discriminator loss term based on the sample portion and the potential portion of the input pair processed by the discriminator neural network and a single discriminator loss term based on only one of the sample portion or the potential portion of the input pair. The training can be performed according to the above program.

[0030] Generally speaking, inference is the process of determining potential values ​​that describe specific input data items.

[0031] The method may further include classifying the input data item based on a potential value representing the input data item. For example, a classifier may be trained using the output of the encoder neural network as a representation of the input data item to perform classification. The classifier may be a linear classifier or other type of classifier.

[0032] The method may also include performing an action using the agent based on representing the potential value of the input data item.

[0033] The potential values ​​generated from the data items by the trained encoder neural network can be used to classify the data, for example as part of an image or audio signal processing classification system. The potential values ​​generated from the data items provide a representation of the data items. Multiple data items can be processed by the trained encoder neural network to generate potential values, and, for example, can be used to train a classifier. For example, the data items used to train the classifier can be labeled with corresponding labels from a plurality of labels, and the classifier can be trained to classify the potential value representations generated from unclassified data items having one of the plurality of labels.

[0034] The potential values ​​generated by the trained encoder neural network can additionally or alternatively be used to search for data items similar to the provided query data item. For example, for each data item in a plurality of stored data items, a potential value can be generated and stored. Subsequently, the query data item can be used to query the stored data items by generating a potential value for the query data item and searching for data items based on the potential values ​​stored for the stored data items and the potential value for the query data item. For example, the search can be based on a similarity measure, or the search can use a neural network trained based on the stored data items.

[0035] The classification system may provide a classification of any suitable data items and, for example, may be used to classify / find / recommend audio clips, images, videos, or products, for example, based on an input sequence that may represent one or more query images, videos, or products.

[0036] The data items may be data representing a still or moving image (i.e., a sequence of images), in which case the individual numerical values ​​contained in the data items may represent pixel values, such as the values ​​of one or more color channels of a pixel. The training images used to train the neural network may be real-world images captured by a camera.

[0037] Alternatively, the data items may be data representing sound signals, such as amplitude values ​​of an audio waveform (e.g., natural language; in this case, the training examples may be samples of natural language, such as speech from a human speaker recorded by a microphone). In another possibility, the data items may be text data, such as text strings or other representations of words and / or sub-word units in a machine translation task. Thus, the data items may be one-dimensional, two-dimensional, or higher-dimensional.

[0038] Alternatively, the latent values ​​may define text strings or spoken sentences or their encodings, and the generator neural network may generate images corresponding to the text or speech (text-to-image synthesis). In principle, the opposite situation is possible, where the data items are images and the latent values ​​represent text strings or spoken sentences. Alternatively, the latent values ​​may define text strings or spoken sentences or their encodings, and the generator network may then generate corresponding text strings or spoken sentences in different languages.

[0039] In particular, where the data item is a data sequence (e.g. a video sequence), the generator may also generate the data item (e.g. a video) autoregressively, in particular given one or more previous video frames.

[0040] The generator network can generate sound data, such as speech, in a similar manner. This may be conditioned on the audio data and / or other data, such as text data. In general, the target data can define local and / or global features of the generated data items. For example, for audio data, the generator neural network can generate an output sequence based on a series of target data values. For example, the target data may include global features (also when the generator network is used to generate a sequence of data items), which may include information defining the voice or speech style of a particular person's voice or the identity of the speaker or language. The target data may additionally or alternatively include local features (i.e., different for a sequence of data items), which may include language features derived from the input text, optionally with intonation data.

[0041] In another example, the target data may define the motion or state of a physical object, such as the action and / or state of a robotic arm. The generator neural network may then be used to generate data items that predict future images or video sequences seen by a real or virtual camera associated with the physical object. In these examples, the target data may include one or more previous images or video frames seen by the camera. Such data may be useful for reinforcement learning, for example to facilitate planning in a visual environment. More generally, the system learns to encode probability densities (i.e., distributions) that can be used directly for probabilistic planning / exploration.

[0042] By employing target data defining noisy or incomplete images, the generator neural network can be employed for image processing tasks such as denoising, deblurring, image completion, etc. The encoder neural network can be employed for image compression. The system can similarly be used to process signals other than those representing images.

[0043] In general, the input target data and the output data items can be any kind of digital data. Therefore, in another example, the input target data and the output data items can each include tokens that define sentences in a natural language. For example, the generator neural network is then used in the system for machine translation or generation of sentences that represent the concepts represented in the potential values ​​and / or additional data. The potential values ​​can be used additionally or alternatively to control the style or emotion of the generated text. In a further example, the input and output data items can typically include speech, video, or time series.

[0044] The generator neural network can be used to generate further examples of data items for training another machine learning system. The generator network can be used to generate new data items that are similar to data items in the training data set. The set of potential values ​​can be determined by sampling from a potential distribution of potential values. If the generator network has been trained conditioned on additional data (e.g., labels), new data items can be generated conditioned on the additional data (e.g., labels provided to the generator network). In this way, additional labeled data items can be generated, for example to supplement the lack of unlabeled training data items.

[0045] Neural networks can be configured to receive any kind of numeric data input and generate any kind of scoring, classification, or regression output based on the input.

[0046] For example, if the input to a neural network is an image or features that have been extracted from an image, the output generated by the neural network for a given image may be a score for each object category in a set of object categories, where each score represents an estimated likelihood that the image contains an object belonging to that category.

[0047] As another example, if the input to the neural network is an internet resource (e.g., a web page), a document or a portion of a document, or features extracted from an internet resource, a document or a portion of a document, the output generated by the neural network for a given internet resource, document or portion of a document may be a score for each topic in a set of topics, where each score represents an estimated likelihood that the internet resource, document or portion of a document is about that topic.

[0048] As another example, if the input to a neural network is features of the impression context of a particular ad, the output generated by the neural network may be a score representing an estimated likelihood that the particular ad will be clicked.

[0049] As another example, if the input to the neural network is features of personalized recommendations for a user, e.g., features characterizing the context of the recommendation, e.g., features characterizing previous actions taken by the user, the output generated by the neural network can be a score for each content item in a set of content items, where each score represents an estimated likelihood that the user will respond positively to the recommended content item.

[0050] As another example, if the input to a neural network is a sequence of text in one language, the output generated by the neural network may be a score for each text segment in a set of text segments in another language, where each score represents an estimated likelihood that the text segment in the other language is a correct translation of the input text into the other language.

[0051] As another example, if the input to the neural network is a sequence representing a spoken sentence, the output generated by the neural network may be a score for each text segment in a set of text segments, each score representing an estimated likelihood that the text segment is a correct transcription of the sentence.

[0052] The subject matter described in this specification can be implemented in specific embodiments to achieve one or more of the following advantages.

[0053] In general, generative adversarial networks are not able to perform reasoning without modification. The present disclosure trains an encoder neural network that is able to perform the reverse operation (i.e., reasoning) of a generator neural network, and thus, is able to generate a set of potential values ​​representing a data item. The set of potential values ​​can be used for other tasks, such as classification or under the control of an agent, such as in a reinforcement learning system.

[0054] Additionally, by jointly training the generator neural network and the encoder neural network in this way, the generated latent values ​​are more effective at capturing important changes in the data distribution rather than trivial changes, such as contrast differences between pixels. This enables improved accuracy on downstream tasks such as the classification task mentioned above.

[0055] The training method also enables the use of larger generator and encoder neural networks that can model more complex data distributions. Specifically, the loss function including a joint discriminator loss term and a single discriminator loss term enables large-scale generator and encoder neural networks to be effectively and efficiently trained. Given the improved overall training speed, the training method requires fewer computing resources - such as processor time and power usage - to complete the training.

[0056] The training method can also be performed using only unlabeled data and does not require the data to be corrupted in any way as in self-supervised methods.

[0057] The training method can also be performed using a distributed system, for example, the generator neural network, the encoder neural network, and the discriminator neural network can reside on different processing systems of the distributed system and be trained in parallel.

[0058] In embodiments, latent values ​​can provide representations of the data items on which the system is trained that tend to capture the high-level semantics of the data items rather than their low-level details, with training encouraging the encoder neural network to model the former rather than the latter. Thus, latent values ​​can naturally capture the "class" of a data item, despite being trained using unlabeled data. This greatly expands the number of potentially available training data items, and thus the detailed semantics that can be captured.

[0059] Once trained, the latent values ​​can be used by subsequent systems for subsequent tasks, which can be greatly simplified because useful semantic representations are already available. As described above, the subsequent system can be configured to perform almost any task, including but not limited to image processing or visual tasks, classification tasks, and reinforcement learning tasks. As previously described, some such subsequent systems typically require labeled training data items, and can be further trained using labeled training data to fine-tune the latent value representations that have been derived using unlabeled training data. Such systems can learn faster, use less memory and / or computing resources, and ultimately perform better by using the systems and methods described herein to determine the latent value representations on which they can work.

[0060] In some embodiments, the above-mentioned system / method may provide a potential value representation to a subsequent reinforcement learning system. For example, such a reinforcement learning system may be used to train an agent policy neural network through reinforcement learning to control the agent to perform a reinforcement learning task when interacting with the environment. For example, in response to an observation, the reinforcement learning system may select an action to be performed by the agent and cause the agent to perform the selected action. Once the agent has performed the selected action, the environment transitions to a new state, and the reinforcement learning system may receive a reward, typically a numerical value. The reward may indicate whether the agent has completed the task, or the progress of the agent in completing the task. For example, if the task specifies that the agent should navigate to a target location through the environment, the reward for each time step may have a negative value once the agent reaches the target location, otherwise it may have a zero value. As another example, if the task specifies that the agent should explore the environment, the reward for the time step may have a positive value when the agent navigates to a previously unexplored location at the time step, otherwise it may have a zero value.

[0061] In some such embodiments, the environment is a real environment and the agent is a mechanical agent that interacts with the real environment, such as a robot or an autonomous or semi-autonomous land, air, or sea vehicle that navigates through the environment.

[0062] In these embodiments, for example, observations may include one or more of: images, object position data, and sensor data captured as the agent interacts with the environment, such as sensor data from image, distance or position sensors or data from actuators.

[0063] For example, in the case of a robot, observations may include data characterizing the current state of the robot, such as one or more of: joint positions, joint velocities, joint forces, torques or accelerations - e.g., gravity compensated torque feedback - and the global or relative pose of an object held by the robot. In the case of a robot or other mechanical agent or vehicle, observations may likewise include one or more of: position, linear or angular velocity, force, torque or acceleration, and the global or relative pose of one or more parts of the agent. Observations may be defined in 1D, 2D, or 3D, and may be absolute and / or relative observations. For example, observations may also include sensed electronic signals, such as motor currents or temperature signals; and / or image or video data, such as from a camera or LIDAR sensor, such as data from a sensor of the agent or from a sensor positioned separately from the agent in the environment.

[0064] In these embodiments, the action may be a control input for controlling a robot, such as a torque for a joint of a robot or a high-level control command or an autonomous or semi-autonomous land, air, or sea vehicle, such as a torque for a control surface or other control element of a vehicle or a high-level control command. In other words, for example, an action may include position, velocity, or force / torque / acceleration data for one or more joints of a robot or a portion of another mechanical agent. The action data may additionally or alternatively include electronic control data, such as motor control data, or more generally data for controlling one or more electronic devices in an environment, the control of which may have an effect on the observed state of the environment. For example, in the case of an autonomous or semi-autonomous land, air, or sea vehicle, the action may include actions for controlling navigation, such as steering, and motion, such as braking and / or acceleration of the vehicle.

[0065] In some other embodiments, the environment is a simulated environment, and the agent is implemented as one or more computers that interact with the simulated environment. The simulated environment can be a motion simulation environment, such as a driving simulation or a flight simulation, and the agent can be a simulated vehicle that navigates through the motion simulation. In these embodiments, the action can be a control input for controlling a simulated user or a simulated vehicle. In this way, the robotic reinforcement learning system may undergo partial or complete simulation training before being used on a real-world robot. In another example, the simulated environment can be a video game, and the agent can be a simulated user playing the video game. Typically, in the case of a simulated environment, observations can include simulated versions of one or more of the aforementioned observations or observation types, and actions can include simulated versions of one or more of the aforementioned actions or action types.

[0066] In the case of an electronic agent, observations may include data from one or more sensors that monitor parts of a plant or service facility, such as current, voltage, power, temperature, and other sensors and / or electronic signals that represent the functionality of electronic and / or mechanical items of equipment. In some other applications, the agent may control actions in a real-world environment that includes items of equipment, such as in a data center, in a power / water distribution system, or in a manufacturing plant or service facility. The observations may then be related to the operation of the plant or facility. For example, observations may include observations of power or water used by the equipment or observations of power generation or distribution control or observations of resource use or waste generation. Actions may include actions that control or impose operating conditions on items of equipment in the plant / facility and / or cause changes in settings in the operation of the plant / facility, for example, to adjust or turn on / off actions of components of the plant / facility. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 shows an encoder and generator that can be trained;

[0068] Figure 2 An identifier is shown, which is used in an example of some principles of the present disclosure to Figure 1 The encoder and generator are trained;

[0069] Figure 3 Shows the Figure 1 The encoder and generator of Figure 2 The steps of the method for jointly training the discriminator of

[0070] Figure 4 The steps of a first method using a trained encoder are shown;

[0071] Figure 5 The steps of a second method using a trained encoder are shown; and

[0072] Figure 6 The steps of a method using a trained encoder are shown.

[0073] Like reference numbers and designations indicate like elements throughout the various drawings. DETAILED DESCRIPTION

[0074] Figure 1 Schematically illustrated are encoder neural networks 11 ("encoder") and generator neural networks 12 ("generator") that can be trained by the training method disclosed herein. The method employs a training database that includes many instances of data items (e.g., images or portions of sounds or text). Each data item is represented by a corresponding vector x (i.e., a collection of combinations of multiple data values). The data items have a distribution P x .

[0075] The encoder 11 is a neural network defined by a plurality of numerical network parameters that are adjusted during the training of the encoder 11. The encoder 11 performs a function defined at any time by the current values ​​of the set of network parameters. To generate the output The output is a set of potential values, that is, multiple potential values. The number of elements in x is usually the same for all data items and can be represented by N, which is an integer greater than 1. Similarly, The number of potential values ​​in is typically the same for all sets of potential values ​​and can be represented by n. N can be larger (e.g., at least 10 times larger, typically at least 1000 times larger) than The number of potential values ​​n in . For example, each data item is an image consisting of one or more values ​​for each pixel in at least 10,000 pixels, and The number of potential values ​​n in may be less than 1000 or less than 200. Let γ denote the network parameters of encoder 11, which may also be represented as

[0076] Generator 12 is defined by a number of numerical network parameters that are adjusted during training of generator 12. Generator 12 receives as input a set of potential values ​​("potentials"). The set of potential values ​​includes a plurality of potential values ​​and is represented by a vector z. Typically, the number of potential values ​​in z is equal to n, where n is an integer greater than 1. The potential set is drawn from a distribution P z For example, the distribution may be the same for each component of z and may be a simple continuous distribution such as an isotropic Gaussian N(0,I). The generator 12 performs a function defined at any time by the current value of the set of network parameters (given by , to generate data items Denote the network parameters of the generator 11 by Ξ, and the encoder 11 can also be represented by express.

[0077] Given the potential prior P z The generator 12 models the conditional distribution P(x|z) of the data item x. Given a latent input z sampled from the data distribution P x Given a data item x sampled from , the encoder 11 models the inverse conditional distribution P(z|x) to predict the potential z. In an embodiment of the present disclosure, the encoder 11 and the generator 12 are implemented using the generator and discriminator architecture from the "BigGAN" framework.

[0078] For example, in the BiGAN (bidirectional GAN) framework (Jeff Donahue, Philipp and Trevor Darrell, “Adversarial feature learning” submitted in arXiv:1605.09782 at ICLR 2017, and Vincent Dumoulin, Ishmael Belghazi, Ben Poole, Olivier Mastropietro, Alex Lamb, Martin Arjovsky, and Aaron Courville, “Adversarially Learned Inference” submitted in arXiv:1606.00704 at ICLR 2017, the disclosures of which are incorporated herein by reference), according to this example of the present disclosure, the encoder 11 and the generator 12 are trained using a discriminator neural network (“discriminator”) that takes as input data item-potential pairs (also referred to herein as “input pairs” of the discriminator). The input pairs received by the discriminator at different times may be (e.g., alternating) representations of the input x to the encoder and the corresponding output Data items - potential pairs or represents the input z to generator 12 and the corresponding output Data items - potential pairs Therefore, the "sample part" of each pair is either x or And the "potential part" of each pair is or z. The discriminator learns to distinguish between x and encoder 11 with the potential distribution P from generator 12 z Yes Specifically, the input to the discriminator is and The encoder 11 and the generator 12 are trained to obtain the two joint distributions P x,ε(x) and In order to “fool” the discriminator, the two types of data item-latent pairs are sampled from these two joint distributions respectively.

[0079] Training is performed using a training data database of training data items x, such that P x represents the distribution x over the training data items in the training data database (e.g., where the data items are images, a database of example images; the example images may be real-world images captured with one or more cameras, or alternatively synthetic images generated by a computer). That is, the generator 12 is configured to generate data items based on the set of potential values ​​z These data items are the distribution P of the database representing the training data items x of the sample.

[0080] Possible discriminators 21 proposed in this disclosure are as follows Figure 2 As shown in Figure 2, and different from the discriminator used in BiGAN, it receives a continuous data item-potential pair, where each pair is from the data distribution P x Yes or from generator 12 and potential distribution P z Yes For example, the discriminator may receive each type of data item-potential pair alternately; or successively receive a batch of multiple data item-potential pairs of one type, and then successively receive a batch of multiple data item-potential pairs of another type.

[0081] The discriminator 21 comprises a sample discriminator subnetwork 211 which receives (only) samples generated by the generator network 12, depending on which type of data item-potential pair is the input to the discriminator 21. or training data item x. That is, the sample discriminator network 211 does not receive a set of potential values ​​z and or based on their data. The sample discriminator subnetwork 211 performs a function defined at any time by the current value of the set of trainable numerical network parameters Ψ (given by F Ψ Projection θ x (can be considered as having x A linear neural network layer with network parameters given by ) is applied to the output of the sample discriminator subnetwork 211 to produce x Denoted as a unary (one component) sample discrimination score (for both types of data item-potential pairs). In particular, in the case where the data items are images, the sample discriminator subnetwork 211 can be implemented as a neural network with one or more input layers being a convolutional neural network.

[0082] The discriminator 21 comprises a latent discriminator subnetwork 212 which receives (only) the latent information used by the generator 12 to generate the sample, depending on which type of data item-potential pair is the input to the discriminator 21. The potential value set z or the potential value set generated by the encoder 11 based on the training data item x That is, the potential discriminator sub-network 212 does not receive samples and training data items x and data based on them. The latent discriminator subnetwork 212 performs a function defined at any time by the current value of the set of trainable numerical network parameters Φ (given by H Φ Projection θ z (a linear neural network layer) is applied to the output of the latent discriminator subnetwork 212 to produce z Denote the unary latent discriminant score (for both types of data item-latent pairs) . The latent discriminator subnetwork 212 may optionally be implemented as a perceptron, such as a multi-layer perceptron.

[0083] The discriminator 21 further includes a joint discriminator subnetwork 213 that receives the outputs of the sample discriminator subnetwork 211 and the latent discriminator subnetwork 212. The joint discriminator subnetwork 213 performs a function defined at any time by the current value of the set of trainable numerical network parameters Θ (denoted by J Θ Projection θ xz (Linear Neural Network Layer) is applied to the output of the joint discriminator subnetwork 212 to produce xz Denote the unary joint discriminant score (for both types of data items - potential pairs) . The joint discriminator subnetwork 213 may optionally be implemented as a perceptron, such as a multi-layer perceptron.

[0084] Three Score S xz , S x and S z The summation unit 22 sums the values ​​to produce a loss value (or “loss”) denoted by l. Thus, the loss l is calculated using the following terms: (i): the unary sample score S x , the one-element sample score S x Based only on the sample part of the input pair, and the unary sample score S of the adaptive network 211, 212, 213 of the discriminator 21 x Depends only on the output of the sample discriminator network 211; (ii) the unary potential score S z , the unary potential score S z Based only on the latent part of the input pair, and the unary latent score S of the adaptive network 211, 212, 213 of the discriminator 21 zdepends only on the output of the latent discriminator network 212; and (iii) the joint score S xz , the joint score S xz The data distribution and the latent distribution are linked and are based on the sample part and the latent part of the input pairs and on the outputs of all three adaptive networks 211 , 212 , 213 .

[0085] The loss function used to perform training can be generated from the loss value generated by the summing unit 22. When calculating the loss value, the summing unit 22 can apply the summation result to different symbols according to the type of the input pair (i.e., whether it is a sample of the distribution generated by the generator neural network and the corresponding potential value set used to generate the sample, or alternatively, whether it is a training data item of the training data item set and the potential value set generated by the encoder neural network based on the training data item). Optionally, in some cases, the loss value can be modified using a hinge function ("hinge") before or after the summation.

[0086] Specifically, based on the scalar discriminator score function and the corresponding per-sample loss and Discriminator loss function (“Discriminator loss”) and the encoder-generator loss function The following definition:

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094] where y is either -1 or +1 (i.e., y∈{-1,+1}) and h(t)=max(0,1-t) is a "hinge" that is optionally used to regularize the discriminator. D In the case of , the hinge can be viewed as modifying the operation of summation unit 22. Therefore, the loss function and There are two single discriminator loss terms, which are based on the unary sample score S x and S z(each based on only one of the sample or potential parts of the input pair); and a joint discriminator loss term based on the joint score S xz The expected value of (based on the sample and latent parts of the input pair processed by the discriminator neural network). In the case of the discriminator loss function, the calculation of the expected value takes into account the hinge.

[0095] ε and The corresponding network parameter sets γ and Ξ are optimized to minimize the loss function And the projection θ x ,θ z and θ xz and the corresponding network parameter sets Ψ, Φ and Θ of F, H and J are optimized to minimize the loss function expect is estimated via Monte Carlo sampling of small batches.

[0096] Note that you can replace l D Alternative discriminator loss used According to the three loss terms (i.e., The sum of only calls the “hinge” h once. However, it is experimentally found that the above definition of l clamping each of the three loss terms separately D This approach doesn't perform very well in comparison.

[0097] We experimentally find that the discriminator 21 leads to better representation learning results (e.g. compared to BiGAN) without affecting generation. This is achieved by explicitly enforcing the property that the marginal distributions of x and z match at the global optimum (i.e., the distribution generated by the encoder 11 matches the distribution of P z matches, and the distribution of data items generated by generator 12 matches P x Match), unary single discriminator score S x and S z Guide the optimization in the “right direction”. For example, in the context of image generation, the unary loss term on x matches the loss used by the GAN algorithm and provides a learning signal that only steers the generator 12 towards a distribution consistent with the image distribution P x matches independently of its underlying input.

[0098] According to the present disclosure, several variations of the discriminator 21 are possible. For example, depending on which type of data item-potential pair is the input to the discriminator 21, the joint discriminator sub-network 213 can directly (i.e., substantially without prior modification, e.g., by an adaptive component) receive the samples generated by the generator network 12. and is used by generator 12 to generate samples The potential value set z or training data item x and the potential value set generated by the encoder 11 based on the training data item x Thus, the joint discriminator sub-network 213 receives data item-potential pairs directly, rather than via the single discriminator sub-networks 211 , 212 .

[0099] In a further variation, based on a single discriminator score S x and S z The loss term of any one of the above can be omitted from the loss function. Optionally, in this case, the corresponding one of the sample discriminator subnetwork 211 or the potential discriminator subnetwork 212 can be omitted from the discriminator 21. In this case, the joint discriminator subnetwork 213 can directly (i.e., substantially without prior modification, e.g., by an adaptive component) receive (i) the sample and training data items x or (ii) sets of potential values ​​z and The corresponding one in .

[0100] Steering Figure 3 , a method 300 for jointly training the encoder 11, the generator 12, and the discriminator 21 is shown. The method can be performed by one or more computers in one or more locations. The method 300 includes repeatedly executing a set of steps 301 to 305 of the method until a termination criterion is reached (e.g., the number of iterations reaches a predetermined value). Each execution of the set of steps 301 to 305 is called an iteration. Steps 304 and 305 can be performed by an update unit (e.g., a general-purpose computer) configured to modify the encoder 11, the generator 12, and the discriminator 21. Initially, the network parameters of the encoder 11, the generator 12, and the discriminator 21 are set to corresponding initial values, for example, these initial values ​​may all be the same.

[0101] In step 301, the generator 12 obtains a batch of potential values ​​and generates corresponding sample data items based on its current network parameters.

[0102] In step 302, the encoder 11 obtains a small batch of training data items and generates a corresponding set of potential values ​​for each one based on its current network parameters.

[0103] In step 303, the data item-potential pairs generated in step 301 and the data item-potential pairs generated in step 302 are successively input to the discriminator 21, and the summing unit 22 generates a corresponding loss for each data item-potential pair.

[0104] In step 304, the updating unit estimates based on the loss output by the summing unit 22 in step 303 and The current value of .

[0105] In step 305, the updating unit updates at least one network parameter of one or more of the encoder 11, the generator 12, and the discriminator 21 according to conventional machine learning techniques such as back propagation. The update of the network parameters γ and Ξ is to reduce the loss function And the projection θ x ,θ z and θ xz The update of the network parameters Ψ, Φ and Θ of F, H and J (i.e., the network parameters of the corresponding linear network layer) is to reduce the loss function Optionally, in alternating iterations of the set of steps 301 to 305, step 305 may be used as (i) a pair of ε and The network parameters γ and Ξ are jointly updated or (ii) the projection θ x ,θ z and θ xz and the network parameters Ψ, Φ and Θ of F, H and J are updated. Note that in the variant, the projection θ x ,θ z and θ xz It may be selected (eg, randomly) prior to execution of method 300 and not changed during the process.

[0106] For example, for each iteration except the first, by and The current value of is compared with the value obtained in the last iteration to estimate and The gradient of the loss function with respect to the corresponding network parameters (i.e., Relative to ε and The gradient of the network parameters γ and Ξ and the loss function Relative to the projection θ x ,θ z and θ xz The updating may be performed by calculating the gradients of the network parameters Ψ, Φ and Θ of F, H and J, and computing the updated value of each network parameter based on the corresponding gradient relative to the parameter. In the first iteration, the updated value may be obtained randomly. Alternatively, steps 301 to 304 may be performed in each iteration, not only for the current values ​​of the network parameters of the encoder 11, the generator 12 and the discriminator 21, but also for one or more supplementary sets of these network parameters, which are respectively shifted by a certain amount from the corresponding current values ​​of the network parameters, and in step 305, and The gradient with respect to the corresponding network parameter can be obtained by dividing the value of the current network parameter by and The estimated parameters are compared with a complementary set of these network parameters.

[0107] Steering Figure 4 , a first method 400 using a trained encoder 11 is shown. The method can be performed by one or more computers in one or more locations. In step 401, the encoder receives a data item (e.g., including an image (such as a real-world image captured by a camera) and / or an image sequence (such as a video sequence captured by a video camera and / or a sentence captured with a microphone)). In step 402, the encoder 11 processes the input data item to generate a corresponding potential value set representing the data item.

[0108] In step 403, encoder 11 classifies the data item based on the potential values. Classification step 403 can be performed in a variety of ways, for example, by determining in which of a plurality of predetermined regions in the space of potential values ​​the set of potential values ​​obtained in step 402 falls, and classifying the data item into a class associated with the determined region.

[0109] The predetermined region sets may be obtained in a variety of ways. For example, the predetermined regions may have been obtained by performing steps equivalent to steps 401 to 402 on a plurality of images and subjecting the resulting plurality of potential value sets to an automatic clustering algorithm, thereby identifying a plurality of clusters in the space of potential value sets, wherein each potential value set in the plurality of potential value sets is associated with a cluster, and defining each predetermined region based on the potential value set associated with the corresponding cluster.

[0110] Alternatively, the set of predetermined regions in the space of latent values ​​associated with the corresponding predetermined classes may be obtained using a plurality of labeled data items (i.e., each data item is associated with a set of label data indicating one or more of the plurality of predetermined classes into which the data item falls). The encoder 11 performs steps equivalent to steps 401 to 402 for each of the labeled data items, and defines the predetermined region for each of the classes as a region in the space of latent values ​​that contains data items labeled with label data indicating the class. This may be accomplished by training a classification unit (e.g., a neural network classification unit) such as a linear classifier based on the output of the trained encoder 11 for each of the labeled data items and the corresponding label data. Thus, although in method 300 the encoder 11 is trained with unlabeled data items (or in any case, labels are not typically used to train the encoder 11), the labeled data items may subsequently be used to train the classifier unit. This makes it possible to obtain high-quality classification even if the number of available labeled data items is small, for example because they are obtained by manually labeling a portion of a large database of data items. This is because the number of labeled data items required to train the classifier unit is small, assuming that the classifier unit includes far fewer network parameters than the encoder 11.

[0111] Note that regions in the space of potential values ​​associated with corresponding classes may overlap, for example, if the classes are not mutually exclusive. For example, if the classes are "winter", "summer", "mountain", "valley", and if the data item is an image of a mountain or a valley captured in winter or summer, the label data of each data item may indicate a corresponding class in the class "winter" or "summer" and a corresponding class in the class "mountain" or "valley", and the predetermined regions corresponding to the classes "winter" and "summer" may overlap with the predetermined regions corresponding to the classes "mountain" and "valley", respectively, while they may not overlap with each other.

[0112] Experiments associated with method 300 were performed using data items that were 128×128 images and a potential value set for each data item consisting of 120 elements. The structure of the generator 12 and the sample discriminator subnetwork 211 is used in “BigGAN”. The latent discriminator subnetwork 212 and the joint discriminator subnetwork 213 are 8-layer MLPs with ResNet-style skip connections (four residual blocks, each with two layers) and hidden layers of size 2048. The architecture of the encoder 11 is a ResNet-v2-50 ConvNet, followed by a 4-layer MLP (where the size of each layer is 4096) with skip connections (two residual blocks) after the global average pooling output of the ResNet.

[0113] The encoder 11 is configured to generate an output having a distribution N(μ,σ). The encoder 11 comprises a linear output layer which determines the mean μ and is further quantized by a non-negative “soft-plus” nonlinearity Determine the value associated with σ The potential value set z can be generated based on the reparameterized sampling. Specifically, each component of z is generated as z=μ+epsilon*σ, where epsilon is sampled from a unit Gaussian with zero mean. Alternatively, the potential value can be based on a discrete probability distribution.

[0114] Experimentally, it is found that omitting the sample discriminator subnetwork 211 degrades performance more than omitting the latent discriminator subnetwork 212, but the discriminator 21 performs best if it includes both discriminator subnetworks 211, 212. The sample discriminator subnetwork 211 has a large positive impact on the performance of the trained generator 12 and slightly leads to a slight improvement in the performance of the linear classifier. The presence of the latent discriminator subnetwork 212 leads to an improvement in the performance of the linear classifier, which depends on the distribution P z .

[0115] It is found that efficient use of computing resources is achieved when the training data items x in the training database received by the encoder 11 have a higher number of training data items x than the data items generated by the generator 12. In this case, when the type When the data item-potential pair is provided to the discriminator 21, the training data item x is first downsampled to match the data item generated by the generator 12. In this way, with high computational efficiency, it is possible to generate an encoder 11 that allows accurate classification of high-resolution images without requiring the generator 12 and the discriminator 21 to be able to generate / receive high-resolution images, thereby reducing the computational time required to train the encoder 11, the generator 12, and the discriminator 21. Note that in the experiments, the number of potential values ​​generated by the encoder 11 is the same as the number of potential values ​​input by the generator 12 and the discriminator 21.

[0116] Although, as described above, the training data items are typically not labeled, in a further variation of the training method (which may be used in the case where the training data items are associated with corresponding label data), the set of potential values ​​generated by the encoder upon receiving a given data item may include one or more components that label the data item as belonging to one or more categories in a predetermined set of categories. For example, in the case where the data item is a data item-potential pair that is a labeled training data item, the potential values ​​generated by the encoder from the labeled training data item may be a plurality of (e.g., normally distributed) latent variables that are connected to a categorical variable of the data item. The encoder is trained, for example, by adding an additional term to the loss function such that the categorical variable is equal to a value derived from the label data of the training data item. For example, the categorical variable may be a variable that labels the data item as belonging to one of a plurality of predetermined categories (e.g., 1024 categories). In the case where the data item is a data item-potential pair generated by the generator 12, the potential values ​​received by the generator 12 include the above-mentioned kind of potential values ​​(e.g., from a normal distribution) connected to the value of the categorical variable. In this way, the generator neural network can learn to cluster data items and the encoder neural network can learn to classify data items (i.e., the classification is performed by encoder 11, rather than using a classifier based on the output of encoder 11).

[0117] Steering Figure 5 , a second method 500 using a trained encoder 11 is shown. The method can be performed by one or more computers in one or more locations. In step 501, the encoder receives a data item (e.g., including an image (such as a real-world image captured by a camera) and / or an image sequence (such as a video sequence captured by a video camera and / or a sentence captured with a microphone)). In step 502, the encoder 11 processes the input data item to generate a corresponding potential value set representing the data item.

[0118] In step 503, the control unit uses the potential value set to generate control data for controlling the agent, and transmits the control data to the agent so that the agent performs an action. The agent can be any type of agent described above, for example, a robot. The agent interacts with the environment, such as by moving in the environment (the term "movement" is used to include translation of the agent from one position in the environment to another and / or reconfiguration of the agent's components, but does not necessarily include translation of the agent) and / or by moving objects in the environment. The environment can be a real-world environment, and the data items can be data collected by one or more sensors (e.g., a microphone or camera, such as a video camera) and describing the environment.

[0119] The process of generating control data based on a set of potential values ​​can be based on a policy. The policy can be trained by any known technique from the field of reinforcement learning, such as based on a reward value that indicates how much the agent's actions contribute to performing a specific task.

[0120] Steering Figure 6 , a method 600 using a trained encoder 12 is shown. The method can be performed by one or more computers in one or more locations. In step 601, the generator 12 receives a set of potential values. In step 602, the generator 12 processes the set of potential values ​​to generate corresponding data items.

[0121] Suppose the user wishes to compress the data item. This can be done by using an encoder to generate a set of potential values ​​that have significantly fewer elements than the data item (e.g., at least 100 times smaller or at least 1000 times smaller). Then, Figure 6 The method can use the potential value set to regenerate the data item. Although the regenerated data item will be different from the original data item received by the encoder 11, it can contain the same salient information (for example, if the original data item is an image of a dog, the regenerated image is also an image of a dog) and may be difficult to distinguish from the data items obtained from the training database. Experiments have confirmed this effect.

[0122] Assume, in another example, that the user wishes to obtain a data item that combines the features of two (or more) existing data items. In a preliminary step, the user may process the two data items with the encoder 11 to generate corresponding potential value sets, then form a new potential value set based on the two potential value sets (e.g., each potential value in the new set may be the average of the corresponding potential values ​​in the two potential value sets), and use the new potential value set in step 601 of the method. The generator 12 outputs a data item that has the features of the two initial data items and is difficult to distinguish from the data items obtained from the training database.

[0123] In another example, suppose the space of potential values ​​has been defined as above for Figure 4 The discussed approach is divided into regions, where different regions correspond to respective classes. A user wishing to generate data items belonging to one or more classes may select a set of potential values ​​within the respective regions and use it as the set of potential values ​​in step 601 of the method.

[0124] For a system of one or more computers to be configured to perform a specific operation or action, it means that software, firmware, hardware, or a combination thereof that causes the system to perform the operation or action when run has been installed on the system. For one or more computer programs to be configured to perform a specific operation or action, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.

[0125] The embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. However, a computer storage medium is not a propagation signal.

[0126] The term "data processing apparatus" covers various apparatus, devices and machines for processing data, including by way of example a programmable processor, a computer or a plurality of processors or computers. The apparatus may include special-purpose logic circuitry, for example an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the apparatus may also include: code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0127] A computer program (which may also be referred to or described as a program, software, software application, module, software module, script or code) may be written in any form of programming language including compiled or interpreted languages, declarative languages ​​or procedural languages, and it may be deployed in any form including as a standalone program or as a module, component, subroutine or other unit suitable for a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a file that holds portions of other programs or data - for example, one or more scripts stored in a markup language document, a single file dedicated to the program in question, or multiple coordinated files - for example, a file storing portions of one or more modules, subroutines or code. A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed in multiple sites and interconnected by a communication network.

[0128] As used in this specification, an "engine" or "software engine" refers to a software-implemented input / output system that provides an output that is different from the input. An engine can be a coded function block, such as a library, a platform, a software development kit ("SDK"), or an object. Each engine can be implemented on any suitable type of computing device, such as a server, a mobile phone, a tablet computer, a notebook computer, a music player, an e-book reader, a laptop or desktop computer, a PDA, a smart phone, or other fixed or portable device, which includes one or more processors and a computer-readable medium. Additionally, two or more engines can be implemented on the same computing device or on different computing devices.

[0129] The processes and logic flows described in this specification may be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating outputs. The processes and logic flows may also be performed by a dedicated logic circuit system (e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit)), and the device may also be implemented as the dedicated logic circuit system. For example, the processes and logic flows may be performed by a graphics processing unit (GPU), and the device may also be implemented as the graphics processing unit.

[0130] Computers suitable for executing computer programs include, by example, a central processing unit that can be a general-purpose microprocessor or a special-purpose microprocessor or both or any other kind. Typically, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more storage devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data--for example, a disk, a magneto-optical disk, or an optical disk, or be operably coupled to receive data from the mass storage device or transfer data to the mass storage device or perform these two operations. However, a computer does not need to have such a device. In addition, a computer can be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, for example, a universal serial bus (USB) flash drive, just to name a few examples.

[0131] Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including, for example, semiconductor memory devices—e.g., EPROM, EEPROM, and flash memory devices, magnetic disks—e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0132] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device for displaying information to the user—e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor—and a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and the input from the user may be received in any form including sound input, voice input, or tactile input. Additionally, a computer may interact with a user by sending documents to and receiving documents from a device used by a user, for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.

[0133] Embodiments of the subject matter described in this specification may be implemented in a computing system including, for example, a backend component such as a data server, or a middleware component such as an application server, or a frontend component - for example, a client computer with a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such backend components, middleware components, or frontend components. The components of the system may be interconnected by digital data communication in any form or medium, for example, a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), for example, the Internet.

[0134] A computing system may include clients and servers. Clients and servers are generally remote from each other and generally interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0135] Although this specification contains many specific implementation details, these details should not be viewed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be directed to specific embodiments of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations and even initially claimed as the same, one or more features from the claimed combination may be deleted from the combination in some cases, and the claimed combination may involve sub-combinations or variations of sub-combinations.

[0136] Similarly, although the operations are described in a particular order in the figures, this should not be understood as requiring that such operations be performed in the particular order shown or in a continuous order, or that all the operations shown be performed to obtain the desired result. In some cases, multitasking and parallel processing can be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0137] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying drawings do not necessarily need to be in the particular order shown or in a sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing can be advantageous.

Claims

1. A computer-implemented method for training a generator neural network and an encoder neural network, in, The generator neural network is configured to generate data items based on a set of potential values, the data items being samples representing a distribution of a set of training data items; wherein the encoder neural network is configured to generate a set of potential values ​​for a corresponding input data item, wherein the corresponding input data item comprises at least one of: data representing a still or moving image, or data representing a sound signal, or text data; wherein the method comprises jointly training the generator neural network, the encoder neural network, and a discriminator neural network, the discriminator neural network being configured to distinguish between samples of the distribution generated by the generator network and samples of the distribution that are not generated by the generator network, and wherein the discriminator neural network is configured to distinguish input pairs including a sample portion and a potential portion by being processed by the discriminator neural network; wherein the sample part and the potential part of the input pair include samples of the distribution generated by the generator neural network and corresponding potential value sets used to generate the samples, or training data items in the set of training data items and potential value sets generated by the encoder neural network based on the training data items; and wherein the training is based on a loss function comprising a joint discriminator loss term based on the sample portion and the latent portion of the input pair processed by the discriminator neural network and a single discriminator loss term based on only one of the sample portion or the latent portion of the input pair.

2. The method according to claim 1, wherein: The single discriminator loss term is based on the sample portion of the input pair.

3. The method according to claim 2, wherein: The single discriminator loss term includes a sample discrimination score generated based on processing the sample portion of the input pair using a sample discriminator sub-network.

4. The method according to claim 3, wherein: The sample discrimination score is further generated based on applying a projection to the output of the sample discriminator subnetwork.

5. The method according to claim 3, wherein: The sample discriminator subnetwork is based on a convolutional neural network.

6. The method according to claim 1, wherein: The single discriminator loss term is based on the latent portion of the input pair.

7. The method according to claim 6, wherein: The single discriminator loss term includes a latent discriminator score generated based on processing the latent portion of the input pair using a latent discriminator subnetwork.

8. The method according to claim 7, wherein: The latent discriminator score is further generated based on applying a projection to the output of the latent discriminator subnetwork.

9. The method according to claim 7, wherein: The latent discriminator subnetwork is based on a multi-layer perceptron.

10. The method according to claim 1, wherein: The loss function includes multiple single discriminator loss terms.

11. The method according to claim 1, wherein: The joint discriminator loss term includes a joint discriminator score generated using the joint discriminator sub-network.

12. The method according to claim 11, wherein: The joint discriminator subnetwork is configured to process an output of a sample discriminator subnetwork and an output of a latent discriminator subnetwork, wherein the sample discriminator subnetwork is configured to process the sample portion of the input pair and the latent discriminator subnetwork is configured to process the latent portion of the input pair.

13. The method according to claim 11, wherein: The joint discriminator score is further generated based on applying a projection to the output of the joint discriminator sub-network.

14. The method according to claim 11, wherein: The joint discriminator sub-network is based on a multi-layer perceptron.

15. The method according to claim 1, wherein: The loss function is based on the sum of the joint discriminator loss term and the single discriminator loss term.

16. The method according to claim 1, wherein: The loss function includes a hinge function applied to components of the loss function.

17. The method according to claim 1, wherein: The encoder neural network represents a probability distribution, and generating a set of potential values ​​includes sampling from the probability distribution.

18. The method according to claim 17, wherein: The output of the encoder neural network has a mean and a standard deviation that define a normal probability distribution.

19. The method according to claim 17, wherein: The set of potential values ​​is generated based on the reparameterized sampling.

20. The method according to claim 1, wherein: The encoder neural network is based on a convolutional neural network.

21. The method according to any one of claims 1 to 20, wherein: The training further comprises updating the discriminator neural network parameters and alternating between updating the encoder neural network parameters and the generator neural network parameters, wherein the updates are generated based on the loss function.

22. The method according to any one of claims 1 to 20, wherein: The training further includes jointly updating the encoder neural network parameters and the generator neural network parameters.

23. The method according to claim 21, wherein: Alternating the updating of the discriminator neural network parameters and the updating of the encoder neural network parameters and the generator neural network parameters comprises performing multiple updates of the discriminator neural network parameters followed by updates of the encoder neural network parameters and the generator neural network parameters.

24. A method of performing inference using an encoder neural network, the method comprising: Processing an input data item using the encoder neural network to generate a set of potential values ​​representing the input data item, wherein the input data item includes at least one of: data representing a still or moving image, or data representing a sound signal, or textual data; wherein the encoder neural network is jointly trained with a generator neural network and a discriminator neural network, the generator neural network being configured to generate data items based on a set of potential values, the data items being samples representing a distribution of a set of training data items, and the discriminator neural network being configured to distinguish between samples generated by the generator network and samples of the distribution that were not generated by the generator network; wherein the discriminator neural network is configured to distinguish input pairs including a sample portion and a potential portion by being processed by the discriminator neural network; wherein the sample part and the potential part of the input pair include samples of the distribution generated by the generator neural network and corresponding potential value sets used to generate the samples, or training data items in the set of training data items and potential value sets generated by the encoder neural network based on the training data items; wherein the training is based on a loss function comprising a joint discriminator loss term based on the sample portion and the latent portion of the input pair processed by the discriminator neural network and a single discriminator loss term based on only one of the sample portion or the latent portion of the input pair.

25. The method according to claim 24, further comprising: The input data item is classified based on the potential value representing the input data item.

26. The method of claim 24, further comprising: An action is performed with an agent based on the potential value representing the input data item.

27. A method of generating data items using a generator neural network, the method comprising: receiving a set of potential values; processing the set of potential values ​​using the generator neural network to generate a data item; wherein the generator neural network is jointly trained with an encoder neural network and a discriminator neural network, the encoder neural network being configured to generate a set of potential values ​​for corresponding data items, the discriminator neural network being configured to distinguish between samples generated by the generator network and samples of a distribution that are not generated by the generator network, wherein the corresponding input data items used by the encoder neural network to generate the set of potential values ​​include at least one of: data representing a still or moving image, or data representing a sound signal, or text data; wherein the generator neural network is configured to generate data items based on a set of potential values, the data items being samples representing a distribution of a set of training data items; wherein the discriminator neural network is configured to distinguish input pairs including a sample portion and a potential portion by being processed by the discriminator neural network; wherein the sample part and the potential part of the input pair include samples of the distribution generated by the generator neural network and corresponding potential value sets respectively used to generate the samples or training data items in the training data item set and potential value sets generated by the encoder neural network based on the training data items; wherein the training is based on a loss function comprising a joint discriminator loss term based on the sample portion and the latent portion of the input pair processed by the discriminator neural network and a single discriminator loss term based on only one of the sample portion or the latent portion of the input pair.

28. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 27.

29. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 27.

Citation Information

Patent Citations

  • Regularizing machine learning models

    CN108140143A

  • Method and device for establishing cross-domain joint distribution matching model and application thereof

    CN108960324A