Machine learnable system with standardized flows
By using a flexible prior variational autoencoder and a normalized stream function, the shortcomings of the CVAE model in capturing multimodal distributions are improved, the problem of latent variable failure is solved, and the ability of traffic participant generation and prediction in autonomous device controllers is enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2020-07-16
- Publication Date
- 2026-05-22
AI Technical Summary
Existing Conditional Variational Autoencoder (CVAE) models tend to over-regularize when generating multimodal distributions, making it difficult to capture complex multimodal distributions. Furthermore, latent variables become invalid, leading to the generation of unimodal data and poor probability distribution learning, particularly in the generation of traffic participants where unlikely events cannot be generated.
A variational autoencoder with flexible priors is employed, combined with a normalized stream function. The mapping between the target space and the latent space is generated through encoder and decoder functions, and the latent representation is mapped to the basic space using the normalized stream function. Multiple invertible normalized sub-functions are generated by a neural network to improve the learning ability, especially for pattern capture of complex probability distributions.
It effectively improves the difficulty of capturing multimodal distributions in CVAE models, reduces the failure of latent variables, and improves the diversity and accuracy of generated data. It is suitable for traffic participant classification and future behavior prediction in autonomous device controllers such as autonomous vehicles.
Smart Images

Figure CN112241756B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to machine-learnable systems, machine-learnable data generation systems, machine learning methods, machine-learnable data generation methods, and computer-readable media. Background Technology
[0002] Data generation based on learned distributions plays a crucial role in the development of various machine learning systems. For example, data similar to data previously measured using sensors can be generated. Such data can be used to train or test other devices, potentially other machine learning devices. For instance, given sensor data, such as from image sensors or LiDAR, might be available in the form of multiple images (e.g., videos), and it is expected that more of this data can be generated. Similarly, an image database including, for example, traffic participants or road signs can be used to test or train another device (e.g., an autonomous vehicle), and generating additional images of known traffic participants or road signs may be desirable.
[0003] Similarly, predicting the future states of agents or interacting agents in an environment is a critical capability for the successful operation of autonomous agents. For example, in many scenarios, this can be considered a generative problem or a sequence of generative problems. In complex environments such as real-world traffic scenarios, the future is highly uncertain, and therefore requires structured generation, for example, in the form of one-to-many mappings. This involves, for instance, predicting the possible future states of the world.
[0004] In Kihyuk Sohn's "Learning Structured Output Representation using Deep Conditional Generative Models," a Conditional Variational Autoencoder (CVAE) is described. CVAEs are conditional generative models used to generate outputs from Gaussian latent variables. This model is trained within a stochastic gradient variational Bayesian framework and allows for generation using stochastic feedforward inference.
[0005] CVAE can model complex multimodal distributions by factoring the distribution of future states using a set of latent variables, which are then mapped to possible future states. While CVAE is a general-purpose model capable of successfully modeling the future states of the world under uncertainty, it has been found to have drawbacks. For example, CVAE tends to over-regularize, the model has been found to struggle to capture multimodal distributions, and latent variable collapse has been observed.
[0006] In cases of posterior failure, the conditional decoding network forgets low-intensity patterns of the conditional probability distribution. This can lead to poor unimodal generation and learning of probability distributions. For example, in traffic participant generation, patterns corresponding to the conditional probability distribution for unlikely events (such as pedestrians entering / crossing the street) may not appear to be generated at all. Summary of the Invention
[0007] Having an improved system for data generation and a corresponding training system would be advantageous. The invention is defined by the independent claims; the dependent claims define advantageous embodiments.
[0008] In an embodiment, a machine-learnable system is configured to generate an encoder function that maps data in a target space to a latent representation in a latent space, a decoder function that maps the latent representation in the latent space to a target representation in the target space, and a normalized stream function that maps the latent representation to base points in a base space based on conditional data.
[0009] The encoder and decoder functions are arranged to generate parameters that define at least the mean of a probability distribution. The outputs of the encoder and decoder functions are determined by sampling the defined probability distribution with the generated mean, wherein the probability distribution defined by the encoder function has a predetermined variance.
[0010] Both VAE and CVAE models assume a standard Gaussian prior on the latent variables. This prior is found to play a role in generation quality, the tendency of CVAE to over-regularize, its difficulty in capturing multimodal distributions, and latent variable failure.
[0011] In this embodiment, a variational autoencoder with flexible priors is used to model the probability distribution. This improves at least some of these problems, such as the posterior failure problem of CVAE. Furthermore, forcing the encoder portion of the encoder / decoder section to have a fixed variance has been found to significantly improve learning in experiments.
[0012] Machine-learnable systems with stream-based priors can be used to learn the probability distribution of arbitrary data, such as images, audio, video, or other data obtained from sensor readings. Applications for learning generative models include, but are not limited to, traffic participant trajectory generation, generative classifiers, and synthetic data generation, for example, for training or validation purposes.
[0013] In this embodiment, the normalized stream function comprises a sequence of multiple reversible normalized stream subfunctions, one or more parameters of which are generated by a neural network. In particular, it has been found that using nonlinear normalized stream subfunctions for one or more subfunctions improves the ability to learn difficult probability distributions over a latent space, for example, having one or more modes, particularly minor modes.
[0014] Another aspect relates to a machine-learnable generative system configured to generate, for example, by: obtaining base points in a base space, applying an inverse conditionally normalized stream function to the base points conditioned on conditional data to obtain a latent representation, and applying a decoder function to the latent representation to obtain a generative target.
[0015] In an embodiment, a machine-learnable generative system is included in an autonomous device controller. For example, conditional action data may include sensor data from the autonomous device. The machine-learnable generative system can be configured to classify objects in the sensor data and / or predict future sensor data. The autonomous device controller can be configured to make decisions based on the classification. For example, the autonomous device controller can be configured and / or included in an autonomous vehicle, such as a car. For example, the autonomous device controller can be used to classify other traffic participants and / or predict their future behavior. The autonomous device controller can be configured to adapt the control of the autonomous device, for example, if the future trajectory of another traffic participant intersects with the trajectory of the autonomous device.
[0016] Machine-learnable systems and machine-learnable generative systems are electronic. These systems can be included within another physical device or system, such as a technical system, for controlling that physical device or system, for example, controlling its movement. Machine-learnable systems and machine-learnable generative systems can be devices.
[0017] Another aspect is machine learning methods and machine-learnable generative methods. Embodiments of these methods can be implemented on a computer as computer-implemented methods, or in dedicated hardware, or a combination of both. Executable code for embodiments of the methods can be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product includes non-transitory program code stored on a computer-readable medium for executing embodiments of the methods when the program product is executed on a computer.
[0018] In an embodiment, a computer program includes computer program code adapted to perform all or part of the steps of an embodiment of the method when the computer program is run on a computer. Preferably, the computer program is embodied on a computer-readable medium.
[0019] On the other hand, it is a method to make computer programs available for download. This aspect is used when computer programs are uploaded to, for example, Apple's App Store, Google Play Store, or Microsoft's Windows Store, and when computer programs are available for download from such stores. Attached Figure Description
[0020] Additional details, aspects, and embodiments will be described by way of example only with reference to the accompanying drawings. Elements in the figures are illustrated for simplicity and clarity and are not necessarily drawn to scale. In the figures, elements corresponding to those already described may have the same reference numerals. In the drawings,
[0021] Figure 1a An example of an embodiment of a machine-learnable system is illustrated schematically.
[0022] Figure 1b An example of an embodiment of a machine-learnable system is illustrated schematically.
[0023] Figure 1c An example of an embodiment of a machine-learnable data generation system is illustrated schematically.
[0024] Figure 1d An example of an embodiment of a machine-learnable data generation system is illustrated schematically.
[0025] Figure 2a An example of an embodiment of a machine-learnable system is illustrated schematically.
[0026] Figure 2b An example of an embodiment of a machine-learnable system is illustrated schematically.
[0027] Figure 2c An example of an embodiment of a machine-learnable system is illustrated schematically.
[0028] Figure 2d An example of an embodiment of a machine-learnable system is illustrated schematically.
[0029] Figure 3a An example of an embodiment of a machine-learnable data generation system is illustrated schematically.
[0030] Figure 3b An example of an embodiment of a machine-learnable data generation system is illustrated schematically.
[0031] Figure 4a The illustration schematically depicts diverse samples of known systems using k-means clustering.
[0032] Figure 4b The illustration schematically depicts diverse samples using k-means clustering as an example of an embodiment.
[0033] Figure 4c The prior learning in the embodiment is illustrated schematically.
[0034] Figure 4d.1 schematically illustrates examples of several basic facts.
[0035] Figure 4d.2 schematically illustrates several completions based on a known model.
[0036] Figure 4d.3 schematically illustrates several implementations according to the embodiment.
[0037] Figure 5a An example of an embodiment of a neural network machine learning method is illustrated schematically.
[0038] Figure 5b An example of an embodiment of a neural network machine-learnable data generation method is illustrated schematically.
[0039] Figure 6a A computer-readable medium having a writable portion including a computer program, according to an embodiment, is illustrated schematically.
[0040] Figure 6b A representation of a processor system according to an embodiment is shown schematically.
[0041] List of reference numbers in Figure 1-3 :
[0042] 110 Machine-Learning System
[0043] 112 Training data storage device
[0044] 130 processor system
[0045] 131 encoder
[0046] 132 decoder
[0047] 133 Standardized Flow
[0048] 134 training units
[0049] 140 Memory
[0050] 141 Target Space Storage Device
[0051] 142 Potential Space Storage Devices
[0052] 143 Basic Space Storage Device
[0053] 144 Conditional storage device
[0054] 150 communication interface
[0055] 160 Machine-Learning Data Generation System
[0056] 170 processor system
[0057] 172 Decoder
[0058] 173 Standardized Flow
[0059] 180 memory
[0060] 181 Target Space Storage Device
[0061] 182 Potential Space Storage Device
[0062] 183 Basic Space Storage Device
[0063] 184 Conditional storage device
[0064] 190 Communication Interface
[0065] 210 Target Space
[0066] 211 encoding
[0067] 220 Potential Space
[0068] 221 Decoding
[0069] 222 Conditional Standardized Flow
[0070] 230 base space
[0071] 240 Condition Space
[0072] 252 Training target data
[0073] 253 Training Condition Data
[0074] 254 Random Sources
[0075] 255 Conditional Data
[0076] 256 Target data generated
[0077] 257 Target data generated
[0078] 262 Machine-Learning Systems
[0079] 263 Parameter Transfer
[0080] 264 Machine-Learnable Data Generation System
[0081] 272 Machine-Learning Systems
[0082] 274 Machine-learnable data generation system
[0083] 330 Basic Space Sampler
[0084] 331 Conditional Encoder
[0085] 340 Standardized Flow
[0086] 341 Conditional Encoder
[0087] 350 potential spatial elements
[0088] 360 Decoding Network
[0089] 361 Target Space Elements
[0090] 362 conditions
[0091] 1000 computer-readable media
[0092] 1010 writable portion
[0093] 1020 Computer Program
[0094] 1110 (one or more) integrated circuits
[0095] 1120 Processing Unit
[0096] 1122 Memory
[0097] 1124 Application-Specific Integrated Circuit
[0098] 1126 Communication Components
[0099] 1130 Interconnection
[0100] 1140 processor system. Detailed Implementation
[0101] While the invention is permissible in many different forms of embodiments, one or more specific embodiments are shown in the accompanying drawings and will be described in detail herein. It should be understood that this disclosure is to be considered as an example of the principles of the invention and is not intended to limit the invention to the specific embodiments shown and described.
[0102] In the following description, for the sake of understanding, the elements of the embodiments are described in operation. However, it will be clear that the corresponding elements are arranged to perform the functions described as being performed by them.
[0103] Furthermore, the invention is not limited to the embodiments, and the invention lies in each novel feature or combination of features described herein or recited in mutually different dependent claims.
[0104] Figure 1a An example of an embodiment of the machine-learnable system 110 is illustrated schematically. Figure 1c An example embodiment of a machine-learnable data generation system 160 is illustrated schematically. For example, Figure 1a The machine-learnable system 110 can be used to train parameters of a machine-learnable system, such as a neural network, which can be used in a machine-learnable data generation system 160.
[0105] The machine-learnable system 110 may include a processor system 130, a memory 140, and a communication interface 150. The machine-learnable system 110 may be configured to communicate with a training data storage device 112. The storage device 112 may be a local storage device of the system 110, such as a local hard disk drive or memory. The storage device 112 may also be a non-local storage device, such as a cloud storage device. In the latter case, the storage device 112 may be implemented as a storage interface to the non-local storage device.
[0106] The machine-learnable data generation system 160 may include a processor system 170, a memory 180, and a communication interface 190.
[0107] Systems 110 and / or 160 can communicate with each other, external storage devices, input devices, output devices, and / or one or more sensors via a computer network. The computer network can be the Internet, an intranet, a LAN, a WLAN, etc. The system includes a connection interface arranged to communicate as needed, either within or outside the system. For example, the connection interface can include connectors, such as wired connectors (e.g., Ethernet connectors, optical connectors, etc.) or wireless connectors (e.g., antennas, such as Wi-Fi, 4G, or 5G antennas).
[0108] The execution of systems 110 and 160 can be implemented in a processor system (e.g., one or more processor circuits, such as a microprocessor), an example of which is shown in this document. Figure 1b and Figure 1d The diagram illustrates functional units that can be functional units of a processor system. For example, Figure 1b and Figure 1d These can be used as blueprints for the possible functional organization of a processor system. One or more processor circuits are not shown separately from the units in these diagrams. For example, Figure 1b and Figure 1dThe functional units shown can be implemented, in whole or in part, in computer instructions stored at systems 110 and 160 (e.g., in the electronic memory of systems 110 and 160) and can be executed by the microprocessors of systems 110 and 160. In hybrid embodiments, the functional units are implemented partly in hardware, such as as coprocessors (e.g., neural network coprocessors), and partly in software stored and executed on systems 110 and 160. Network parameters and / or training data can be stored locally at systems 110 and 160 or can be stored in cloud storage devices.
[0109] Figure 1b An example of an embodiment of a machine-learnable system 110 is illustrated schematically. The machine-learnable system 110 has access to a training storage device 112, which may include a collection of training target data (x). For example, in an embodiment, the machine-learnable system is configured to generate a model that can generate data similar to the training data, for example, data that appears to be sampled from the same probability distribution from the training data that appears to have been sampled. For example, the training data may include sensor data, such as image data, LiDAR data, etc.
[0110] Training storage device 112 may also or alternatively include training pairs, which include conditional data ( c ) and data generation objectives ( x The machine-learnable system 110 can be configured to train, for example, a model comprising multiple neural networks, to generate a data generation objective based on conditional data. For example, in this case, the system learns the conditional probability distribution of the training data, rather than the unconditional probability distribution.
[0111] Figure 2a An example of a data structure in an embodiment of the machine-learnable system 110 is illustrated schematically. The system 110 operates on the following: target data, such as data generation targets, collectively referred to as target space 210; latent data, such as internal or unobservable data elements, collectively referred to as latent space 220; and base points or base data, collectively referred to as base space 230. Figure 2b An example of a data structure in an embodiment of the machine-learnable system 110 is illustrated schematically. In addition to an additional space 240: a condition space 240 including conditional data, Figure 2b and Figure 2a same.
[0112] Conditional data can be useful for data generation, and especially in predictive applications. For example, conditional data could be the time of day. A data generation task can then be applied to generate data for a specific time of day. For example, a conditional task could be sensor data from different modalities. Then, a data generation task can be applied to generate data suitable for another modality of sensor data. For example, training data could include image data and LiDAR data; using one as the data generation target and the other as a condition, image data can be converted into LiDAR data, and vice versa.
[0113] Conditional probability can also be used for prediction tasks. For example, the condition could be past sensor values, and the data generation task could be future sensor data. In this way, data can be generated based on the possible distribution of future sensor values. Data processing can be applied to both target and conditional data. For example, conditional data could encode the past trajectories of one or more traffic participants, while the data generation target could encode the future trajectories of one or more traffic participants.
[0114] Various spaces can typically be implemented as vector spaces. Representations of target, potential, underlying, or conditional data can be as vectors. Elements of a space are typically called "points," even if the corresponding data (e.g., vectors) do not need to refer to physical points in physical space.
[0115] Figure 2c It schematically shows that it can be used Figure 2a An example of an embodiment of a machine-learnable system implemented using the data structure is shown. A machine-learnable system 262 trained using training target data 252 is illustrated. Once training is complete, the trained parameters can be transferred, at least partially, to a machine-learnable data generation system 264 in parameter transfer 263. Systems 262 and 264 can also be the same system. The machine-learnable data generation system 264 can be configured to receive input from a random source 254, for example, to sample points in a base space. The machine-learnable data generation system 264 can use an inverse normalized stream and a decoder to generate generated target data 256. The generated target data 256 should appear as if they were drawn from the same distribution as the target data 252. For example, data 252 and 256 can both be images. For example, data 252 can be images of traffic participants (such as cyclists and pedestrians), and the generated data 256 should be similar images. Data 256 can be used to test or train another device (e.g., an autonomous vehicle) that should learn to distinguish between traffic participants. Data 256 can also be used to test or debug other devices.
[0116] Figure 2d It schematically shows that it can be used Figure 2b An example of an embodiment of a machine-learnable system implemented using a data structure is shown. A machine-learnable system 272 is illustrated, which is trained using training target data 252 and corresponding conditions 253. Once training is complete, the trained parameters can be transferred, at least partially, to a machine-learnable data generation system 274 in parameter transfer 263. Systems 272 and 274 can be the same system. The machine-learnable data generation system 274 can be configured to receive input from a random source 254, for example, to sample points in a base space and new conditions 255. Condition 255 can be obtained from a sensor, such as the same or similar sensor that produces condition 253. The machine-learnable data generation system 274 can use an inverse conditional normalization stream and a decoder to generate generated target data 257. The generated target data 276 should appear as if they were drawn from the same conditional distribution as the target data 252 and conditions 253. For example, data 252 and 257 can both be images. For example, data 252 could be an image of traffic participants (such as cyclists and pedestrians), while condition 253 could be time of day information. Condition 255 could also be time of day information. The generated data 256 should be an image similar to image 252, conditioned on the time of day information in 255. Data 257 can be used similarly to data 256; however, condition 255 can be used to generate test data 256 specific to certain conditions. For example, images of road participants at specific times of day (e.g., at night, in the morning, etc.) could be generated. Conditions can also be used for prediction tasks, such as trajectory prediction.
[0117] System 110 may include storage devices for storing elements of various spaces. For example, system 110 may include a target space storage device 141, a potential space storage device 142, a basic space storage device 143, and optionally a conditional storage device 144 to store one or more elements of a corresponding space. The space storage device may be part of an electronic storage device, such as a memory.
[0118] System 110 may be configured with encoder 131 and decoder 132. For example, encoder 131 implements the decoding of target space 210 ( X Data generation target in ) x Mapped to latent space 220 ( Z The latent representation in ) The encoder function () Decoder 132 implements the latent space 220 (). Z The latent representation in ) z Mapped to target space 210 ( X The target representation in () The encoder and decoder functions are ideally close to being each other's inverses when fully trained, although this ideal will typically not be fully realized in practice.
[0119] The encoder and decoder functions are stochastic in the sense that they produce parameters of a probability distribution from which they sample the output. That is, the encoder and decoder functions are nondeterministic functions in the sense that they may return different results each time they are called, even with the same set of input values and even if the function definition remains unchanged. Note that the parameters of the probability distribution itself can be deterministically computed from the input.
[0120] For example, the encoder and decoder functions can be arranged to generate parameters that at least define the mean of a probability distribution. The outputs of the encoder and decoder functions are then determined by sampling the defined probability distribution with the generated mean. Other parameters of the probability distribution can be obtained from different sources.
[0121] In particular, it was found that it is especially advantageous if the encoder function does not generate variance. The mean of the distribution can be learned, but the variance can be a predetermined fixed value, such as an identity or a multiple of an identity. This does not require symmetry between the encoder and decoder; for example, the decoder could learn the mean and variance of, for example, a Gaussian distribution, while the corresponding encoder only learns the mean but not the variance (which is fixed). The fixed variance result in the encoder improved the learning in the experiments by as much as 40% (see experiments below).
[0122] In this embodiment, the data generation targets in target space 210 are elements that the model can be asked to generate once it has been trained. The distribution of the generated targets may or may not depend on conditional data referred to as conditions. The conditions can collectively be considered as condition space 240.
[0123] The generated targets or conditions (e.g., elements of target space 210 or condition space 240) can be low-dimensional, such as one or more values organized in a vector, such as velocity, one or more coordinates, temperature, classification, etc. The generated targets or conditions can also be high-dimensional, such as the output of one or more data-rich sensors (e.g., image sensors), such as LiDAR information. Such data can also be organized in a vector. In embodiments, the elements in spaces 210 or 240 can themselves be the output of an encoding operation. For example, systems outside the model described herein, either wholly or partially, can be used to encode future and / or past traffic conditions. Such encoding itself can be performed using a neural network including, for example, an encoder network.
[0124] For example, in an application, the output of a set of temperature sensors distributed throughout the engine (e.g., 10 temperature sensors) can be generated. The generated target and space 210 can be low-dimensional, such as 10-dimensional. The conditions can encode past usage information of the motor, such as operating time, total energy used, external temperature, etc. The conditions can have more than one dimension.
[0125] For example, in an application, future trajectories of traffic participants are generated. The target can be a vector describing the future trajectory. Conditions can encode information about past trajectories. In this case, both the target and condition spaces can be multidimensional. Therefore, in an embodiment, the dimensions of the condition and / or target spaces can be 1 or more, such as 2 or more, 4 or more, 8 or more, 100 or more, etc.
[0126] In this application, the target includes data from sensors (e.g., LiDAR or image sensors) that provide multidimensional information. The conditions may also include sensors (e.g., LiDAR or image sensors) that provide multidimensional information, such as other modalities. In this application, the system learns to predict LiDAR data based on image data, or vice versa. This is useful, for example, for generating training data, such as for training or testing autonomous driving software.
[0127] In this embodiment, the encoder and decoder functions form a variational autoencoder. The variational autoencoder learns a latent representation z of data x given by an encoding network z = ENC(X). This latent representation z can be reconstructed back into the original data space by a decoding network y = DEC(z), where y is the decoded representation of the latent representation z. The encoder and decoder functions can be similar to the variational autoencoder, except that no fixed priors are assumed on the latent space. Instead, a probability distribution on the latent space is learned in a normalized or conditionally normalized stream.
[0128] Both the encoder and decoder functions can be nondeterministic in the sense that they produce parameters of a (typically deterministic) probability distribution from which the output is sampled. For example, the encoder and decoder functions can generate the mean and variance of a multivariate Gaussian distribution. Interestingly, at least one of the encoder and decoder functions can be arranged to generate the mean, determining the function output by sampling a Gaussian distribution having a mean and a predetermined variance. That is, the function can be arranged to generate only the desired portion of the distribution's parameters. In an embodiment, the function generates only the mean of the Gaussian probability distribution but not the variance. The variance can then be a predetermined value, such as an identity, for example... ,in These are hyperparameters. Sampling of the function output can be accomplished by sampling from a Gaussian distribution with a generated mean and a predetermined variance.
[0129] In this embodiment, this is done for the encoder function, not the decoder function. For example, the decoder function might generate the mean and variance for sampling, while the encoder function only generates the mean but uses a predetermined variance for sampling. There could be an encoder function with a fixed variance, but a decoder function with the variance of a learnable variable.
[0130] We found that using fixed variance helps resist latent variable failure. In particular, we found that fixed variance in the encoder function is effective in resisting latent variable failure.
[0131] Encoder and decoder functions can include neural networks. By training the neural network, the system can learn to encode relevant information in the target space into an internal, latent representation. Typically, the dimension of the latent space 220 is lower than the dimension of the target space 210.
[0132] System 110 may further include normalized streams, sometimes simply referred to as streams. Normalized streams will latently represent ( z Mapped to base space 230 ( e The base space is the fundamental point in the target space. Interestingly, when the encoder maps points in the target space to the latent space, the probability distribution of the points in the target space is introduced into the probability distribution in the latent space. However, the probability distribution in the latent space can be very complex. The machine-learnable system 110 may include a normalized flow 133 (e.g., a normalized flow unit) to map elements of the latent space to yet another space—the base space. The base space is configured to have a predetermined probability distribution or one that can otherwise be easily computed. Typically, the latent space ( Z ) and basic space ( E They have the same dimension. Preferably, the flow is reversible.
[0133] In this embodiment, the normalized flow is a conditionally normalized flow. The conditional flow also maps from the latent space to the base space, but learns conditional probability distributions instead of unconditional probability distributions. Preferably, the flow is invertible, that is, for a given condition... c The flow relative to the base point e and potential points z It is reversible.
[0134] Figure 2a The encoding operation from the target space to the latent space is shown as arrow 211; the decoding operation from the latent space to the target space is shown as arrow 221, and the conditional flow is shown as arrow 222. Note that the flow is reversible, indicated by bidirectional arrows. Figure 2bThe condition space 240 is shown, where one or both of the flow and the base space can depend on the condition.
[0135] Conditionally normalized stream functions can be configured to optionally conditionally act on data ( c Given ) as a condition, the latent representation ( z Mapped to the base space ( E The basic point in ) or The normalized flow function depends on the latent point (). z ), and optionally also depends on the conditional action data ( c In this embodiment, the flow is a deterministic function.
[0136] Normalized flows can be achieved by utilizing parameterized invertible mappings. unknown distribution Transform into a known probability distribution To learn about potential space Z The probability distribution of the dataset. Mapping. It is called a normalized flow; Refers to learnable parameters. The probability distribution is known. Typically, it is a multivariate Gaussian distribution, but it can also be some other distribution, such as a uniform distribution.
[0137] Z raw data points z probability yes ,in ,Right now . It is an invertible mapping Jacobian determinant, Considering the change in probability quality due to invertible mappings. Given... Because invertible mappings can be computed. Output e And probability distribution As is known through construction, the standard multivariate normal distribution is frequently used. Therefore, by calculating data points... z Transformation value e ,calculate And compare the result with the Jacobian determinant. Multiply to calculate data points z probability It is easy. In the embodiment, the standardized flow... It also depends on the conditions, thus making standardization a conditional standardization flow.
[0138] One way to implement a normalized flow is as a sequence of multiple invertible sub-functions, called layers. These layers constitute the normalized flow. For example, in one embodiment, each layer includes a nonlinear flow that intersects with the hybrid layer.
[0139] A normalized stream can be transformed into a conditionally normalized stream by making its invertible subfunctions dependent on conditions. In fact, any regular (unconditional) normalized stream can be adapted to become a conditionally normalized stream by replacing one or more of its parameters with the output of a neural network that takes the conditions as input and outputs the layer's parameters. The neural network can have additional inputs, such as latent variables, or the output of previous layers, or parts thereof. Conversely, a conditionally normalized stream can be transformed into an unconditionally normalized stream by removing or fixing the conditional inputs. The examples below include subfunctions for conditionally normalized streams, which can, for example, learn dependencies on conditions; these examples can be transformed by removing the conditional inputs. c "Input is used to modify unconditionally normalized flows."
[0140] For example, in order to reversible mappings Modeling can be performed by constructing multiple layers or coupling layers. The Jacobian determinant of multiple stacked layers. J It is merely the product of the Jacobian determinants of the individual layers. Each coupling layer i From the previous layer i -1 gets the variable The variable that serves as input (or, in the case of the first layer), and produces the transformation. The variables of the transformation Including layers i The output of each individual coupling layer. This may include affine transformations, the coefficients of which depend at least on the conditions. One way to do this is to split the variable into left and right parts and set, for example,
[0141]
[0142]
[0143] In these coupling layers, the layer i The output is called Each It can be composed of the left and right halves, for example For example, these two halves could be vectors. A subset of. Half. One can remain unchanged, while the other half... It can be modified, for example, by an affine transformation with scaling and offset, which may depend solely on... The left half can have half, fewer, or more coefficients. In this case, because... Only depends on instead of The elements in the stream can be inverted.
[0144] Due to this construction, the Jacobian determinant of each coupled layer is simply a scaling network. The product of the outputs. Furthermore, the inverse of this affine transformation is easily computed, which facilitates easy sampling from the learned probability distribution used for the generative model. By allowing the invertible layers to include parameters given by a conditionally dependent learnable network, the flow can learn very useful complex conditional probability distributions. . scale (Scaling) and offset The output of (offset) can be a vector. Multiplication and addition operations can be component-wise. Each layer can have a function for... scale and offset Two networks.
[0145] In this embodiment, the left and right halves can be switched after each layer. Alternatively, an arrangement of layers can be used, for example, A random but fixed arrangement of elements. Other reversible layers can be used in addition to, or in place of, permutation and / or affine layers. Using left and right halves helps to make the flow reversible, but other learnable and reversible transformations can be used instead.
[0146] A permutation layer can be a reversible permutation of entries from vectors fed by the system. The permutation can be randomly initialized but remains fixed during training and inference. Different permutations can be used for each permutation layer.
[0147] In this embodiment, one or more of the affine layers are replaced with nonlinear layers. It has been found that nonlinear layers are better at transforming the probability distribution in the latent space into a normalized distribution. This is particularly true when the probability distribution in the latent space has multiple patterns. For example, the following nonlinear layers can be used.
[0148]
[0149]
[0150] As shown above, operations on vectors can be performed component-by-component. The non-linear example above uses a neural network: offset , scale , C () D ()and G Each of these networks can depend on parts and conditions of the output of the previous layers. cThe network can output vectors. Other useful layers include convolutional layers, such as 1×1 convolutional layers, and layers that multiply with an invertible matrix M.
[0151] .
[0152] This matrix can be the output of a neural network; for example, this matrix can be... .
[0153] Another useful layer is the activation layer, where the parameters do not depend on the data, for example...
[0154] .
[0155] Activation layers can also have conditional parameters, for example,
[0156] .
[0157] network s ()and o () can produce a single scalar or vector.
[0158] Another useful layer is the shuffling layer, or permutation layer, where coefficients are arranged according to a permutation. When initializing the layer for the first time for the model, the permutation can be chosen randomly but remains fixed thereafter. For example, the permutation may not depend on the data or training.
[0159] In the embodiments, there are multiple layers, such as 2, 4, 8, 10, 16 or more. If each layer is followed by an arrangement of layers, then this number can be twice as large. The flow maps from the latent space to the base space, or vice versa, because the flow is reversible.
[0160] The number of neural networks involved in the normalized flow can be as large as or greater than the number of learnable layers. For example, the affine transformation example given above can use two layers. In embodiments, the number of layers in the neural network can be limited to, for example, one or two hidden layers. In Figure 2, the effect of the condition space 240 on the normalized flow has been illustrated using a diagram oriented towards flow 222.
[0161] For example, in an embodiment, the conditional normalization flow may include multiple layers of different types. For example, the layers of the conditional normalization flow may be organized into blocks, with each block comprising multiple layers. For example, in an embodiment, a block may include nonlinear layers, convolutional layers, scaled activation layers, and shuffling layers. For example, there may be multiple such blocks, such as two or more, four or more, sixteen or more, etc.
[0162] Note that the number of neural networks involved in conditionally normalized flows can be quite high, for example, more than 100. Furthermore, the networks can have multiple outputs, such as vectors or matrices. These networks can be learned using methods such as maximum likelihood learning.
[0163] For example, a nonlinear unconditionally normalized flow can also have multiple blocks with multiple layers. For example, such a block can include a nonlinear layer, an activation layer, and a shuffling layer.
[0164] Therefore, in the embodiment, it can have: vector space 210 X , it is n 220-dimensional, used to generate targets, such as future trajectories; latent vector space 220 Z , it is d 230-dimensional, for example, for latent representation; vector space 230 E It can also be d It is dimensional and has a fundamental distribution, such as a multivariate Gaussian distribution. Furthermore, the optional vector space 240 can represent conditions, such as past trajectories, environmental information, etc. The normalized flow 222 is... Figure 2a The flow operates between spaces 220 and 230; the conditionally normalized flow 222 is conditional on elements from space 240. Figure 2b It operates between space 220 and space 230.
[0165] In the embodiments, the base space allows for easy sampling. For example, the base space may be a multivariate Gaussian distribution on which (e.g., The probability distribution in the base space is a vector space. In this embodiment, the probability distribution in the base space is a predetermined probability distribution.
[0166] Another option is to conditionally define the distribution of the underlying space, preferably while still allowing for easy sampling. For example, one or more parameters of the underlying distribution can be generated by a neural network, which will at least be used by some other neural network. The conditions (e.g., the distribution can be) (The input is taken as the neural network.) It can be learned along with other networks in the model. For example, neural networks. A value (e.g., the mean) can be calculated, and a fixed distribution can be shifted through that value (e.g., the mean), for example, by being added to that value (e.g., the mean). For example, if the basic distribution is a Gaussian distribution, then the conditional basic distribution could be... ,For example For example, if the distribution is uniform over intervals such as [0, 1], then the conditional distribution can be... or To keep the mean equal to .
[0167] System 110 may include training unit 134. For example, training unit 134 may be configured to train an encoder function, a decoder function, and a normalized stream on a set of training pairs. For example, training may attempt to minimize the reconstruction loss of the cascade of the encoder and decoder functions, and minimize the difference between the probability distribution in the base space and the cascade of the encoder and normalized stream functions applied to the set of training pairs.
[0168] Training can follow these steps (in this case, to generate sensor data from traffic participants, for example, for training or testing purposes).
[0169] 1. Encode the sensor data from the training samples using an encoder model. For example, this function can map the sensor data from the target space X to a distribution in the latent space Z.
[0170] 2. Sample points in the latent space Z from the encoder's predicted mean and from a given fixed variance.
[0171] 3. Use a decoder to decode the sampled points. This is the process of returning from the latent space Z to the target space X.
[0172] 4. Calculate ELBO
[0173] a. Using sensor data to calculate the likelihood under the flow prior.
[0174] b. Use decoded sensor data to calculate data likelihood loss.
[0175] 5. Perform gradient descent steps to maximize ELBO.
[0176] The example above does not use conditions. Another example follows these steps (in this case, to generate future trajectories, which may depend on conditions):
[0177] 1. Encode future trajectories from training samples using an encoder model. For example, this function can map future trajectories from the target space X to a distribution in the latent space Z.
[0178] 2. Sample points in the latent space Z from the encoder's predicted mean and from a given fixed variance.
[0179] 3. Use a decoder to decode the future trajectory. This is the return from the potential space Z to the target space X.
[0180] 4. Calculate ELBO
[0181] a. Use future trajectories to compute the likelihood of the flow prior; the flow also depends on the conditions corresponding to the future trajectories.
[0182] b. Use the decoded trajectory to calculate the data likelihood loss.
[0183] 5. Perform gradient descent steps to maximize ELBO.
[0184] Training can also include the following steps
[0185] 0. Encode the conditions from the training pairs using a conditional encoder. This function can compute the mean of the (optional) conditional basis distribution in the basis space. The encoding can also be used for conditionally normalized streams.
[0186] For example, this training could include maximizing the lower bound of evidence, ELBO. Given conditional data... c ELBO can be targeted at training objectives. x probability or conditional probability The lower bound of ELBO. For example, in an embodiment, ELBO can be defined as...
[0187]
[0188] in, It is a probability distribution and Kullback-Leibler divergence, probability distribution Defined by the fundamental distribution and the conditionally normalized flow. For this purpose, the conditionally normalized flow is used to... Transform it into a probability distribution that is easier to evaluate, such as the standard normal distribution. The normalized flow can represent a class of distributions that are much richer than the standard prior in the latent space.
[0189] When using normalized flow to transform the basic distributions into more complex priors for the latent space, the formula for the KL part of ELBO can be as follows:
[0190]
[0191] in It is a conditionally standardized flow and It is the Jacobian of the conditionally normalized flow. By using this conditional prior based on complex flows, autoencoders can more easily learn complex conditional probability distributions because they are not restricted by the simple Gaussian prior assumptions over the latent space. The above formula can be adapted to the unconditional case, for example, by omitting the dependence on the condition.
[0192] In this embodiment, the encoder function, decoder function, and normalized stream are trained together. The training can be performed in batches or partially in batches.
[0193] Figure 1d An example embodiment of a machine-learnable generative system 160 is illustrated schematically. The machine-learnable generative system 160 is similar to system 110, but does not need to be configured for training. For example, this means that system 160 may not require encoder 131 or access to training storage device 112 or training unit. On the other hand, system 160 may include sensors and / or sensor interfaces for receiving sensor data, for example, to construct conditions. For example, system 160 may include decoder 172 and normalization stream 173. Normalization stream 173 can be configured for the reverse direction from the base space to the latent space. Normalization stream 173 may be a conditional normalization stream. Like system 110, system 160 may also include storage devices for various types of points, such as target space storage device 181, latent space storage device 182, base space storage device 183, and optional condition storage device 184, to store one or more elements of the corresponding space. The space storage device may be part of an electronic storage device, such as a memory.
[0194] Decoder 172 and stream 173 can be trained by a system such as system 110. For example, system 160 can be configured to train under given conditions as follows ( c Determine the generated targets in the target space. x )
[0195] - For example, by sampling the base space (e.g., using a sampler), the base points in the base space can be obtained. e ).
[0196] - Based on conditional data ( c The inverse conditional normalized stream function is applied to the base points as a condition. e ) to obtain potential representation ( z ),as well as
[0197] - Apply the decoder function to the latent representation ( z ) to obtain the generation target ( x ).
[0198] The use of conditions can be omitted. Note that decoder 172 can output the mean and variance instead of directly outputting the target. In the case of mean and variance, in order to obtain the target, it is necessary to sample from the defined probability distribution (such as a Gaussian distribution).
[0199] Each time a base point is obtained in the base space ( eThis allows for the generation of corresponding targets. In this way, a set of multiple generated targets can be aggregated. Several ways exist for using generated targets. For example, given a generated target, control signals can be calculated, which are used, for example, by autonomous devices such as autonomous vehicles. For instance, the control signal could be to avoid traffic participants, for example, in all generated future scenarios.
[0200] Multiple generation targets can be processed statistically; for example, they can be averaged, or the top 10% of generation can be selected, and so on.
[0201] Figure 3a An example of an embodiment of a machine-learnable generative system is illustrated schematically. Figure 3a The illustration shows the processes that can be performed in the embodiments used for generation.
[0202] A base spatial sampler 330 samples from a base distribution. Using the sampled base points from sampler 330, the base points are mapped to points in the latent spatial elements 350 via a normalization stream 340. The parameters of the normalization stream can be generated by one or more networks 341. The latent spatial elements 350 are mapped to target spatial elements 361 via a decoding network 360; this may also involve sampling. In an embodiment, the decoding network 360 may include a neural network.
[0203] Figure 3b An example of an embodiment of a machine-learnable generative system is illustrated schematically. Figure 3b The illustration shows the processes that can be performed in the embodiments used for generation.
[0204] Figure 3b Condition 362 is illustrated, for example, past sensor information from which future sensor information is to be generated. Condition 362 can be input to conditional encoder 341. Conditional encoder 341 can include one or more neural networks to generate parameters for layers in the conditional flow. Additional inputs to conditional encoder 341 may exist, such as base points and intermediate points of the conditional normalization flow. The networks in encoder 341 can be deterministic, but they can also output probability distribution parameters and use sampling steps.
[0205] Condition 362 can be input to conditional encoder 331. Conditional encoder 331 may include a neural network to generate parameters for a probability distribution from which base points can be sampled. In an embodiment, conditional encoder 331 generates a mean, which is used to sample base points with a fixed variance. Conditional encoder 331 is optional. Base sampler 360 may use an unconditional predetermined probability distribution.
[0206] The base spatial sampler 330 samples from the base distribution. This can be done using parameters generated by the encoder 331. The base distribution can be fixed instead. The conditional encoder 331 is optional.
[0207] Using parameters from the layer of encoder 341 and sampled base points from sampler 330, the base points are mapped to points in latent space element 350 via normalization stream 340. Latent space element 350 is then mapped to target space element 361 via decoding network 360; this may also involve sampling.
[0208] In this embodiment, the decoding network 360, the conditional encoder 331, and the conditional encoder 341 may include neural networks. The conditional encoder 341 may include multiple neural networks.
[0209] like Figure 3a Alternatively, similar processing for the application phase shown in 3b can be performed during the training phase, and there may be an additional mapping from the target space to the latent space using the encoder network.
[0210] An example application of learning conditional probability distributions is predicting the future location x of traffic participants. For instance, this can be achieved from conditional probability distributions. Mid-sampling. This allows sampling of the most likely future traffic participant location x given a feature f and a future time t. The car can then drive to a location where no location sample x has been generated, since that location is most likely to have no other traffic participants.
[0211] In embodiments, such as during training pairs or during application, conditional data ( c This includes past trajectory information of traffic participants, and the target is generated therein. x This includes information on the future trajectories of traffic participants.
[0212] Encoding past trajectory information can be performed as follows. Past trajectory information, such as past trajectory data, can be encoded into a first fixed-length vector using a neural network, such as a recurrent neural network (e.g., LSTM). Environmental map information can be encoded into a second fixed-length vector using a CNN. Interacting traffic participant information, such as interactive traffic participant information, can be encoded into a third fixed-length vector. One or more of the first, second, and / or third vectors can be cascaded into conditions. Interestingly, the neural network used for encoding the conditions can also be trained together with system 110. Encoding the conditions can also be part of a network that encodes parameters of the flow and / or the underlying distribution. Networks used to encode conditional information—in this case, past trajectory and environmental information—can be trained together with the remaining networks used in the model; they can even share parts of their networks with other networks, for example, networks encoding conditions, such as those for the underlying distribution or conditional flow, can share parts of their network bodies with each other.
[0213] Trained neural network devices can be applied in autonomous device controllers. For example, the conditional data of the neural network can include sensor data from the autonomous device. Target data can be future aspects of the system, such as generated sensor outputs. The autonomous device can perform movement at least partially autonomously, for example, modifying movement based on the device's environment without the user specifying the modification. For example, the device can be a computer-controlled machine, such as a car, robot, vehicle, household appliance, power tool, manufacturing machine, etc. For example, the neural network can be configured to classify objects in sensor data. The autonomous device can be configured to make decisions based on the classification. For example, if the network can classify objects in the environment surrounding the device and can, for example, classify other traffic near the device—such as people, cyclists, cars, etc.—stop, slow down, turn, or otherwise modify the device's movement.
[0214] In various embodiments of systems 110 and 160, the communication interface can be selected from a variety of alternatives. For example, the interface can be a network interface to a local area network or wide area network (e.g., the Internet), a storage interface to an internal or external data storage device, a keyboard, an application interface (API), etc.
[0215] Systems 110 and 160 may have a user interface, which may include known components such as one or more buttons, a keyboard, a display, a touchscreen, etc. The user interface may be configured to accommodate user interaction for configuring the system, training a network on a training set, or applying the system to new sensor data, etc.
[0216] The storage device can be implemented as an electronic storage device such as flash memory or a magnetic storage device such as a hard disk. The storage device may include multiple discrete memories, which together constitute storage devices 140 and 180. The storage device may include temporary storage, such as RAM. The storage device may be a cloud storage device.
[0217] System 110 can be implemented in a single device. System 160 can be implemented in a single device. Typically, systems 110 and 160 each include a microprocessor that executes appropriate software stored at the system; for example, this software may have been downloaded and / or stored in a corresponding memory, such as volatile memory like RAM or non-volatile memory like flash memory. Alternatively, the system can be implemented entirely or partially in programmable logic, such as a field-programmable gate array (FPGA). The system can be implemented entirely or partially as a so-called application-specific integrated circuit (ASIC), such as an integrated circuit (IC) customized for its specific purpose. For example, the circuitry can be implemented in CMOS, for example, using a hardware description language such as Verilog, VHDL, etc. In particular, systems 110 and 160 may include circuitry for neural network evaluation.
[0218] The processor circuitry can be implemented in a distributed manner, for example, as multiple sub-processor circuits. The storage device can be distributed across multiple distributed sub-storage devices. Part or all of the memory can be electronic memory, magnetic memory, etc. For example, the storage device can have volatile and non-volatile components. Parts of the storage device can be read-only.
[0219] Several additional optional refinements, details, and embodiments are described below. The notation below differs slightly from that above in that it utilizes " x "Indication conditions, using " z "Indicates the elements of the potential space, and utilizes" y "Indicates the elements of the target space."
[0220] Conditional priors can be learned using conditional normalized flows. Priors based on conditional normalized flows can be learned using a simple basis distribution. Start, and then you can use a reversible normalized flow. of n The layer will distribute the basic distribution Transform into a more complex prior distribution on the latent variables ,
[0221] (2).
[0222] Given basic density and each layer of the transformationi Jacobi The log-likelihood of the latent variable z can be expressed using the variable change formula.
[0223] (3).
[0224] One option is to consider a spherical Gaussian as the fundamental distribution. This allows for easy sampling from the underlying distribution and therefore conditional priors. This enables the learning of complex multimodal priors. Multiple layers of nonlinear flow can be applied on top of the basic distribution. It has been found that nonlinear conditionally normalized flow allows for conditional priors. It is highly multimodal. Nonlinear conditional flow also allows for complex conditions to be applied based on past trajectories and environmental information.
[0225] The KL divergence term used for training may not have a simple closed-form expression for conditional flow-based priors. However, the KL divergence can be computed by evaluating the likelihood over the underlying distribution rather than complex conditional priors. For example, it can be done using:
[0226] (4)
[0227] in It is the entropy of the variational distribution. Therefore, ELBO can be expressed as:
[0228] (5).
[0229] To learn complex conditional priors, the will distribution in (5) can be jointly optimized. and conditional priors The variational distribution attempts to match the conditional prior, and the conditional prior attempts to match the variational distribution, such that ELBO(5) is maximized and the data is well interpreted. This model will be referred to in this paper as... conditional flow -VAE (or CF-VAE).
[0230] In an embodiment, The variance can be fixed at C. This results in a weaker inference model, but the entropy term becomes constant and no longer requires optimization. In detail, it can be done using...
[0231] (6).
[0232] Furthermore, the maximum possible shrinkage also becomes bounded, thus setting an upper bound on the log-Jacobi. Therefore, during training, this encourages the model to focus on interpreting the data and prevents degeneration, where entropy or the log-Jacobi term dominates over the log-likelihood of the data, leading to more stable training and preventing... conditional flow- The latent variables of VAE are invalid.
[0233] In this embodiment, the decoder function can be conditional. This allows the model to more easily learn an effective decoder function. However, in this embodiment, the conditional effect of the decoder on condition x is removed. This is possible because the conditional prior is learned—the latent prior distribution... Information specific to x can be encoded—unlike standard CVAEs that use data-independent priors. This ensures that the latent variable z encodes information about future trajectories and prevents failure. In particular, this prevents situations where the model might ignore sub-modes and only model the main mode of the conditional distribution.
[0234] In the first example application, this model was used for trajectory generation. Past trajectory information can be encoded into a fixed-length vector using an LSTM. For efficiency, the conditional encoder can be shared between the conditional stream and the decoder. CNNs can be used to encode environmental map information into fixed-length vectors. The CVAE decoder can utilize this information as a condition. To encode information about interactive traffic participants / agents, convolutional social pooling from A. Atanov et al.'s "Semi-conditional normalizing flows for semi-supervised learning" can be used. For example, for efficiency, it can utilize... Convolution is used to exchange LSTM trajectory encoders. Specifically, convolutional social pooling can pool information using a grid superimposed on the environment. This grid can be represented using tensors, where past trajectory information of traffic participants is aggregated into tensors indexed corresponding to the grid in the environment. Past trajectory information can be encoded using LSTM before being aggregated into grid tensors. For computational efficiency, trajectory information can be directly aggregated into tensors, and then... Convolution is used to extract trajectory-specific features. Finally, it can be applied... Several layers of convolution, for example, to capture interactive perceptual contextual features of traffic participants in a scene. .
[0235] As mentioned earlier, conditional flow architecture can include streams. There are several layers, with dimensionality mixing between them. Conditional context information can be aggregated into a single vector. This vector can be used to conditionally act at one or more layers or at each layer to affect the conditional distribution. Modeling.
[0236] Another example is illustrated using the MNIST sequence dataset, which comprises handwritten stroke sequences of MNIST bits. For evaluation, complete strokes were generated given the first ten steps. This dataset is interesting because the distribution of stroke completions is highly multimodal, with a wide variation in the number of patterns. Given an initial stroke of 2, completing 2, 3, or 8 is possible. On the other hand, given an initial stroke of 1, the only possible completion is 1 itself. Priors based on data-related conditional flow perform very well on this dataset.
[0237] Figure 4a and Figure 4b The diagram shows a variety of samples using k-means clustering. The number of clusters was manually set to the expected number of digits based on the initial stroke count. Figure 4a Using the known model BMS-CVAE, and Figure 4b Use examples. Figure 4c The priors learned in the embodiments are shown.
[0238] The table below compares the two embodiments (starting with CF) with known methods. Evaluations were performed on MNIST sequences, and negative CLL scores are given: lower is better.
[0239] method -CLL CVAE 96.4 BMS-CVAE 95.6 CF-VAE 74.9 CF-VAE - 77.2
[0240] The table above uses variational posteriors with fixed variance. The CF-VAE, and its conditional flow (CNLSq) version, utilizes a flow-based CF-VAE with affine coupling. The conditional log-likelihood (CLL) metric is used for evaluation, and the same model architecture is used across all baselines. The LSTM encoder / decoder has 48 hidden neurons, and the latent space is 64-dimensional. The CF-VAE uses a standard Gaussian prior. The CF-VAE outperforms known models by more than 20%.
[0241] Figure 4c The captured patterns and learning-conditional-flow-based priors are illustrated. To enable visualization, the density of the conditional-flow-based priors is projected to 2D using TSNE, followed by kernel density estimation. It can be seen that the conditional-flow-based priors... The number of patterns in the data reflects the data distribution. The number of patterns in the data. In contrast, BMS-CVAE cannot fully capture all patterns—its generation is pushed to the mean of the distribution due to the standard Gaussian prior. This highlights the advantage of using priors based on data correlation streams to capture the highly multimodal distribution of handwritten strokes.
[0242] Next, it was found that if the posterior condition is not fixed... The variance, for example in the encoder, results in a 40% performance degradation. This is because entropy, or the log-Jacobi term, dominates during optimization. It was also found that using priors based on affine conditional flows leads to a performance degradation (77.2 vs. 74.9 CLL). This illustrates the advantage of nonlinear conditional flows in learning highly nonlinear priors.
[0243] Figure 4d.1 schematically illustrates several examples from the MNIST sequence dataset. These are the basic fact data. The bold portions of the data starting with an asterisk are the data to be completed. Figure 4d.2 shows the completion according to the BMS-CVAE model. Figure 4d.3 shows the completion according to the embodiment. It can be seen that the completion according to Figure 4d.3 follows the basic facts much more closely than those completions in Figure 4d.2.
[0244] Another example utilizes the Stanford Drone dataset, which includes trajectories of traffic participants such as pedestrians, cyclists, and cars from video captured by drones. The scenes are densely packed with traffic participants, and the layout contains many intersections, resulting in highly multi-model traffic participant trajectories. Evaluation uses 5-fold cross-validation and a single-criteria training-test split. The table below shows the results compared to several known methods for one embodiment.
[0245] method mADE mFDE Social GANs
[14] 27.2 41.4 MATF GAN
[37] 22.5 33.5 SoPhie
[28] 16.2 29.3 Target generation [7] 15.7 28.1 CF-VAE 12.6 22.3
[0246] A CNN encoder is used to extract visual features from the last observed RGB image of the scene. These visual features act as additional conditions on the conditionally normalized flow. The CF-VAE model with RGB input performs best—outperforming existing technologies by over 20% (Euclidean distance). (4 seconds). Conditional flow can use visual scene information to fine-tune the learned conditional priors.
[0247] Another example uses the HighD dataset, which comprises vehicle trajectories recorded using drones over highways. The HighD dataset is challenging because only 10% of the vehicle trajectories contain lane changes or interactions—a single dominant pattern exists alongside several sub-patterns. Therefore, it is challenging for schemes predicting a single average future trajectory (e.g., targeting the dominant pattern) to outperform. For example, simple feedforward (FF) models perform well. This makes the dataset even more challenging because VAE-based models frequently suffer posterior failures when a single pattern dominates. VAE-based models trade off the cost of ignoring sub-patterns by making the posterior latent distribution fail as a standard Gaussian prior. Experiments confirm that predictions from known systems such as CVAEs are typically linear continuations of trajectories; that is, they exhibit failures to the dominant pattern. However, the predicted trajectories according to this embodiment are much more diverse and cover events such as lane changes; they include sub-patterns.
[0248] CF-VAE significantly outperforms known models, demonstrating that posterior failure does not occur. To further combat posterior failure, the additional conditional effect of past trajectory information on the decoder is removed. Furthermore, the addition of contextual information about interacting traffic participants further improves performance. Conditional CNLSq streams can effectively capture complex conditional distributions and learn complex data-related priors.
[0249] Figure 5a An example of an embodiment of a neural network machine learning method 500 is illustrated schematically. Method 500 may be computer-implemented and includes...
[0250] - Access the set of (505) training pairs, which include conditional data ( c ) and generating objectives ( x ); for example, it comes from electronic storage devices.
[0251] -Use encoder functions to divide the target space ( X The generated target in ) x Mapping (510) to the latent space ( Z The latent representation in ) ),
[0252] -Use the decoder function to extract the latent space ( Z The latent representation in ) z Mapping (515) to the target space ( X The target representation in () ),
[0253] -Use conditionally normalized stream functions to conditionally apply data ( c ) as a condition for the latent representation ( zMapping (520) to the base space ( E The basic point in ) The encoder function, decoder function, and conditionally normalized stream function are machine-learnable functions, and the conditionally normalized stream function is invertible; for example, the encoder, decoder, and conditional stream function can therefore include neural networks.
[0254] - Train (525) encoder functions, decoder functions and conditionally normalized streams on a set of training pairs. This training involves minimizing the reconstruction loss of the cascade of encoder and decoder functions and minimizing the difference between the probability distribution in the base space and the cascade of encoder and conditionally normalized stream functions applied to the set of training pairs.
[0255] Figure 5b An example of an embodiment of a neural network machine-learnable generation method 550 is illustrated schematically. Method 550 may be computer-implemented and includes...
[0256] - Obtain (555) conditional data ( c ),
[0257] - The target is generated by determining (560) as follows.
[0258] - Obtain the base point in the (565) base space ( e ),
[0259] -Use conditional data ( c ) as a condition to base point ( e Mapping to the latent space ( Z The latent representation in ) The inverse conditionally normalized stream function of ) is used to conditionally normalize the data ( c ) conditionally map the fundamental points (570) to the latent representation, and
[0260] -Use the potential space ( Z The latent representation in ) z Mapped to the target space ( X The target representation in () The decoder function is used to map the latent representation (575) to obtain the generated target. The decoder function and the conditionally normalized stream function may have been learned according to a machine-learnable system as described in this paper.
[0261] For example, machine learning methods and machine-learnable generative methods can be computer-implemented methods. For example, accessing training data and / or receiving input data can be accomplished using communication interfaces such as electronic interfaces, network interfaces, and memory interfaces. For example, storing or retrieving parameters, such as network parameters, can be accomplished using electronic storage devices such as memory or hard drives. For example, applying the neural network to the training data and / or adjusting the stored parameters to train the network can be accomplished using electronic computing devices such as computers. Encoders and decoders can also output the mean and / or variance, rather than directly as the output. In the case of mean and variance, sampling from a defined Gaussian distribution is necessary to obtain the output.
[0262] During training and / or application, a neural network can have multiple layers, which may include, for example, convolutional layers. For instance, a neural network can have at least 2, 5, 10, 15, 20, or 40 hidden layers, or more. The number of neurons in the neural network can be, for example, at least 10, 100, 1000, 10000, 100000, 1000000, or more.
[0263] As will be apparent to those skilled in the art, many different ways are possible to perform the method. For example, the steps may be performed in the order shown, but the order of the steps may vary or some steps may be performed in parallel. Furthermore, other method steps may be inserted between steps. The inserted steps may represent a refinement of the method as described herein, or they may be unrelated to the method. For example, some steps may be performed at least partially in parallel. Moreover, a given step may not be fully completed before the next step begins.
[0264] Embodiments of the method can be executed using software, which includes instructions for causing a processor system to perform methods 500 and / or 550. The software may include only those steps taken by a specific sub-entity of the system. The software can be stored in a suitable storage medium, such as a hard disk, floppy disk, memory, optical disk, etc. The software can be transmitted as a signal along a wired or wireless route or using a data network (e.g., the Internet). The software can be made available for download and / or for remote use on a server. Embodiments of the method can be executed using a bitstream arranged to configure programmable logic (e.g., a field-programmable gate array (FPGA)) to perform the method.
[0265] It will be appreciated that the invention is also extended to computer programs adapted for practicing the invention, particularly computer programs on or in a carrier. The program may be in the form of source code, object code, intermediate source code, and object code such as partially compiled form, or in any other form suitable for use in an implementation of an embodiment of the method. Embodiments relating to computer program products include computer-executable instructions corresponding to each processing step of at least one of the illustrated methods. These instructions may be subdivided into subroutines and / or stored in one or more files that may be statically or dynamically linked. Another embodiment relating to computer program products includes computer-executable instructions corresponding to each component of at least one of the illustrated systems and / or products.
[0266] Figure 6a A computer-readable medium 1000 with a writable portion 1010 according to an embodiment is shown. The writable portion 1010 includes a computer program 1020, which includes instructions for inducing a processor system to perform machine learning methods and / or machine-learnable generative methods. The computer program 1020 may be embodied on the computer-readable medium 1000 as a physical label or by means of magnetization of the computer-readable medium 1000. However, any other suitable embodiments are contemplated. Furthermore, it will be appreciated that although the computer-readable medium 1000 is shown herein as an optical disc, the computer-readable medium 1000 may be any suitable computer-readable medium, such as a hard disk, solid-state storage, flash memory, etc., and may be non-recordable or recordable. The computer program 1020 includes instructions for inducing a processor system to perform the methods of training and / or applying machine-learnable models—for example, including training or applying one or more neural networks.
[0267] Figure 6b The diagram illustrates a schematic representation of a processor system 1140 according to an embodiment of a machine learning system and / or a machine-learnable generative system. The processor system includes one or more integrated circuits 1110. Figure 6bThe diagram schematically illustrates the architecture of one or more integrated circuits 1110. Circuit 1110 includes a processing unit 1120 (e.g., a CPU) for running computer program components to perform methods according to embodiments and / or implement modules or units thereof. Circuit 1110 includes a memory 1122 for storing programming code, data, etc. A portion of the memory 1122 may be read-only. Circuit 1110 may include a communication element 1126, such as an antenna, a connector, or both, etc. Circuit 1110 may include an application-specific integrated circuit 1124 for performing some or all of the processing defined in the methods. Processor 1120, memory 1122, application-specific IC 1124, and communication element 1126 may be interconnected to each other via interconnect 1130 (e.g., a bus). Processor system 1110 may be arranged using antennas and / or connectors for contact and / or contactless communication, respectively.
[0268] For example, in one embodiment, the processor system 1140 (e.g., a training and / or application device) may include processor circuitry and memory circuitry, with the processor arranged to execute software stored in the memory circuitry. For example, the processor circuitry may be an Intel Core i7 processor, an ARM Cortex-R8, etc. In another embodiment, the processor circuitry may be an ARM Cortex-M0. The memory circuitry may be ROM circuitry or non-volatile memory, such as flash memory. Alternatively, the memory circuitry may be volatile memory, such as SRAM. In the latter case, the device may include a non-volatile software interface, such as a hard disk drive, a network interface, etc., arranged to provide the software.
[0269] It should be noted that the embodiments mentioned above are illustrative and not limiting of the invention, and those skilled in the art will be able to devise many alternative embodiments.
[0270] In the claims, any reference marks placed between parentheses should not be construed as limiting the claims. The use of the verb "comprising" and its variations does not exclude the presence of elements or steps other than those recited in the claims. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. Expressions such as "at least one of..." when preceding a list of elements indicate the selection of all elements or any subset of elements from that list. For example, the expression "at least one of A, B, and C" should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all A, B, and C. The invention can be implemented by means of hardware comprising several different elements, and by means of a suitably programmed computer. In a device claim enumerating several components, several of these components can be embodied by the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously.
[0271] In the claims, references enclosed in parentheses refer to reference numerals in the drawings of exemplary embodiments or formulas of embodiments, thereby increasing the comprehensibility of the claims. These references should not be construed as limiting the claims.
Claims
1. A machine-learnable system for training a decoder function, an encoder function, and a normalized stream function, the machine-learnable system being used in conjunction with a sensor data generation system, wherein the normalized stream is a nonlinear conditionally normalized stream function that maps a latent representation to base points in a base space based on conditionally applied data, the system comprising: - A training storage device comprising a set of training target data, said target data including sensor data, said sensor data including image data or LiDAR data. - Processor system, which is configured for - An encoder function that maps target data in the target space to a latent representation in the latent space. - A decoder function that maps latent representations in the latent space to target data representations in the target space, wherein the encoder and decoder functions are arranged to generate parameters that define at least the mean of a probability distribution, and the outputs of the encoder and decoder functions are determined by sampling the defined probability distribution having the generated mean, wherein the probability distribution defined by the encoder function has a predetermined variance. - A normalized stream function, which maps a latent representation to base points in a base space with a predetermined probability distribution. The encoder function, decoder function, and normalized stream function are machine-learnable functions. The normalized stream function is invertible. The processor system is further configured to... - Train an encoder function, a decoder function, and a normalized stream on a set of training target data. The training includes minimizing the reconstruction loss of the cascade of the encoder and decoder functions and minimizing the difference between the predetermined probability distribution in the base space and the cascade of the encoder and normalized stream functions applied to the set of training target data.
2. The machine-learnable system of claim 1, wherein the encoder function and the decoder function define a Gaussian probability distribution.
3. The machine-learnable system of claim 1, wherein the decoder function generates a mean and a variance, the mean and variance being learnable parameters.
4. The machine-learnable system of claim 1, wherein the training includes maximizing a lower bound on evidence, the lower bound on evidence being specific to the training target data. x probability The lower bound of evidence is defined as follows: in, It is a probability distribution and Kulbeck-Leibler divergence, probability distribution Defined by the base distribution and normalized flow.
5. The machine-learnable system as described in claim 1, wherein... - The conditional data in the training pairs includes past trajectory information of traffic participants, and the generated target includes future trajectory information of traffic participants, or - The conditional data in the training pairs includes sensor information, and the generated targets include classification.
6. The machine-learnable system as described in any one of claims 1-5, wherein the conditionally normalized flow function comprises a sequence of multiple reversible normalized flow subfunctions, and one or more parameters of the multiple reversible normalized flow subfunctions are generated by a neural network.
7. The machine-learnable system of claim 6, wherein one or more of the normalized flow subfunctions are nonlinear.
8. A machine-learnable data generation system, the system comprising: - Processor system, which is configured for - The inverse normalized stream function maps fundamental points to latent representations in the latent space, and A decoder function maps latent representations in the latent space to target data representations in the target space. An encoder function is defined corresponding to the decoder function, wherein the encoder and decoder functions are arranged to generate parameters that define at least the mean of a probability distribution. The outputs of the encoder and decoder functions are determined by sampling the defined probability distribution having the generated mean, wherein the probability distribution defined by the encoder function has a predetermined variance. The decoder function and the normalized stream function have been learned according to the machine-learnable system as described in claim 1. - Target data is determined as follows - Obtain the base point in the base space. - Apply the inverse normalized stream function to the fundamental points to obtain the latent representation. - Apply the decoder function to the latent representation to obtain the target data.
9. The machine-learnable data generation system as described in claim 8, wherein... - Base points are sampled from the base space according to a predetermined probability distribution.
10. The machine-learnable data generation system of claim 9, wherein the base points are sampled multiple times from the base space, and wherein at least a portion of the corresponding plurality of target data is averaged.
11. A machine learning method, the method comprising: - Access a stored collection of training target data, including sensor data, such as image data or LiDAR data. - Use encoder functions to map target data in the target space to latent representations in the latent space. - A decoder function is used to map the latent representation in the latent space to the target data representation in the target space, wherein the encoder and decoder functions are arranged to generate parameters that define at least the mean of a probability distribution, and the outputs of the encoder and decoder functions are determined by sampling the defined probability distribution having the generated mean, wherein the probability distribution defined by the encoder function has a predetermined variance. - By using a normalized stream function, the latent representation is mapped to base points in a base space with a predetermined probability distribution. The encoder function, decoder function, and normalized stream function are machine-learnable functions, and the normalized stream function is invertible. The normalized stream function is a nonlinear conditional normalized stream function that maps the latent representation to base points in the base space based on conditionally applied data. - Train an encoder function, a decoder function, and a normalized stream on a set of training target data. The training includes minimizing the reconstruction loss of the cascade of the encoder and decoder functions and minimizing the difference between the predetermined probability distribution in the base space and the cascade of the encoder and normalized stream functions applied to the set of training target data.
12. A method for generating machine-learnable data, the method comprising: - Use inverse normalized stream functions to map fundamental points to latent representations in the latent space. - A decoder function is used to map a latent representation in the latent space to a target data representation in the target space. An encoder function is defined corresponding to the decoder function, wherein the encoder and decoder functions are arranged to generate parameters that define at least the mean of a probability distribution. The outputs of the encoder and decoder functions are determined by sampling the defined probability distribution having the generated mean, wherein the probability distribution defined by the encoder function has a predetermined variance. The decoder function and the normalized stream function have been learned according to the machine-learnable system as described in claim 1. - Target data is determined as follows - Obtain the base points in the base space with a predetermined probability distribution. - Apply the inverse normalized stream function to the fundamental points to obtain the latent representation. - Apply the decoder function to the latent representation to obtain the target data.
13. A temporary or non-temporary computer-readable medium comprising data representing instructions that, when executed by a processor system, cause the processor system to perform the method according to claim 11 or 12.