Machine Learning Systems
By introducing two latent spaces into the neural network, the output is indirectly constructed in the second latent space, the problem of difficulty in modeling uncertainty in existing neural networks when processing complex and high-dimensional data is solved, and more effective training and reasoning is achieved.
Patent Information
- Application Number
- CN202010534293.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-14
- Filing Date
- 2020-06-12
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2040-06-12
AI Technical Summary
Existing neural networks are difficult to effectively model uncertainty when processing complex and high-dimensional data, and the training and inference cost is high.
By introducing two latent spaces, the first latent space is used to determine the reference points associated with the input data instance, thereby indirectly constructing the output in the second latent space. This approach avoids the assumption of a priori distribution directly on global parameters and reduces computational complexity.
This method can learn more efficiently the distribution structure of the input data, reduce the cost of training and inference, and provide better modeling of uncertainty.
Smart Images

Figure CN112084836B_ABST
Abstract
Description
Technical Field
[0001] The presently disclosed subject matter relates to machine learning systems, autonomous device controllers, machine learning methods, and computer-readable media. Background Art
[0002] Neural networks are a ubiquitous paradigm for approximating almost any kind of function. Their highly flexible parameter form combined with large amounts of data allow accurate modeling of the underlying task, a fact that often leads to state-of-the-art prediction performance. While predictive performance is certainly an important aspect, in many safety-critical applications such as self-driving cars, one also desires accurate uncertainty estimates about the predictions.
[0003] Bayesian neural networks have been attempts to give neural networks the ability to model uncertainty; they assume a prior distribution over the network weights, and by reasoning, they can represent its uncertainty in the posterior distribution. However, for such complex models, the choice of prior is quite difficult, as understanding the interaction of parameters with the data is a nontrivial task. As a result, priors are often adopted for computational convenience and tractability. Furthermore, due to the high dimensionality and complexity of the posterior, reasoning over the weights of a neural network can be a daunting task.
[0004] An alternative way to "bypass" the aforementioned problems is to adopt random processes. They directly assume a distribution over functions (e.g., neural networks) without the necessity of taking a prior distribution over global parameters (such as neural network weights). Gaussian processes (GPs) are examples of random processes. Unfortunately, Gaussian processes have two limitations. For example, the underlying model is not very flexible for high-dimensional problems, and the training and inference costs are quite high, as they scale cubically with the size of the dataset. Summary of the invention
[0005] To solve these and other problems, a machine learning system as defined in the claims is proposed. The machine learning system is configured to map input data instances to outputs according to a system mapping.
[0006] In an embodiment, an output is derived from a second potential input vector in a second latent space by applying a third function to the second potential input vector. A second potential input vector may be determined for the input data instance from second potential reference vectors obtained for a plurality of reference data instances, the plurality of reference data instances being identified as parents of the input data instance. Identifying the parent reference data instance may be based on a similarity between a first potential input vector obtained for the input data instance and a first potential reference vector obtained for a set of reference data instances.
[0007] Interestingly, the output is derived from a second latent vector in a second latent space. The input instance is not directly mapped to the second latent vector; instead, it is constructed from a reference instance related to it. The first latent space can be used to determine this relationship. This indirect process has the advantage that the system is forced to learn the structure of the input distribution very well.
[0008] Preferably, the various mappings are random rather than deterministic. If necessary, the mappings can be repeated and averaged to improve accuracy. Various machine learning functions can be applied to the first, second and third functions. However, neural networks have been shown to work particularly well.
[0009] Machine learning systems can be used for the control of physical systems. For example, the input data instances and reference data instances can include sensor data, particularly images. The output can include output labels; for example, the system mapping can be a classification task. Other applications are possible. For example, the system can be trained to predict variables that are difficult to measure directly, for example, the output can include physical quantities of the physical system, such as temperature, pressure, etc.
[0010] In an embodiment, the second potential input vector ( ) further depends on the corresponding reference label. For example, for and reference mark Mapped to the second latent vector A function can be called a function For example, one can use ,in is from the reference tag to Alternatively, except In addition to or instead of , one can also model and train . For example, the second function may be configured to take as input a reference data instance and a corresponding reference label to produce a second latent reference vector. This embodiment exploits the fact that two latent spaces may be used. In embodiments where a direct mapping into the second latent space is not used for unseen input data instances, their predicted values may be improved by including information about the corresponding labels directly in the second function. In embodiments where a direct mapping into the second latent space is used for unseen input data instances, modeling consistency may be improved by reusing the same second function for reference points and for training and unseen points. The latter may be achieved, for example, by introducing a new mapping as above to complete.
[0011] In an embodiment, the third function may be configured to take as input the second potential input vector and the first potential input vector to produce an output.Including the first potential input vector increases the exploration capability of the system.
[0012] The relationships between training points and reference points and / or among reference points can be implemented as learning a dependency graph over the latent representations of points in a given dataset. In doing so, they define a Bayesian model without explicitly assuming a prior distribution over the underlying global parameters; they instead take a prior over the relational structure of the given dataset, which is a much simpler task. The model is scalable to large datasets.
[0013] Machine learning systems can be used in autonomous device controllers. For example, machine learning systems can be used to classify objects in the vicinity of an autonomous device.
[0014] For example, a controller may use a neural network to classify objects in sensor data and use the classification to generate control signals for controlling an autonomous device. The controller may be part of an autonomous device. The autonomous device may include a sensor configured to generate sensor data that may be used as an input to a neural network.
[0015] The machine learning system is an electronic system. The system may be included in another physical device, such as a technical system, for controlling the physical device, such as its movement.
[0016] One aspect of the presently disclosed subject matter is a machine learning approach.
[0017] Embodiments of the method may be implemented on a computer as a computer-implemented method, or implemented in dedicated hardware, or in a combination of the two. Executable code for embodiments of the method may be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product may include non-transitory program code stored on a computer-readable medium for performing embodiments of the method when the program product is executed on a computer.
[0018] In an embodiment, the computer program comprises computer program code, which is suitable for executing all or part of the steps of the embodiment of the method when the computer program is run on a computer. Preferably, the computer program can be embodied on a computer readable medium.
[0019] Another aspect of the presently disclosed subject matter provides a method of making a computer program available for download. This aspect is used when the computer program is uploaded to, for example, Apple's App Store, Google's Play Store, or Microsoft's Windows Store, and when the computer program is available for download from such a store. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Additional details, aspects, and embodiments of the presently disclosed subject matter will be described by way of example with reference to the accompanying drawings. The elements in the various figures are illustrated for simplicity and clarity and are not necessarily drawn to scale. In the various figures, elements corresponding to elements already described may have the same reference numerals. In the accompanying drawings,
[0021] Figure 1 schematically illustrates an example of a Venn diagram of sets related to an embodiment,
[0022] Figure 2 Schematically showing examples of spaces and figures related to the embodiments,
[0023] Figure 3a schematically illustrating an example of an embodiment of a machine learning system configured for training a machine learnable function,
[0024] Figure 3b schematically illustrates an example of an embodiment of a machine learning system configured for applying a machine learnable function,
[0025] Figure 4 Schematically illustrates examples of various data in embodiments of machine learning systems and / or methods,
[0026] Figure 5 schematically shows an example of a diagram related to an embodiment,
[0027] Figure 6a-6c Schematically showing an example of the predicted distribution for a regression task for a baseline method,
[0028] Figure 6d-6e Schematically illustrates an example of a predicted distribution for an embodiment,
[0029] Figure 7 An example of an embodiment of a flow chart for an embodiment of a machine learning method is schematically shown.
[0030] Figure 8a schematically shows a computer-readable medium having writable means comprising a computer program according to an embodiment,
[0031] Figure 8b A representation of a processor system according to an embodiment is schematically shown. DETAILED DESCRIPTION
[0032] While the presently disclosed subject matter is susceptible of embodiment in many different forms, there are one or more specific embodiments shown in the drawings and will be described in detail herein, with the understanding that this disclosure is considered an example of the principles of the presently disclosed subject matter and is not intended to limit the presently disclosed subject matter to the specific embodiments shown and described.
[0033] In the following, for the sake of understanding, the elements of the embodiments are described in operation. However, it will be clear that the corresponding elements are arranged to perform the functions described as being performed by them.
[0034] Furthermore, the presently disclosed subject matter is not limited to the embodiments and the presently disclosed subject matter lies in each and every novel feature or combination of features described herein or recited in mutually different dependent claims.
[0035] Prediction tasks frequently occur in computer vision. For example, autonomous device control, such as for autonomous cars, depends on decisions based on reliable classifications. One problem with predictions made by machine learning systems is that there is often a small difference between predictions for situations in which the system is well trained and predictions for situations in which the system is not well trained. For example, consider a neural network trained to classify road signs. If the network is presented with new, previously unseen road signs, the neural network will most likely make a confident and possibly correct classification. However, if an image outside the distribution of images used for training, such as an image of a cat, is presented to the neural network, the known neural network tends to still confidently predict the road sign for the image. This is undesirable behavior, so there is a need for a machine learning system that fails gently when such an image (which is called an "out of distribution" ood image) is presented to the machine learning system. For example, in an embodiment, for points outside the data distribution, the prediction can fall back to a certain default value or be close to a certain default value.
[0036] Other advantages include the possibility of providing the model with prior knowledge, such as specifying an inductive bias. For example, nearby instances may be considered more likely to be the parents of a given instance. This is much more intuitive than specifying a prior distribution over a global latent variable. The encoding of the neighborhood may depend on the specific type of input data. For example, for images, a distance metric such as the L2 metric may be used. If the input data includes other suitable data, other distance metrics may be included. For example, in an embodiment, the input data includes sensor data from multiple sensors. For temperature, a temperature difference may be used, for pressure, a pressure difference may be used, and so on. However, the temperature and pressure data may be appropriately scaled to take into account the different scales at which they appear in the modeled system. Multiple distance metrics may be combined into a single distance, for example, by a possibly weighted sum of squares.
[0037] Another advantage is that the model is forced to learn the dependency structure between data instances. Furthermore, such dependencies can be visualized. Such visualization allows debugging of the model. Optionally, embodiments can compute and visualize these dependencies. If in a so-called G graph (an example of which is in Figure 5 ) or the clusters shown in Figure A (these figures are further explained below) do not correspond to reality, the modeler can use this information to further control the machine learning system, for example by increasing or decreasing the number of reference points or by changing the prior knowledge. A further advantage is that in embodiments, an appropriate Bayesian model can be obtained, for example, which is exchangeable and consistent under marginalization. This allows probabilistic modeling and inference of latent information without having to specify a prior distribution over the global latent variables.
[0038] A new machine learning system is provided, which can be used to train and / or apply machine learnable functions. In an embodiment, two different latent spaces, such as multidimensional vector spaces, are used. The elements in these spaces are respectively referred to as first latent vectors and second latent vectors. Vectors can also be referred to as points in vector space. Outputs such as predictions or classifications are performed based on the second latent vectors. A special set of reference data instances that can be extracted from the same distribution as the input data instances are mapped into the second latent space, which are second latent reference vectors.
[0039] For training instances that are different from the reference instances, or for unseen input data instances, an indirect scheme is used. Both the training data instances and the reference instances are mapped to a first latent space. Thus, the reference data instance corresponds to a vector in the first latent space, e.g., by a first function, and corresponds to a vector in the second latent space, e.g., by a second function.
[0040] In the first latent space, reference points are selected that are relevant to the training input instances, e.g., they are similar to the training input instances. Points in the second latent space are then derived from the corresponding points in the second latent space, e.g., averaging them, possibly with a weighted average. Finally, the output is computed from the determined points in the second latent space. Thus, the system computes the output indirectly from the input by selecting relevant so-called parents. This forces the system to develop a better understanding of the structure of the input distribution. This contributes to a better understanding of the input distribution and to a better ability to recognize unseen inputs that are poorly reflected in the input data.
[0041] Note that one or more or all of the mappings can be stochastic. Rather than deterministically determining the output, a probability distribution is determined and then the mapping results are sampled from the probability distribution. For example, this can be done for the mapping from the input instance to the first latent space, for the mapping of the reference instance to the second latent space, for any of the mappings used to determine the parent, and for the mapping from the second latent space to the output.
[0042] Figure 1 An example of a Venn diagram of sets related to an embodiment is schematically shown.
[0043] Figure 1 The training input is shown in . For supervised or partially supervised learning, all or part of the training inputs can be accompanied by prediction targets. For simplicity, we refer to the prediction targets as labels. The prediction targets can also be used to predict variables such as the temperature of the system, or to predict multiple labels, such as for classifying inputs, etc. The training inputs can also be unsupervised, for example, in an embodiment, the machine learning system can be trained as an autoencoder, for example, to compress the input or use it as input to a further learning system. For simplicity, it will be assumed that the training inputs are provided with labels, but this is not required.
[0044] Figure 1 Also shown is the reference set . Reference Set can be drawn from the same distribution as the training input. It can also be associated with the prediction target in whole or in part. We will refer to M as The training points in are called training points and The white background corresponds to , which is The complement of .
[0045] Figure 3aAn example of an embodiment of a machine learning system 300 configured for training a machine learnable function is schematically shown. Figure 4 Examples of various data in embodiments of machine learning systems and / or methods are schematically shown. For example, the machine learning system 300 may be configured to utilize Figure 4 The illustrated embodiment.
[0046] Figure 4 An exemplary embodiment is shown with the understanding that many variations are possible, for example, as described herein. Figure 4 The task performed by the embodiment of is to train the system to classify the input. However, this embodiment can be adapted for unsupervised learning. For example, for unsupervised learning, the loss function can be determined from another source, for example, by training the system as an autoencoder, or in a GAN framework, etc. Note that Figure 4 A subset of can be used to apply the system to unseen inputs.
[0047] exist Figure 4 Shown at 412 is the observed input: And the corresponding tags For example, the observed input may include sensor data, such as image data. For example, the label may classify the content of the image data. For example, the image may represent a road sign, and the label may represent the type of road sign.
[0048] refer to Figure 4 The explained system learns to map input data instances to outputs. The outputs may include labels. The outputs may include vectors whose elements give probabilities for multiple labels. For example, where a single label prediction is desired, the vectors may sum to one. In general, input data instances may be referred to as "points," meaning points in the space of all possible input data instances.
[0049] In this embodiment, the observed input is split 413 into two sets: a training set 422 and a reference set 424. For example, the reference point: For example, training point 422 may be: The splitting can be done randomly, for example, a random part can be selected as a reference point. Figure 1 As shown in , more general settings are possible, however, for simplicity, the above situation will be assumed here.
[0050] The training data may include any sensor data, including one or more of the following: video, radar, LiDAR, ultrasound, motion, etc. The resulting system can be used to calculate control signals for controlling a physical system. For example, the control signals can be used for computer-controlled machines such as robots, vehicles, household appliances, power tools, manufacturing machines, personal assistants, or access control systems. For example, a decision can be made based on the classification results. For example, in addition to outputs such as classification results, or instead of outputs such as classification results, the system can generate a confidence value. A confidence value such as predicted entropy expresses whether the model believes that the input is likely to come from a specific class, which is indicated by low entropy, for example; or whether the model is uncertain, such as it believes that all classes are equally likely, which is indicated by high entropy, for example.
[0051] The confidence value can be used in controlling the physical system. For example, if the confidence is low, the control can switch to a more conservative mode, such as reducing speed, reducing maneuverability, engaging human control, stopping the vehicle, etc.
[0052] A machine learning system trained for classification can be included in a system for conveying information, such as a supervisory system or an imaging system, such as a medical imaging system. For example, a machine learning system can be configured to classify sensor data, such as to grant access to a system, such as through facial recognition. Interestingly, the machine learning system takes into account ambiguity in the observed data due to a better indication of confidence.
[0053] The system 300 can communicate, for example, for receiving input, training data, or for transmitting output, and / or for communicating with an external storage device or input device or output device. This can be done through a computer network. The computer network can be the Internet, an intranet, a LAN, a WLAN, etc. The computer network can be the Internet. The system includes a connection interface that is arranged to communicate within or outside the system as required. For example, the connection interface can include a connector, such as a wired connector such as an Ethernet connector, an optical connector, etc., or a wireless connector such as an antenna, such as a Wi-Fi, 4G or 5G antenna, etc.
[0054] Execution of system 300 is implemented in a processor system, such as one or more processor circuits, examples of which are shown herein. Figure 3a and Figure 3b The functional units shown may be functional units of a processor system. For example, Figure 3a and Figure 3b Can be used as a blueprint for possible functional organization of a processor system. The processor circuit(s) are not shown separately from the units in these figures. For example, Figure 3a and Figure 3b The functional units shown in the figure may be implemented in whole or in part in computer instructions stored at the system 300, for example, in an electronic memory of the system 300, and executable by a microprocessor of the system 300. In a hybrid embodiment, the functional units are implemented partially in hardware, for example as a coprocessor such as a neural network coprocessor, and partially in software stored and executed on the system 300. Parameters of the machine learnable function, such as the neural network and / or the training data may be stored locally at the system 300, or may be stored in cloud storage.
[0055] System 300 may be distributed across multiple devices, such as multiple devices for storage and / or computing. System 300 may be implemented in a single device. For example, a device may include a processor system for computing and may include local storage or an interface for external storage such as cloud storage.
[0056] For example, machine learning system 300 may include training storage 310. Training storage 310 may store training input data instances, or "training points," and corresponding labels. Figure 3a Input data instances 312 and corresponding input labels 314 are shown. For example, the machine learning system 300 may include a reference storage 320. For example, the reference storage 320 may store reference data instances or reference points. Typically, the training data is not needed after training, but reference information, such as reference points or information derived from the reference points, may be needed.
[0057] Reference storage 320 may store reference labels for all, some, or none of the reference portions. For simplicity, it is assumed that labels are available for all reference inputs, but this is not required. For example, Figure 3a Reference data instances 322 and corresponding reference tags 324 are shown.
[0058] Typically, the split of the inputs observed in the training and reference data only needs to be done once. The number of reference points depends on the distribution and can be found empirically. In general, it is desirable that the reference points reflect the distribution of the training inputs well. As a starting point, the number of reference points can be chosen to be a multiple of the number of labels being predicted, such as 20 or 30 times more, etc. Different distributions may use more or fewer reference points to obtain optimal performance.
[0059] To train the system, the training input 422 and the reference point 424 are passed through a first function 431 is mapped to the latent space As a result, one obtains a latent space Points in 432 and Space Potential points in 434. Point 434 is obtained from a reference data instance, while point 432 is obtained from a training data instance.
[0060] In an embodiment, the function 431 is a machine learnable function, such as a neural network. For example, an image can be mapped to a vector, which can have a smaller dimension. Latent space It can be a multi-dimensional vector space.
[0061] From the input space To the latent space The mapping in Figure 2 In the leftmost picture, a space including both training points and reference points is shown. In the middle picture, these two types of points are mapped into the latent space .
[0062] For example, the system 300 may include a first function 330, such as a first function unit, etc., which is configured to apply a first function Unit 330 may, for example, implement a neural network For example, the first function 330 may be applied to training inputs, such as input data instances, and reference inputs, such as reference data instances. For example, these data instances may include images.
[0063] space The latent vector in is called the first latent vector. Interestingly, the first latent vector can be used to determine a reference point that can be considered as a parent. Determining the parent can be based on the similarity between the first latent vectors. Interestingly, determining the parent can be interpreted as determining a directed graph. In the case where the first latent vector is obtained from the training input, a bipartite graph is obtained as follows: Bipartite graph: 442. (Bipartite graph Not to be confused with the first function 431). The edges in the bipartite graph indicate which elements of R are considered to be parents of which elements of M. For example, from Elements in arrive Elements in The directed edge indicates yes The parent of . Figure 4 Two arrows toward graph 442 are shown in , because the bipartite graph depends on the first latent vector obtained from M and the first latent vector obtained from R.
[0064] The first latent vector 434 obtained from 424 may also have a parent. However, in this case, the first latent vector 434 corresponds to a reference data instance, and the parent also corresponds to a reference data instance. Therefore, the induced graph is a graph of R. We find that when care is taken to ensure that the graph is a directed acyclic graph (DAG), an improved system is obtained.
[0065] Figure 2 The rightmost image shows the bipartite graph between the training points and the reference points. and the DAG graph between the reference point .
[0066] To create a graph from the first latent vector, edges may be drawn using a similarity function. In an embodiment, the similarity function generates a probability distribution for the edges of the graph, such as a Bernoulli distribution for each edge. In particular, a conditional Bernoulli distribution may be used. For example, the distribution over the graph may be given by a uniformly distributed 1D variable U and a conditional Bernoulli distribution E|U for the edges given U. The similarity function may be a kernel function. Thus, the graph and Figure And thus the parent for a given first latent vector can be sampled from the distribution, rather than being determined deterministically. In an embodiment, FIG. and Figure One or both of may have an edge that is determined deterministically; for example, an edge may be assumed if the similarity function exceeds a value.
[0067] For example, given a first latent vector, the parent identifier 350 can be configured to select, for example, sample from R (e.g., from a reference point). For example, the parent identifier 350 can be configured to select, for example, sample from R (e.g., from a reference point) based on the first latent input vector and the first function ( ) is obtained from the reference data instance set, and the reference data instance set ( ) are identified as parents of the input data instance. Therefore, the reference data instance may correspond to the first latent vector, the second latent vector, the graph Points and graphs in The training data instance can correspond to the first latent vector and the graph A second latent vector can be calculated for each training data instance.
[0068] During training, a first latent vector for R and M elements may be generated by applying the first function 330. After training, the first latent vector elements of R may be pre-computed and stored. For example, the parent identifier 350 may be configured to input a latent vector to calculate a similarity with a first latent vector obtained for a reference input data instance, and to probabilistically select a parent based on the similarity. The parent identifier 350 may be configured to use the information about the parent in the case where the input itself is an input to R. The first latent vector of It is DAG.
[0069] In practice, training may be performed in batches. In an embodiment, a batch may include all reference points and some training points. Graph G may be fully computed, e.g., sampled, and graph A may be partially computed, e.g., sampled, e.g., induced so far by the training points in the batch. If the reference set is too large, one may also use mini-batches for the reference set itself. Using the full set of reference points has the advantage of avoiding biased gradients, but batch training of R points allows for a large reference set, which may be used, for example, to train a larger model.
[0070] For example, when used for training on images, data augmentation can be used. For example, random geometric transformations such as translation, rotation, scaling, etc. can be applied to prevent overfitting. In an embodiment, due to the random transformations, the reference points in each batch are slightly different.
[0071] The reference point 424 may be, for example, represented by a second function 435 is mapped to the second latent space 436. Second Function 435 It can also be a machine-learnable function. The second function 435 can be implemented in the second function 340 of the system 300, such as a second function unit. In particular, the second function It could be a neural network.
[0072] To train the system, the system mapping may be applied to the specific point 452. Mapping the specific point 452 to an output according to the system mapping may include determining that the point 452 is at If point 452 is a training point, this can be done using a bipartite graph , if point 452 is points in the graph, this can be Point 452 is mapped to a point in the second latent space through its parent.
[0073] against The second latent space for a specific point in The elements of can be calculated from the reference point of their parent For example, we can take the average of the second latent vectors of the parent generation. For training, we can Even if an element of R already has a The elements of Z can also be calculated by averaging the second latent vectors of the parents, just as we can for In this way, we obtain Vector in Alternatively, computing a vector in the second latent space for a point in R may be done by averaging the parents from a graph G including the second latent vector for the point itself.
[0074] In an embodiment, the second function is also applied to the training point to directly obtain a point in the second latent space. Then, the directly mapped point can be updated using the second latent vector of its parent. For example, it can be used to average the second latent vector of its parent. The weighting factor can adjust the weight of the parent relative to the directly mapped vector. For example, a weighting factor of 1 can only use the second latent vector of the parent, while a weighting factor of 0 can only use the second latent vector obtained directly from the input instance. In an embodiment, the weighting factor is greater than 0, or at least greater than 0.1, at least greater than 0.5, and so on.
[0075] Interestingly, z is obtained indirectly because the first latent space is used to identify relevant reference points, i.e., parents, which are then used to determine vectors to represent points of M with points in Z. Thus, the system is forced to learn the structure of the reference points. For example, system 300 may include a latent vector determiner 360 configured to determine a second latent vector for a point given a second latent vector of a parent of the point. Note that from The dimensions are mapped to Output dimensions, where is the number of all possible classes. and Accordingly The dimensions are mapped to and Dimensions. Increase or decrease and The dimensions lead to different model behaviors.
[0076] Finally, the third function 437 Can be applied to vectors To generate output. The third function It can be implemented in the third function 370, for example, a third function unit. The third function can also be a neural network. In the implementation, the first, second and third function units can reuse a large amount of code, such as a neural network implementation code. It is not necessary for all first, second and third functions to be neural networks. For example, the first function can be mapped to a feature vector. The latter mapping can be handmade. In an embodiment, the first, second and third functions can be any learnable classification algorithm, such as a random forest, a support vector machine, etc. However, neural networks have been proven to be universal. In addition, using neural networks for all three functions unifies the training on the three functions. All three functions can be deep networks, such as convolutional networks, etc. However, it is not necessary for all networks to have equal sizes. For example, in an embodiment, the function The neural network can have a and Few layers.
[0077] Finally, for example, from The predicted label 464, such as the output of unit 370, can be used to calculate a training signal 466 for a trainable portion of the system 300. For example, the system 300 may include a training unit 380 configured to train the system 300. For example, after mapping, the training input data instance is mapped to the output according to the system mapping, and the training unit 380 can be configured to derive the training signal from the error derived from the output and adjust the machine learnable function according to the training signal. The learning algorithm can be specific to the machine learnable function, for example, the learning algorithm can be back propagation. The error can include the difference between the output and the predicted target, or can be derived from the difference between the output and the predicted target. The error can also be obtained by further calculation, for example, for an autoencoder or a recurrent-GAN, for example, a further learnable function can be used to reconstruct all or part of the input data instance. The error can be calculated based on the difference between the reconstructed input data instance and the actual input data instance.
[0078] In an embodiment, the first function ( ) and the second function ( ) and at least one of the third functions ( ) are machine learnable functions. In an embodiment, all three functions , and are all machine learnable functions. For example, they could all three be neural networks.
[0079] Note that training can also be done on reference points, although this is not required. In fact, the reference points do not need to have labels for the system to work, although improved predictions have been found in experiments when labels are provided for the reference points.
[0080] It is found to be advantageous if the system is stochastic, for example, the parent can be determined by drawing from a distribution, such as a Bernoulli distribution. For example, the first function ( ), the second function ( ) and / or a third function ( ) can be a probability distribution from which the corresponding output is sampled.
[0081] There are a number of variations that can be applied in the system 300. For example, to derive the second latent vector, the second function B can use the reference labels as input in addition to the reference points themselves. Figure 3a Indicated by a dashed line from label 324 to function 340 in Figure 4 426 to the latent space 436. This variation is possible because two latent spaces are used. For unknown inputs, only a mapping to the first latent space is required, and thus the second latent space can use the labels in the mapping. Having the labels as inputs in the second latent space increases the predictive power of the second latent space.
[0082] To include label 426 when mapping to the second latent vector, one can use a machine learnable function, such as a neural network, which takes as input both the reference data instance and the label. In a variant, two machine learnable functions are used, such as two neural networks. The first machine learnable function can map the reference data instance to the second latent space, and the second machine learnable function can map the label to the second latent space; then, possibly with weighting applied, the two second latent vectors can be added or averaged. The weighting factor can be a learnable parameter.
[0083] Another variation that can be used with or instead of the previous variation is to provide the first potential vector of point 452 as an additional input to the third function 370. This option is Figure 3a 370. It is shown as a dotted line from the first function 330 to the third function 370. It is found that adding this additional vector increases the extrapolation capability of the system. On the other hand, without this additional vector, the system recovers low confidence predictions more strongly when equipped with the ood input. Depending on the situation, both can be preferred.
[0084] An advantage of an embodiment is that the system reverts to a default output or close to a default output when faced with an ood input. For example, given an ood input, the input may have no parent and therefore be mapped to a default vector in the second latent space. The latter is then mapped to a default output or close to a default output. Measuring how close the output will be to the default output can be used to determine confidence in the output.
[0085] For example, in an embodiment, the output of unit 370 may include multiple output probabilities for multiple labels. For example, if the input is a road sign, the output of unit 370 may be a vector indicating the probability for each possible road sign that can be recognized by the system. If a particular output is recognized with confidence, such as a particular road sign, one probability will be high and the remaining probabilities will be low, such as close to zero. If no particular output is recognized with confidence, many or all probabilities will be approximately equal. Therefore, the confidence in the output can be derived by measuring how close the output is to consistency - the closer to consistency, the less confident, and the further away from consistency, the more confident. The confidence that the system has in its output can be reflected in a variety of ways; one way to do this is to calculate the entropy of the generated probability.
[0086] In an embodiment, a confidence value may be derived from a unit 370 that does not use the first latent vector as input. The confidence value indicates how well the input is represented among the training inputs. However, the output itself may be taken from a unit 370 that uses the first latent vector as input, so that the output has better extrapolation.
[0087] Figure 3b An example of an embodiment of a machine learning system configured for applying a machine learnable function is schematically illustrated.
[0088] After the system has been trained, the system can be applied to new inputs, such as input data instances that were not seen during training. The machine learning system after training may be smaller than the system used for training. Figure 3b An example of a reduction system applying system mapping is shown in FIG.
[0089] Figure 3b An input interface 332 is shown which is arranged to obtain an input data instance. For example, the first function in the first function unit 330 can be applied to an input data instance to obtain a first latent input vector in a first latent space.
[0090] For example, Figure 3a As shown in , the reference storage device 320 may include reference data instances and corresponding labels. From these first and second potential reference vectors, the first and second potential reference vectors may be determined in the same manner as in training. The advantage of this scheme is that the first and second potential reference vectors may be appropriately sampled. For example, Figure 3a The storage device 320 can also be used as Figure 3b320 in the reference storage device. Alternative solutions are possible, although some accuracy may be lost, especially in the case where the application system mapping is repeated multiple times to increase accuracy. For example, the reference storage device 320 may include a first potential reference vector and a second potential reference vector that are pre-calculated using the trained first and second functions. Shown are a first potential reference vector 326 and a second potential reference vector 328.
[0091] Using the first latent reference vector, the bipartite graph corresponding to the first latent input vector For example, . For example, a similarity function such as a kernel function can determine the probability that an edge between the first potential reference vector and the reference point exists; the edge can then be sampled accordingly. In this way, a parent set of the new first potential input instance can be obtained. For example, this can be done by the parent identifier 350.
[0092] In the case of having a parent, the latent vector determiner 360 and the third function 370 may determine the latent vector in the second latent space in the same manner as may be done during training, for example by averaging the second latent reference vectors corresponding to the parent and applying the third function 370 to the average. .
[0093] As in training, the first potential input vector, its parents, and the output vector can be determined by sampling the probability distribution. To increase accuracy, this can be done multiple times and the results can be averaged.
[0094] A trained machine learning system, such as a neural network system, can be applied in an autonomous device controller. For example, the input data to the system can include sensor data of an autonomous device. An autonomous device can perform movement at least partially autonomously, for example, modifying the movement according to the environment of the device without user-specified modifications. For example, the system can be a computer-controlled machine such as a car, a robot, a vehicle, a household appliance, a power tool, a manufacturing machine, etc. For example, the system can be configured to classify objects in sensor data. An autonomous device can be configured for decisions that depend on classification. For example, the system can classify objects in the environment surrounding the device, and, for example, if other traffic is classified near the device - for example, people, cyclists, cars, etc., the system can stop, or slow down, or turn, or otherwise modify the movement of the device.
[0095] In various embodiments of the system 300, the communication interface may be selected from a variety of alternatives. For example, the interface may be a network interface to a local area network or a wide area network (such as the Internet), a storage interface to an internal or external data storage device, a keyboard, an application program interface (API), etc.
[0096] The system 300 may have a user interface, which may include well-known elements such as one or more buttons, a keyboard, a display, a touch screen, etc. The user interface may be arranged to accommodate user interaction in order to configure the system, train on a training set, or apply to new sensor data.
[0097] The storage device may be implemented as an electronic memory, such as a flash memory, or a magnetic memory, such as a hard disk, etc., such as storage devices 310, 320 or a storage device for learnable parameters, etc. The storage device may include a plurality of discrete memories that together constitute the storage device. The storage device may include a temporary memory such as RAM. The storage device may be a cloud storage.
[0098] The system 300 may be implemented in a single device. Typically, the system 300 includes a microprocessor that executes appropriate software stored at the system; for example, the software may have been downloaded and / or stored in a corresponding memory, such as a volatile memory such as RAM or a non-volatile memory such as flash memory. Alternatively, the system may be implemented in whole or in part in programmable logic, for example as a field programmable gate array (FPGA). The system may be implemented in whole or in part as a so-called application specific integrated circuit (ASIC), such as an integrated circuit (IC) customized for its specific use. For example, the circuit may be implemented in CMOS, for example using a hardware description language such as Verilog, VHDL, etc. In particular, the system 300 may include circuits for neural network evaluation.
[0099] The processor circuit may be implemented in a distributed manner, for example as a plurality of sub-processor circuits. The storage device may be distributed over a plurality of distributed sub-storage devices. Some or all of the memory may be electronic memory, magnetic memory, etc. For example, the storage device may have volatile and non-volatile portions. Portions of the storage device may be read-only.
[0100] Figure 7 An example of an embodiment of a flow chart for an embodiment of a machine learning method 700 is schematically shown. The machine learning method (700) is configured to map an input data instance to an output according to a system mapping, the system mapping consisting of a plurality of functions. The machine learning method comprises
[0101] - obtaining (710) an input data instance,
[0102] -Store (720) a reference data instance collection ( ),
[0103] -According to the first function ( ) maps ( 730 ) the input data instance to a first latent space ( ),
[0104] - based on the first potential input vector and the first function ( ) is obtained from the reference data instance set, and the reference data instance set ( ) are identified as parents of the input data instance.
[0105] -From the second function ( ) obtains a second latent reference vector, determining (750) a second latent space ( ) in the second potential input vector ( ), the second potential reference vector is from the plurality of reference data instances identified as parents,
[0106] - By passing the third function ( ) is applied to the second latent input vector, from the second latent input vector ( ) determines the output of (760), where the first function ( ) and the second function ( ) and at least one of the third functions ( ) is a machine learnable function.
[0107] For example, accessing training data and / or receiving input data may be accomplished using a communication interface, such as an electronic interface, a network interface, a memory interface, etc. For example, storing or retrieving parameters may be accomplished from an electronic storage device, such as a memory, a hard drive, etc.
[0108] For example, machine learning systems and / or methods can be configured to apply and / or train one or more neural networks. For example, applying a sequence of neural network layers to data of training data and / or adjusting stored parameters to train the network can be accomplished using an electronic computing device such as a computer.
[0109] During training and / or during application, e.g. for the function , or The neural network may have a plurality of layers, which may include one or more projection layers, and one or more non-projection layers such as convolutional layers, etc. For example, the neural network may have at least 2, 5, 10, 15, 20, or 40 hidden layers, or more, etc. The number of neurons in the neural network may be, for example, at least 10, 100, 1000, 10000, 100000, 1000000, or more, etc.
[0110] As will be clear to those skilled in the art, many different ways of performing the method are possible. For example, the order of the steps may be performed in the order shown, but the order of the steps may vary or some steps may be performed in parallel. In addition, other method steps may be inserted between the steps. The inserted steps may represent a refinement of the method such as described herein, or may be unrelated to the method. For example, some steps may be performed at least partially in parallel. In addition, a given step may not be fully completed before the next step begins.
[0111] Embodiments of the method may be performed using software that includes instructions for causing a processor system to perform method 700. The software may include only those steps taken by a specific sub-entity of the system. The software may be stored in a suitable storage medium such as a hard disk, a floppy disk, a memory, an optical disk, etc. The software may be sent as a signal by wire or wireless or using a data network (e.g., the Internet). The software may be made available for downloading on a server and / or remote use. Embodiments of the method may be performed using a bitstream that is arranged to configure a programmable logic, such as a field programmable gate array (FPGA), to perform the method.
[0112] Additional details, embodiments, and variations are set forth below. Embodiments of machine learning systems and / or methods may implement embodiments of machine learning processes, which are referred to herein as functional neural processes (FNPs). In the following examples, a supervised learning setting is assumed, in which a tuple of a given point is given. ,in are the input covariates and is a given label. Let yes A sequence of observed data points. This provides an example of partially supervised learning. FNP can be adapted to fully unsupervised learning.
[0113] At a high level, FNP can be achieved by first select a set of reference points and then use the function around those points Based on the probability distribution above, we assume that arrive Function More specifically, let is such a reference set, and let is another set, such as not in Now let Is from Any finite random set of , for example, that constitutes the observed input. To facilitate the explanation, two more sets are introduced; , which contains of points, and , which contains and All points in . Figure 1 A Venn diagram is provided in . A possible construction of the model is shown below, which consists of Figure 2 It can be shown that it is possible to construct an embodiment corresponding to an infinite exchangeable random process.
[0114] The first step of FNP can be to Each of Independently embedded into latent representation
[0115] (3)
[0116] in can be any distribution, such as a Gaussian or delta-peaked distribution, where its parameters such as mean and variance are given by This function can be any function as long as it is flexible enough to provide For this reason, one can employ neural networks, as their representational power has been demonstrated on a variety of complex high-dimensional tasks such as natural image generation and classification.
[0117] The next step could be to constructs a dependency graph between the points in For example, in GP, such a correlation structure can be expressed as a kernel function that measures the similarity between two inputs. is encoded in the covariance matrix. In FNP, a different approach is taken. Given the latent embedding obtained in the previous step , one can construct Two directed graphs of dependencies between points in ; Directed Acyclic Graph (DAG) between the points in and from arrive The bipartite graph of . These graphs can be represented as random binary adjacency matrices, where for example Corresponding to the vertex is the vertex The distribution of a bipartite graph can be defined as
[0118] (4)
[0119] in Provides a point Depends on reference set midpoint Note that the embedding of each node can preferably be a vector rather than a scalar, and furthermore, The prior distribution above can be represented by the initial vertex The latter allows one to maintain sufficient information about the vertices and construct more informative graphs.
[0120] The DAG between the points in is a bit tricky. For example, to avoid cycles, one could take The topological sorting of the vectors in . Parameter-free scalar projection of To define the sort, for example, when hour .function can be defined as , where each individual can be a monotonic function (such as the log-CDF of the standard normal distribution); in this case, one can guarantee that when individually for all dimensions People get under hour, . This ordering can then be used in the following formula
[0121] (5)
[0122] which results in a random adjacency matrix , the random adjacency matrix can be rearranged into a triangular structure (e.g., a DAG) with zeros on the diagonal. Note that this uses vectors instead of scalars, and that each vertex The prior on the representation of depends on By appropriately defining the and of , can include any inductive bias one wants about functions. For example, one can to encode an inductive bias about how neighboring points should depend on each other. This actually works pretty well. Figure 5 It is shown that FNP can learn . Figure 5 Examples of diagrams related to the embodiments are schematically shown. Figure 5 The following is about the MNIST data The DAG above, which is propagating The mean and thresholding of The edges with a probability less than 0.5 in are then obtained. It can be seen that the system learns meaningful graphs by connecting points with the same class. Note that in an embodiment, for example, during training, ,picture can be sampled, for example for each training batch.
[0123] The dependency graph has been obtained , , a dependency graph can be constructed to induce the , To this end, one can use and The structure of Each target variable The predictive distribution of is parameterized by the local latent variable To achieve this, the local latent variable Summarizes from The selected parent point and its target in context
[0124]
[0125] (6)
[0126] in , is accordingly , The return point , Notice that one can guarantee that the decomposition of the conditional in Eq. 6 is valid, because Coupled DAG corresponds to a directed tree. In this example, for example for a total exchangeable model, permutation invariance is desirable, so one can write The distribution above, for example Defined as Each dimension independent Gaussian distribution.
[0127] (7)
[0128] in and is a vector-valued function whose There are The codomain of the data tuple. It can be in The normalization constant in the case of The inverse of the number of parents of the point. When the point has no parents, an additional small value is used to avoid division by zero. By observing Equation 6, one can see that for a given The prediction is only indirectly through the graph , depends on the input covariates ,picture , yes Intuitively, it encodes the inductive bias for predictions about points “far away” — for example, with A very small probability of connecting to the reference set - will default to The standard normal prior is not informative above, so it is for A factorized Gaussian distribution was chosen for simplicity, and this is not a limitation. Any distribution is valid as long as it defines a probability density that is invariant with respect to the permutations of the parents.
[0129] However, Equation 6 can also hinder extrapolation, something that neural networks can do very well. In cases where extrapolation is important, one can adjust the Prediction ( The potential embedding of ) to add a direct path. This can be used as a middle ground, allowing In general, it provides a knob, because one can change and The dimension of is used to interpolate between GP and neural network behavior.
[0130] Now, by putting everything together, we get the FNP model and FNP The overall definition of the model:
[0131] (8)
[0132] (9)
[0133] The first term makes a prediction based on Equation 6, and the second term further Note that in addition to marginalization over latent variables and graphs, one can also marginalize over non-observation datasets. In this example, the reference set can be selected as the dataset , so the additional integration can be omitted. In general, marginalization can provide a mechanism to include unlabeled data into a model that can be used, for example, to learn better embeddings Or "impute" the missing labels. Having defined the models in Eqs. 8 and 9, one can show that they both define effectively permutation-invariant random processes and thus correspond to Bayesian models.
[0134] Two models have been defined, and their parameters The dataset can be used, for example is fitted, and the model can then be used for novel inputs Make a prediction. For simplicity, assume , although this is not required. The following description focuses on FNP because Notice that in this case, one gets .
[0135] Fitting the model parameters using maximum marginal likelihood is difficult because the necessary integration / summation of Eq. 8 is intractable. For this reason, variational inference can be employed to maximize the probability that the model parameters and the variational parameters of The marginal likelihood of
[0136] (10).
[0137] For a tractable lower bound, one can for example assume a variational posterior distribution Factorization ,in This leads to
[0138] (11)
[0139] The lower bound is decomposed into Item and corresponding to Item For large datasets , efficient optimization of this bound is possible. While the first term may not be amenable to small batches in general, the second term is amenable to small batches. As a result, one can use, for example, The size of the mini-batch is scaled by .
[0140] In fact, for and For all distributions above , one can use a diagonal Gaussian distribution, and for , , one can use concrete / Gumbel-softmax relaxation during training. This way, for example, the parameters can be jointly optimized by taking advantage of gradient-based optimization by employing pathwise derivatives obtained using the reparameterization trick , In addition, most parameters Possible differences between models and inference networks is related, because the regularization property of the lower bound can reduce the model parameters More specifically, for , , the neural network trunk can share two output heads, one for each distribution. The above priors can also be used in Points in aspects to be parameterized; , Both can be defined as ,in , Is for provides functions for the mean and variance, and is the linear embedding of the label. Tags From its potential code where the prior is not globally the same but can be conditioned on the parent of that particular point. This conditioning can transform the iid into a more general exchangeable model.
[0141] For the unseen points To perform predictions, one can take the posterior predictive distribution of FNP. More specifically, one can show that by using Bayes' rule, the predictive distribution of FNP has the form
[0142] (12)
[0143] in is the representation given by the neural network, and Can be marked from A binary vector of which points are the parents of the new point. For example, one can first use a neural network to project the reference set and the new point into a latent space on, and then by Bias it to Predictions are made based on the parent .
[0144] Figure 6a-6e An example of a prediction distribution for a regression task is schematically shown.
[0145] Several experiments were performed to validate the effectiveness of FNP, two of which are described in this paper. Comparisons were made against three baselines: a standard neural network (denoted as NN), a neural network trained and evaluated with Monte Carlo (MC) dropout, and a neural process (NP) architecture. For the first experiment, the inductive biases that can be encoded in FNP were explored by visualizing the prediction distribution on a one-dimensional (1d) regression task. For the second experiment, the prediction performance and the uncertainty quality that FNP can provide on the MNIST benchmark image classification task were measured. Experiments were also performed on CIFAR 10.
[0146] Figure 6a Involves MC-pressure difference. Figure 6b Involves neural processes. Figure 6c Involves Gaussian processes. Figure 6d and Figure 6e In the embodiments. Figure 6e In the case of , the first latent vector of the input data instance is used as an additional input to the third function. The predicted distribution for the regression task according to different models is considered. The shaded area corresponds to 3 standard deviations.
[0147] The generation process corresponds to Draw 12 points from Plot 8 points, and then parameterize the objective as ,in This generates nonlinear functions with "gaps" in the middle of the data where one would want uncertainty to be high. A heteroscedastic noise model was used for all models. A Gaussian process (GP) with an RBF kernel was also included. For the global potential of NP, 50 dimensions were used, and for FNP Potential, using dimensions. For the reference set , 10 random points were used for FNP, and the full data set was used for NP.
[0148] People can see that The FNP with RBF function has similar behavior to the GP. This is not the case for MC-Dropout or NP, where more linear behavior is seen in terms of uncertainty and false overconfidence in the region in the middle of the data. However, they do appear to extrapolate better, whereas the FNP and GP default to a flat zero prediction outside of the data. seems to combine the best of two worlds, as it allows for extrapolation and GP-like uncertainty, although for The modification of the free bits at the boundaries of helps encourage the model to rely more on these specific latent variables. Empirically, it is observed that Adding more capacity can make FNP to move closer to the behavior observed for MC-Dropout and NP. In addition, the model parameters The fact that the number of may cause FNP to overfit can lead to a reduction in prediction uncertainty.
[0149] For the second task, image classification of MNIST and CIFAR 10 was considered. For MNIST, the LeNet-5 architecture was used, which has two convolutional layers and two fully connected layers, while for CIFAR, a VGG-like architecture was used, which has 6 convolutional layers and two fully connected layers. In both experiments, the 300 random points as the target for FNP and the target for NP , for comparability, up to 300 points were randomly selected from the current batch for background points during training, and the same 300 points were used as FNP for evaluation. The dimension is , while for NP, the dimension of the global variable for MNIST is And for CIFAR it is .
[0150] As a proxy for the quality of uncertainty, an out-of-distribution (ood) detection task is given; given the fact that FNP is a Bayesian model, one would expect that the epistemic uncertainty of FNP would increase in areas where there is no data (such as the ood datasets). The metrics reported are the average entropy on those datasets, and the area under the ROC curve (AUCR) that determines whether a point is in or out of distribution based on the predicted entropy. Note that it is simple to increase the first metric by just learning a trivial model, but this will be detrimental to the AUCR; in order to have a good AUCR, the model must have low entropy on the in-distribution test set, but high entropy on the ood dataset. For the MNIST model, the ood datasets considered are: non-MNIST, Fashion MNIST, Omniglot, Gaussian distribution and uniform distribution Noise; for CIFAR 10, the sets used are: SVHN, fixed size Pixels of tinyImagenet, iSUN, and similar Gaussian and uniformly distributed noise. A summary of the results for MNIST can be seen in Table 1.
[0151] Table 1: Accuracy and uncertainty on MNIST from 100 posterior prediction samples. For all datasets, the first column is the mean prediction entropy, while for the ood dataset, the second column is the AUCR, and for within-distribution it is the test error in %.
[0152]
[0153] It can be seen that the two FNP models have comparable accuracy to the baseline model, while having higher average entropy and AUCR on the ood dataset. It appears to perform better than FNP. One can see that FNP almost always has better AUCR than all baselines considered. Interestingly, out of all the noise-free ood datasets, Fashion MNIST and SVHN are observed to be the most difficult to distinguish on average across all models. It is also observed that sometimes, noisy datasets on all baselines can act as “adversarial examples”, resulting in lower entropy than the in-distribution test set. FNP does have a similar effect on CIFAR 10, for example - although to a much smaller extent - the effect of FNP on uniformly distributed noise. It should be mentioned that other advances in ood detection are orthogonal to FNP and can further improve the performance.
[0154] About FNP In The trade-off between For The larger capacity leads to better uncertainty, which in turn seems to improve accuracy. These observations are based on the as a condition for learning meaningful .
[0155] The above experiment is used for FNP and FNP The architecture of is constructed as follows. A neural network backbone is used to obtain the input The intermediate hidden representation of , and then parameterize two linear output layers, one leading to parameters, and one that results in , both of which are fully factorized Gaussian distributions. Function for Bernoulli probability is set to RBF, e.g. ,in is optimized to maximize the lower bound. Throughout the training process, the temperature of the binary concrete / Gumbel-softmax relaxation is kept at and used the log CDF of the standard normal distribution as the of For the classifier , using the corresponding or The linear models are operated on top of . During training, a single Monte Carlo sample was used for each batch in order to estimate the bounds on FNP. Single samples were used for NP and MC-dropout. All models were implemented in PyTorch and run across five Titan X (Pascal) GPUs (one GPU per model).
[0156] NN and MC-Difference have the same trunk and classifier as FNP. NP is designed with the same neural network trunk to provide The intermediate representation of To obtain the global embedding ,Label Connected to obtain ,Will Projection to a linear layer dimensions, and then computes the average of each dimension across contexts. Global latent variables The parameters of the distribution over are then determined by acting on The linear layer on top of is given by After sampling, the A linear classifier operating on top of . The initial transformed regression experiments used 100 ReLUs for both NP and FNP models via a single layer of MLP, while the regressor used a linear layer for NP (larger capacity leads to reduction in overfitting and prediction uncertainty) and a single hidden layer MLP with 100 ReLUs for FNP. The MC-Dropout network used a single hidden layer MLP of 100 units and dropout was applied at the hidden layer with a rate of 0.5. In all neural network models, heteroskedastic noise was calculated according to is parameterized, where is the neural network output. For GP, the kernel length scale is optimized according to the marginal likelihood. A soft-free bits modification of the bounds is applied to help The optimization is beneficial, where the initial average across all dimensions and batch elements for FNP allows Free bit, and for FNP yes Spare bits, both are slowly annealed to zero within the 5k update process.
[0157] For the MNIST experiments, the model architecture is 20C5 - MP2 - 50C5 - MP2 - 500FC - Softmax, where 20C5 corresponds to a convolutional layer with 20 output feature maps with a kernel size of 5, MP2 corresponds to a max pooling with size 2, 500FC corresponds to a fully connected layer with 500 output units, and Softmax corresponds to the output layer. The penultimate layer of the network provides For the MC-Dropout network, a dropout of 0.5 is applied to each layer. The number of points in is set to , which is obtained by judging the performance of NP and FNP models on MNIST / non-MNIST pairs. For a FNP mini-batch of 100 points from M, we always set the full appended to each of those batches. For NP, a batch size of 400 points was used, where, for comparability with FNP, up to 300 randomly selected points from the current batch were used for background points during training, and the same 300 points as for FNP were used for evaluation. The training epochs for FNP, NN, and MC-Dropout networks were upper bounded to 100 epochs, and the training epochs for NP were upper bounded to 200 epochs since it performs fewer parameter updates per epoch than FNP. Optimization was performed with Adam using default hyperparameters. Early stopping was performed based on accuracy on the validation set, and no additional regularization was applied. Finally, soft-idle bit modification was applied to the bounds to help , where, throughout the training process, averaging across all dimensions and batch elements allows Free space.
[0158] The architecture used for CIFAR 10 experiments is 2x(128C3)-MP2 - 2x(256C3)-MP2 - 2x(512C3)-MP2 - 1024FC - Softmax with batch normalization after each layer (except the output layer). Similar to the MNIST experiments, the The initial representation of is provided by the penultimate layer of each network. Hyperparameters were not optimized for these experiments, but the same number of reference points, number of spare bits, number of epochs, regularization, and early stopping criteria were used as on MNIST. For MC-Dropout networks, a dropout is applied at a rate of 0.2 at the beginning of each stack of convolutional layers that share the same output channels, and at a rate of 0.5 before each fully connected layer. The initial learning rate of is optimized using Adam. The initial learning rate decays to 1 / for NN, MC-Dropout, and FNP every thirty epochs. , and for each NP period declines to 1 / During training, data augmentation was performed by random cropping with 4 pixels of padding and random horizontal flipping for both reference and other points. No data augmentation was used during test time. Images were further normalized by subtracting the mean and dividing by the standard deviation of each channel, calculated across the training dataset.
[0159] As mentioned above, the objective of FNP can be adapted to small batches, where the batch size is based on the reference set The procedure for FNP is the same as that for FNP The expansion of may be similar. The bound of FNP can be expressed in two terms:
[0160] (twenty three)
[0161] One of them has The term corresponding to the variational bounds of the data points in , and the second , which corresponds to when For the conditional The midpoint limit of Eq. 23 The term cannot in general be decomposed into independent sums, since The DAG structure in Item can; from The conditional iid property of and the structure of the variational posterior can be expressed as Independent sums:
[0162] (twenty four).
[0163] You can use Small batches of points in order to approximate the inner sum and thus obtain a mini-batch dependent Unbiased estimate of the population bound of :
[0164] (25)
[0165] Thus we get the following dependency on mini-batches An unbiased estimate of the population bound of
[0166] (26).
[0167] In practice, this may limit us to using relatively small reference sets, since training may become relatively expensive; in this case, an alternative would be to also subsample the reference set, and only train Reweight appropriately. This may provide a biased gradient estimator, but has been found to be effective.
[0168] It will be appreciated that the subject matter disclosed herein also extends to computer programs, particularly computer programs on or in a carrier, which are suitable for putting the subject matter disclosed herein into practice. The program may be in the form of source code, object code, code intermediate source, and object code such as partially compiled form, or in any other form suitable for use in the implementation of an embodiment of the method. An embodiment associated with a computer program product includes computer executable instructions corresponding to each processing step of at least one of the methods described. These instructions may be subdivided into subroutines and / or stored in one or more files that may be statically or dynamically linked. Another embodiment associated with a computer program product includes computer executable instructions corresponding to each device of at least one of the systems and / or products described.
[0169] Figure 8a A computer readable medium 1000 according to an embodiment is shown, which has a writable component 1010 including a computer program 1020, which includes instructions for causing a processor system to perform a machine learning method, such as a method of training and / or applying a neural network. The computer program 1020 can be embodied on the computer readable medium 1000 as a physical mark or by magnetization of the computer readable medium 1000. However, any other suitable embodiments are also conceivable. In addition, it will be appreciated that although the computer readable medium 1000 is shown here as an optical disk, the computer readable medium 1000 can be any suitable computer readable medium such as a hard disk, a solid state memory, a flash memory, etc., and can be non-recordable or recordable. The computer program 1020 includes instructions for causing a processor system to perform a method of training and / or applying a neural network.
[0170] Figure 8b In the schematic diagram is shown a representation of a processor system 1140 according to an embodiment of a machine learning method, such as training and / or applying a neural network, and / or according to a machine learning system. The processor system comprises one or more integrated circuits 1110 . Figure 8bThe architecture of one or more integrated circuits 1110 is schematically shown in FIG. The circuit 1110 includes a processing unit 1120, such as a CPU, which is used to run a computer program component to perform a method according to an embodiment and / or implement its modules or units. The circuit 1110 includes a memory 1122 for storing programming code, data, etc. Part of the memory 1122 may be read-only. The circuit 1110 may include a communication element 1126, such as an antenna, a connector, or both, etc. The circuit 1110 may include a dedicated integrated circuit 1124 for performing part or all of the processing defined in the method. The processor 1120, the memory 1122, the dedicated IC 1124, and the communication element 1126 may be connected to each other via an interconnect 1130 (e.g., a bus). The processor system 1110 may be arranged for contact and / or contactless communication using an antenna and / or a connector accordingly.
[0171] For example, in an embodiment, a processor system 1140, such as a training and / or application device, may include a processor circuit and a memory circuit, wherein the processor is arranged to execute software stored in the memory circuit. For example, the processor circuit may be an Intel Core i7 processor, an ARM Cortex-R8, etc. In an embodiment, the processor circuit may be an ARM Cortex M0. The memory circuit may be a ROM circuit or a non-volatile memory such as a flash memory. The memory circuit may be a volatile memory, such as an SRAM memory. In the latter case, the device may include a non-volatile software interface such as a hard drive, a network interface, etc., which is arranged to provide software.
[0172] It should be noted that the above-mentioned embodiments illustrate rather than limit the presently disclosed subject matter, and that those skilled in the art will be able to design many alternative embodiments.
[0173] In the claims, any reference marks placed between brackets should not be interpreted as limiting the claims. The use of the verb "comprise" and its conjugations does not exclude the presence of elements or steps other than those stated in the claims. The article "one" or "an" before an element does not exclude the presence of multiple such elements. Statements such as "at least one of them" when preceding a list of elements represent the selection of all elements or any subset of elements from the list. For example, the statement "at least one of A, B, and C" should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all A, B, and C. The subject matter currently disclosed can be implemented by hardware including several different elements and by a computer with suitable programming. In a device claim that lists several components, several of these components can be embodied by the same item of hardware. The mere fact that certain measures are recorded in mutually different dependent claims does not indicate that the combination of these measures cannot be used advantageously.
[0174] In the claims, references in parentheses refer to reference signs in the drawings of exemplary embodiments or formulas of the embodiments, thereby increasing the understandability of the claims. These references should not be construed as limiting the claims.
[0175] Reference Numbers List:
[0176] 300 Machine Learning Systems
[0177] 310 Training Storage Device
[0178] 312 Input Data Example
[0179] 314 Input Tags
[0180] 320 Reference storage device
[0181] 322 Reference Data Examples
[0182] 324 Reference Tags
[0183] 326 First potential reference vector
[0184] 328 Second potential reference vector
[0185] 332 An input interface arranged to obtain an input data instance
[0186] 330 First Function
[0187] 350 Parent identifier
[0188] 340 Second Function
[0189] 360 Latent Vector Determiner
[0190] 370 The third function
[0191] 380 training units
[0192] 412 Observed Input: , and tags
[0193] 413 Split
[0194] 422 Training Points:
[0195] 424 Reference point:
[0196] 426 Reference point label
[0197] 431 First Function
[0198] 432 Potential Space :point
[0199] 434 Potential Space :point
[0200] 435 Second Function
[0201] 436 Potential Space
[0202] 437 The third function
[0203] 442 Bipartite Graph:
[0204] 444 DAG:
[0205] 452 Training Points (from or )
[0206] 454 The parent in or get)
[0207] 456 Latent vector obtained from the parent latent vector
[0208] 464 from The predicted label
[0209] 466 Training signals for functions A, B and / or C
[0210] 610 true function
[0211] 620 Mean function
[0212] 622 Observations
[0213] 630 Mean function
[0214] 634 Observation
[0215] 640 Mean function
[0216] 642 Observations
[0217] 650 Mean function
[0218] 652 Observation
[0219] 654 Reference Point
[0220] 660 Mean function
[0221] 662 Observation
[0222] 664 Reference Point
[0223] 1000 Computer readable medium
[0224] 1010 Writable components
[0225] 1020 Computer Programs
[0226] 1110 Integrated Circuit(s)
[0227] 1120 Processing Unit
[0228] 1122 Memory
[0229] 1124 ASIC
[0230] 1126 Communication Components
[0231] 1130 Interconnect
[0232] 1140 processor system.
Claims
1. A machine learning system configured to map an input data instance to an output according to a system mapping, the system mapping being composed of a plurality of functions, the machine learning system comprising: an input interface arranged to obtain an input data instance, A storage device configured to store: A set of reference data instances (R), A processor circuit configured to: Mapping an input data instance to a first latent input vector in a first latent space (U) according to a first function (A), identifying a plurality of reference data instances in the set of reference data instances (R) as parents of the input data instance based on similarities between the first potential input vector and a first potential reference vector obtained from the set of reference data instances according to the first function (A), determining a second latent input vector (z) in a second latent space (Z) for the input data instance from a second latent reference vector obtained from the plurality of reference data instances identified as parents according to a second function (B), An output is determined from a second latent input vector (z) by applying a third function (C) to the second latent input vector, where At least one of the first function (A) and the second function (1B) and the third function (C) are machine learnable functions, and The input data instance and the reference data instance each include image data.
2. The machine learning system of claim 1, wherein The output includes the output label, the system map is a classification task, and / or The output includes variables, and the system mapping is a regression task.
3. The machine learning system of claim 1 or 2, wherein at least one of the first function, the second function, and the third function is a neural network.
4. The machine learning system of claim 1 or 2, wherein: A similarity between the first latent input vector and the first latent reference vector is determined according to a kernel function.
5. The machine learning system of claim 1 or 2, further configured to train a machine learnable function, wherein the processor circuit is configured to: Mapping training input data instances to outputs according to the system mapping, deriving a training signal from the error derived from the output, Adapt a machine learnable function based on a training signal.
6. The machine learning system of claim 1 or 2, further configured to train a machine learnable function, wherein the processor circuit is configured to: A reference data instance is selected from a set of reference data instances (R), and multiple reference data instances in the set of reference data instances (R) are identified as parents of the selected reference data instance based on a similarity between a first potential reference vector and a first potential reference vector obtained from the selected reference data instance according to a first function (A).
7. The machine learning system of claim 1 or 2, wherein: At least one of the first function (A), the second function (B), the third function (C), and the kernel function is configured to generate a probability distribution from which a corresponding output is sampled.
8. A machine learning system as described in claim 1 or 2, wherein the second potential input vector (Z) determined for the reference data instance further depends on the corresponding reference label.
9. A machine learning system as described in claim 1 or 2, wherein the third function (C) is configured to take the second potential input vector and the first potential input vector as input to produce an output.
10. The machine learning system of claim 1 or 2, wherein The output includes a plurality of output labels and a plurality of corresponding output probabilities, and the processor circuit is configured to calculate a confidence value from the plurality of output probabilities.
11. The machine learning system of claim 1 or 2, wherein An initial second latent reference vector is obtained for the input data instance by applying the second function (BB) to the input data instance, and a second latent input vector (z) in the second latent space (Z) is obtained for the input data instance from: the initial second latent reference vector, and the second latent reference vector obtained for the reference data instance identified as the parent.
12. The machine learning system of claim 1 or 2, wherein the training is divided into batches and the processor system is configured to Sample a directed acyclic graph (G) from reference data instances, Select a batch of training points and sample the corresponding part of the bipartite graph (A), A systematic mapping is applied to the reference data instances and the batch of training points to obtain a plurality of training signals.
13. An autonomous device controller having a machine learning system as described in any one of claims 1 to 12, wherein the input data instance includes sensor data of the autonomous device, the system mapping is configured to classify objects in the sensor data, and the autonomous device controller is configured to control the autonomous device depending on the classification.
14. A machine learning method (700) configured to map an input data instance to an output according to a system mapping, the system mapping being composed of a plurality of functions, the machine learning method comprising: Get the input data instance, Stores a collection of reference data instances (R), Mapping an input data instance to a first latent input vector in a first latent space (U) according to a first function (A), identifying a plurality of reference data instances in the set of reference data instances (R) as parents of the input data instance based on similarities between the first potential input vector and a first potential reference vector obtained from the set of reference data instances according to the first function (A), determining a second latent input vector (z) in a second latent space (Z) for the input data instance from a second latent reference vector obtained from the plurality of reference data instances identified as parents according to a second function (B), An output is determined from a second latent input vector (z) by applying a third function (C) to the second latent input vector, wherein at least one of the first function (A) and the second function (B) and the third function (C) are machine learnable functions, and wherein the input data instance and the reference data instance each include image data.
15. A transitory or non-transitory computer readable medium having data representing instructions which, when executed by a processor system, cause the processor system to perform the method of claim 14.
Citation Information
Patent Citations
System and method for generating automated response to an input query received from a user in a human-machine interaction environment
CN108459784A
A transductive and / or adaptive max margin zero-shot learning method and system
WO2018161217A1