Method and apparatus for operating a learning system

By guiding the selection of image representations, predicting labels and determining uncertainty with predictors, and combining acquisition functions and model fit metrics to optimize image selection, the problem of heavy labeling process in training artificial neural networks is solved, achieving efficient and low-cost learning results.

CN111914992BActive Publication Date: 2025-09-30ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010381572.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-10
Filing Date
2020-05-08
Publication Date
2025-09-30
Estimated Expiration
2040-05-08

AI Technical Summary

Technical Problem

The process of identifying and labeling the labeled image sets required for training artificial neural networks in the prior art is cumbersome and expensive, making it difficult to efficiently utilize unlabeled images for learning.

Method used

By guiding the selection of image representations, using oracles to determine labels, predictors to determine uncertainty and rewards, and selecting images based on uncertainty and rewards, we use probabilistic predictors and acquisition functions for active learning to reduce dependence on labeled data and combine model fit metrics to optimize image selection.

Benefits of technology

It achieves efficient training of artificial neural networks without the need for large amounts of labeled data, improves the flexibility and efficiency of the learning system, and reduces training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111914992B_ABST
    Figure CN111914992B_ABST
Patent Text Reader

Abstract

Methods and apparatus for operating a learning system are provided. The learning system (100) comprises at least one processor and at least one memory for instructions executable by the processor, the at least one processor being adapted to execute instructions for: selecting an image representation (x) from a plurality of image representations n ); Determine a first representation (y) of at least one first label for the image; Determine a second representation (y) of at least one second label in the prediction n ); Based on image representation (x n ), the first representation (y) and the second representation (y n ) determines the second representation (y n )’s prediction uncertainty; depends on the image representation (x n ) and / or the second representation (y n ) determining a reward (R); and selecting another representation of another image depending on the uncertainty and the reward (R), wherein the system (100) is adapted to determine at least one control signal for a video-based or image-based system, wherein the system (100) is adapted to process the image depending on the sensor output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and apparatus for operating a learning system, in particular a learning system for training an artificial neural network. Background Art

[0002] In supervised learning, artificial neural networks for video and / or image processing are trained using training data consisting of labeled images. Creating a collection of labeled images for training is cumbersome and expensive. It is therefore desirable to identify and label images that are particularly useful for training.

[0003] In one aspect, an acquisition function is used to identify an image from a collection of images. The acquisition function is a deterministic function that selects an image based on, for example, entropy or mutual information defined in terms of information related to image content (such as color, etc.).

[0004] The usefulness of a single acquisition function for a particular set of images depends on the image data. To use particularly useful acquisition functions, for example, S. Ebert, M. Fritz, and B. Schiele's "RALF: A Reinforced Active Learning Formulation for Object Class Recognistion" (CVPR, 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 3626-3633) discloses a method that learns a strategy based on a set of available heuristics to select an acquisition function from a set of predetermined and deterministic acquisition functions. Summary of the Invention

[0005] Further improvements of the learning system are achieved by the computer-implemented method and the learning system according to the invention.

[0006] A computer-implemented method for operating a learning system comprises: selecting an image representation from a plurality of image representations, in particular by a guide; determining a first representation of at least one first label for the image, in particular by an oracle; determining a second representation of at least one second label in the prediction, in particular by a predictor; determining an uncertainty of a prediction of the second representation, in particular by the predictor, based on the image representation, the first representation and the second representation; determining a reward depending on the image representation and / or the second representation; and selecting another representation of another image depending on the uncertainty and depending on the reward; determining at least one control signal for a video-based or image-based surveillance system, a customer monitoring system, a safety engineering system, a driver assistance system, a multifunctional comfort system, a safety-critical assistance system with intervention in vehicle control, a system for fusing driver information or driver assistance, a robotic system, a home or garden work robotic system, a networked or collaborative robotic system, wherein the control signal is determined as an output of the system in response to an input of at least one image captured by the learning system, wherein the system is adapted to process images depending on a sensor output of an image, video, radar, LiDAR (Light Detection and Ranging), ultrasound or motion sensor. Thus, a free-form acquisition heuristic is learned in an active learning setting based on reinforcement feedback, without requiring any labeled dataset to generate rewards. This method is applicable to predictors that report their uncertainty, such as prediction variance. This provides an improved self-learning control method. This control is also very useful in many applications, particularly in the automotive and appliance sectors.

[0007] Advantageously, the reward is determined by comparing a first measure of model fit, particularly for a predictor, and a first measure of model fit, with a second measure of model fit, wherein the first measure is determined as a function of at least one first parameter of the model determined before training the model with the representation and as a function of at least one second parameter of the model determined after training the model with the representation. Thus, a measure of model fit (e.g., marginal likelihood) is used to determine the reward for a new image, both before and after training with the new image, as a function of uncertainty.

[0008] Advantageously, the method comprises training a probabilistic predictor, in particular a Gaussian process or a deep Bayesian neural network, for predicting a second representation based on a randomly selected image representation, and then training the predictor for predicting the second representation based on an image representation selected as a function of the reward. In particular, the randomly selected inputs for supervised learning are small compared to the inputs for supervised learning selected as a function of the reward. This provides efficient learning at a low cost in terms of labeled input points.

[0009] Advantageously, for predictor training, a subset of input points is selected from the plurality of image representations depending on an acquisition function, wherein the acquisition function is determined in particular by the guidance, depending on the reward and the probability state. During the learning process, the acquisition criterion is learned functionally as a function of the inputs by reinforcement learning. This provides a significantly more flexible learning system with respect to image selection.

[0010] Advantageously, the input points are arranged in a permuted order depending on the acquisition function, and a subset of input points is selected from the top of the permuted order, in particular by selecting every k-th input point, where k is a natural number greater than 1. In this way, from the available images, the most useful images are used first for training. Thinning the input points breaks repetitive trends and encourages diversity.

[0011] Advantageously, the acquisition function is learned by a probabilistic policy predictor, in particular a deep Bayesian neural network, depending on the probabilistic state and based on the reward. In this way, reinforced feedback from the processed unlabeled images is used to select the next image to be processed.

[0012] Advantageously, the plurality of image representations are processed recursively, wherein a set of samples is selected from the plurality of image representations, labeled, and added to a plurality of input points comprising the labeled images, wherein the plurality of input points comprising the labeled images are processed for training of the predictor, and wherein the set of samples is removed from the plurality of image representations. This makes the training process very efficient.

[0013] The learning system comprises at least one processor and at least one memory for instructions executable by the processor, wherein the at least one processor is adapted to execute the instructions for selecting an image representation from a plurality of image representations; determining a first representation of at least one first label for the image; determining a second representation of at least one second label in a prediction; determining an uncertainty of a prediction of the second representation, in particular based on the image representation, the first representation, and the second representation; determining a reward depending on the image representation and / or the second representation; and selecting another representation of another image depending on the uncertainty and depending on the reward, wherein the system is adapted to determine at least one control signal for a video-based or image-based surveillance system, a customer monitoring system, a safety engineering system, a driver assistance system, a multifunctional comfort system, a safety-critical assistance system with intervention in vehicle control, a system for fusion of driver information or driver assistance, a robotic system, a home or garden work robotic system, a networked or collaborative robotic system, wherein the control signal is determined as an output of the system in response to an input of at least one image captured by the learning system, wherein the system is adapted to process the image depending on a sensor output of an image, video, radar, LiDAR, ultrasound, or motion sensor. The learning system can be applied to various use cases. The learning system comprises, in particular, a guide adapted to select an image; an oracle adapted to determine a first label for the image; and a predictor adapted to predict a second label for the image and output an uncertainty of the prediction. The oracle uses, for example, a human-machine interface for allowing a human user to label images presented to the human user. The oracle is also adapted in the example to determine a reward. The guide selects an image depending on the uncertainty and the reward. Thus, the predictor learns a free-form acquisition heuristic based on reinforced feedback in an active learning setting without requiring any labeled data set to generate a reward. To this end, the predictor in particular reports its prediction variance as uncertainty. The oracle, guide and predictor can be implemented as one or more computer programs and run on one or more processors.

[0014] Advantageously, the system is adapted to determine a reward depending on a model, in particular for a predictor, and a comparison of a first measure of model fit with a second measure of model fit, wherein the first measure is determined depending on at least one first parameter of the model determined before training the model with the representation and depending on at least one second parameter of the model determined after training the model with the representation. In this aspect, the oracle provides ground-truth labels as part of the environment surrounding the model. The model computes model fit. From a Bayesian perspective, a principle measure of model fit is the marginal likelihood, where the parameters are the respective prediction variances before and after training with new images.

[0015] Advantageously, the system comprises a probabilistic predictor, in particular a Gaussian process or a deep Bayesian neural network, wherein the system is adapted to train a predictor for predicting a second representation based on a randomly selected image representation, and then to train the predictor for predicting the second representation based on an image representation selected depending on the reward. The selected representations are labeled and used to train the predictor using one of many known learning or inference methods. The set of randomly selected input points used for supervised learning is preferably selected to be small compared to the set of input points selected for supervised learning depending on the reward. In this way, the best images from the large pool of input points used for unsupervised learning can be identified and used most efficiently.

[0016] Advantageously, the system is adapted to select a subset of input points from the plurality of image representations for predictor training depending on an acquisition function, wherein the acquisition function is determined depending on the reward and the probability state. Instead of using a manually selected acquisition criterion or function, in particular the maximum entropy or mutual information contained in the input points for unsupervised learning, the acquisition function is learned during the learning process by reinforcement learning, for example by bootstrapping.

[0017] Advantageously, the system is adapted to arrange the input points in a permuted order depending on the acquisition function, and in particular to select a subset of input points from the top of the permuted order by selecting every k-th input point, where k is a natural number greater than 1. This sparsification of the input points, for example by a heuristic bank filter, breaks repetitive trends and encourages diversity.

[0018] Advantageously, the system comprises a probabilistic policy predictor, in particular a deep Bayesian neural network, adapted to learn an acquisition function depending on the probabilistic state and based on the reward, so that the reinforcement feedback from the processed unlabeled images is used by the guide to select the next image to be processed.

[0019] Advantageously, the system is adapted to recursively process the plurality of image representations, wherein a set of samples is selected from the plurality of image representations, labeled, and added to a plurality of input points comprising the labeled images, wherein the plurality of input points comprising the labeled images are processed for training the predictor, and wherein the set of samples is removed from the plurality of image representations. In this aspect, the training of the predictor relies on the representation determined by the oracle. This makes the training process very efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Further advantageous embodiments can be derived from the detailed description and the accompanying drawings. In the drawings:

[0021] Figure 1 A schematic diagram depicting parts of a learning system,

[0022] Figure 2 Describes the steps in a method for operating the learning system,

[0023] Figure 3 Portions of the learning system and devices comprising the learning system are depicted. DETAILED DESCRIPTION

[0024] Figure 1 Portions of an exemplary learning system 100 are schematically depicted.

[0025] The acquisition function is determined using a data-driven approach described below. Since prediction uncertainty is a necessary input to the acquisition heuristic, in the example a deep Bayesian neural network is used as the base predictor. In order to obtain a high-quality estimate of the prediction uncertainty at an acceptable computational cost, a deterministic approximation scheme is described that can both efficiently train the deep Bayesian neural network and compute its posterior predictive density following a series of closed-form operations. In this way, probabilistic information - i.e., the uncertainty provided by the predictions of the deep Bayesian neural network - is used for state design, which leads to a full-scale probability distribution. This distribution is then fed into a probabilistic policy network, such as another Bayesian neural network, which is trained by reinforcing feedback collected from each labeling round in order to inform the system of the success of its current acquisition function. This feedback fine-tunes the acquisition function, resulting in improved performance in subsequent labeling rounds.

[0026] The training uses samples that are recursively processed using epochs comprising batches.

[0027] An exemplary training is based on a large unlabeled collection of images, e.g., 1 million images. Training begins with a small labeled subset of this collection (e.g., 100 images). This subset is, for example, randomly selected from the collection, i.e., without active learning. A predictor, such as a Bayesian neural network or a Gaussian process, is trained based on this small subset. In this example, a subset of images is removed from the collection.

[0028] Then the active learning round begins. For example, 1000 images are selected to be labeled in the labeling round based on the active learning algorithm, which will decide which images will be labeled.

[0029] At each labeling round, the large unlabeled set (e.g., 1 million - 100 images in the first labeling round) is used to train a predictor that was previously initially trained on a small, randomly selected subset of images (e.g., 100 images).

[0030] In a predicted model using a predictor, each unlabeled image of the large unlabeled set is assigned an acquisition score by an acquisition scoring function, and a ranked order of the images is determined by ranking the images with respect to these acquisition scores. The acquisition scoring function is referred to as an acquisition function.

[0031] The top ranked images (e.g., the top 10 images) are presented to the oracle. The oracle provides labels for these images and adds them to a previous small set of images. More specifically, the oracle is adapted to present the images to a human via a human-machine interface, such as a display. The human-machine interface is adapted to detect input, more specifically labels, provided by the human while viewing the images, such as via a keyboard. Since the top ranked images are the most relevant, the learning system presents these to the human user for labeling. Thus, the human user does not have to label all images, but only labels the images that are considered to be most relevant.

[0032] Then, for example after the first labeling round with 100+10=110 labeled images, the predictor is trained with the new small set.

[0033] As the model changes recursively, its acquisition scoring function is also changed.

[0034] This cycle is repeated, for example, until our labeling budget is exhausted.

[0035] To predict, use the data set It consists of the eigenvector and labels y for the C-dimensional binary output labels n ∈{0,1} C It consists of N tuples.

[0036] We parameterize any neural network f(·) with parameters w that follow a certain prior distribution p(w). In this example, we use the following data generation process:

[0037]

[0038] where f c is the c-th output channel of the neural network, where is the normal cumulative distribution function, and Ber(·|·) is the Bernoulli distribution.

[0039] For a test data point x * , using the posterior predictive distribution, which performs model averaging over the latent variables w based on their posterior distributions (e.g., below),

[0040] p(y * |x * ,D)=∫p(y * |x * ,w)p(w|D)dw

[0041] For X={x1,...,x N} and Y={y1,...,y N}

[0042]

[0043] Considering the normal mean field approximation posterior

[0044]

[0045] Where V is the number of synaptic connections in the neural network. For a new observation x * , the following approximation can be used to predict the output label y *

[0046]

[0047] where N(·|·,·) is a normal distribution with two free parameters, the tuple represents the weight w for the network f(·) i The set of variational parameters of (for i=1,...,V) is the weight prior p(w i ) is a set of hyperparameters, and the function and The cumulative mapping from the input layer to the top L moments of the exemplary Bayesian neural network predictor is encoded.

[0048] The data-driven label acquisition described in this paper can perform adaptation during active learning labeling rounds. This is achieved through a reinforcement learning framework, where, in this example, a policy net—a deep Bayesian neural network—is trained by observed rewards from the training environment.

[0049] Exemplary definitions of states, actions, and rewards are given below for this framework.

[0050] In this example, based on the unlabeled input point D u In the initial learning before using unlabeled input points, labeled input points D can also be usedI In this example, the input point is the representation x of the image set. n .

[0051] For efficient processing, in this example, the unlabeled sample set is first arranged into unlabeled input points D via an information-theoretic heuristic, such as the maximum entropy criterion. u As such, the heuristic assigns similar scores to samples with similar characteristics, and consecutive samples in the permutation inevitably have high correlation.

[0052] To break the trend and enhance diversity, in one aspect, according to the sorted order of the input points, starting from the top of the order, every k-th input point is picked until M samples {x1, ..., x M}until.

[0053] These samples are fed into a Bayesian neural network predictor, which determines the posterior predictive density estimate for each sample and the state of the unlabeled sample space with the following distribution:

[0054]

[0055] in is the mean, and is the variance of the activation of the output neuron, which is calculated as described above for the Bayesian neural network predictor. The state S follows a C×M dimensional normal distribution.

[0056] In this example, the actions include labeling the samples from the set {x i1 ,...,x iM} are sampled, for example, according to the probability masses assigned to them.

[0057] In the example, the reward is calculated by the oracle 102 based on a measure of model fit and is calculated for a new data point with a newly labeled pair (x, y) as

[0058]

[0059] where q old (·) is the variational posterior before training with a new data point (x,y), and where q new (·) is the variational posterior after training with a new data point (x,y).

[0060] Optionally, a second component that encourages diversity across the selected labels throughout the labeling round is calculated as

[0061] The payoff in this option depends on R improv and R div The sum of is calculated.

[0062] Additionally, the policy net π(·)—in this example, the second Bayesian neural network—is parameterized by φ ≈ p(φ). The input to the policy net π(·) is the state S. The output of the policy net π(·) is a parameterization of the M-dimensional class distribution over possible actions. The output is, for example, for the action taken at time t. The M binary probabilities are determined

[0063]

[0064] Then based on the class distribution an action is selected from:

[0065]

[0066] where Cat(·) is the category distribution.

[0067] In one aspect, an episodic version of the REINFORCE algorithm according to RJ Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning” (Machine learning, 8(3-4):229–256, 1992) is adapted to train the policy net π(·).

[0068] A moving average over all past returns can be used. In the example, the labeling episode consists of selecting a sequence of points to label, for example with a discount factor of γ = 0.95, after which the BNN is retrained and the policy net π(·) takes an update step. Policy π φ (·|S t ) can be parameterized by a neural network with parameter φ.

[0069] The method iterates between labeling nodes, training policy net π(·) and training deep Bayesian neural network f(·). An exemplary pseudo code is given below:

[0070]

[0071]

[0072] Among them G tis the cumulative reward at time step t of the processed node, e.g.

[0073] Figure 1 An exemplary learning system 100 and an exemplary workflow for operating the system are depicted.

[0074] The illustrated example learning system 100 includes as components a oracle 102, a predictor 104, a bulk filter 106, and a guide 108. Instructions for performing the functions provided by these components can be executed on one or more processors and can be stored on one or more memories. The functions of the individual components are described below. However, the functions can be combined or separated in different ways. The interfaces between these functions are adapted accordingly.

[0075] It is assumed that the image is available for input from memory or from a capture device. For the following description, x n is the input observation, i.e., the independent variable, such as the raw image pixel data representing an image, and y or y n is the output label of interest, such as a numeric variable representing the class or descriptive name for the image.

[0076] The guide 108 is adapted to select an image representation x from a plurality of image representations M n In an example, the guide 108 selects a plurality of images as input points to be labeled by the oracle 102 .

[0077] The oracle 102 is adapted to determine the n In one aspect, the oracle 102 receives the plurality of images as input points and provides a set of labeled data for the predictor 104 to learn on. The oracle 102 includes a human-machine interface to display the first representation y of at least one first label of the image represented by x. n The represented image and determining the first representation y according to the input detected by the human-machine interface.

[0078] The oracle 102 is equipped with an interface, such as a human-machine interface, to provide images (eg, to a human user) and receive corresponding image tags (eg, from the human user).

[0079] The predictor 104 then converts the uncertainty of the prediction—in this example, about the mean Variance ——Provided to the main filter 106.

[0080] The predictor 104 is adapted to determine a second representation y of at least one second label n , and is used to represent x based on the imagen , the first representation y and the second representation y n To determine the second representation y n uncertainty in the forecast.

[0081] The predictor 104 is, for example, a probabilistic predictor, in particular the aforementioned Gaussian process or a deep Bayesian neural network.

[0082] The guide 108 uses an acquisition function (such as the deep Bayesian neural network described above) that utilizes reinforcement feedback from the oracle 102 (such as the reward R = R improv +R div ) and the probability state S to learn how to optimally choose the next point. It is thus able to flexibly adapt itself to the data set at hand.

[0083] The guide 108 communicates to the oracle 102 which input points to label next, thereby restarting the cycle.

[0084] In one aspect, the system 100 is adapted to train the predictor 104 based on randomly selected image representations and then based on image representations selected based on a reward R. The predictor is trained on this set using one of many known learning or inference methods. The randomly selected input is preferably selected to be small compared to the input selected based on the reward.

[0085] The guidance 108 is adapted to select the image representation depending on the uncertainty and depending on the reward R.

[0086] The oracle 102 is adapted to depend on the image representation x n and / or the second representation y n To determine the return R.

[0087] The oracle 102 is adapted to determine a reward R depending on the model and a comparison of a first measure of model fit with a second measure of model fit, wherein the first measure depends on the performance of the model using the representation x n At least one first parameter of the model determined before training the model, such as q old (θ old ) and depends on the use of x n At least one second parameter q of the model determined after training the model new (θ new In this aspect, the reward R is calculated based on the difference between the first metric and the second metric. The difference is for the improvement in model fit, for example, for R as described above. improv returns.

[0088] Additionally, the reward for diversity in the labels of the selected batch can be used to determine the reward Rdiv .

[0089] The guide 108 is adapted to select a subset of input points from a set of input points for training of the predictor 104. The set of input points in this context refers to a plurality of image representations.

[0090] The guidance 108 is in one aspect adapted to select the subset depending on the permuted order of the input points and on an acquisition function that is learned depending on the reward R and the probability state S. In an example, the acquisition function is learned by reinforcement learning during the learning process.

[0091] The arrangement is provided, for example, by the subject filter 106. In an example, a heuristic subject filter 106 is used to determine the probability state S based on the uncertainty.

[0092] In an example, the subject filter 106 is adapted to apply active learning. The subject filter 106 takes a prediction (e.g., mean) for each data point in the unlabeled set. ) and its associated uncertainties (e.g. variance ) and rank them by using heuristics, about obtaining scores. Heuristics are for example mutual information or variance

[0093] The intuition is that all acquired scores will agree on a large portion of the unlabeled set which is not very informative.

[0094] In one aspect, starting from the top of the permutation, data points are selected to obtain a set of M selected data points as an action until an order of k*M is reached. This means that the input points are arranged in a permuted order and the input points are selected according to the acquisition function. In particular, input points are selected from the top of the permuted order by selecting every k-th input point, where k is a natural number greater than 1.

[0095] In one aspect, the mean and variance of the predictions of the Bayesian neural network on these M data points are used to construct a multivariate normal distribution. The normal distribution is used as the probability state S in this aspect. The input point for the acquisition function corresponds to the probability state S in this aspect.

[0096] The guide 108 uses this probability state S to learn to select one input data point to label from the set of M actions provided. As a result of the action it has selected, the guide 108 receives a reward R from the oracle 102.

[0097] The learning system 100 described above is adapted to recursively process images from a collection of unsupervised samples, which include unlabeled images. In one aspect, the learning system 100 is adapted to recursively process a plurality of randomly selected images and labels of a sample collection of labeled images in advance, without active learning, for training the predictor 104. This means that the predictor 104 is trained based on these images and labels.

[0098] The following references Figure 2 To describe a computer-implemented method of operating a learning system. Figure 2 The method comprises the steps of the method. In one aspect, the steps are recursively repeated as described above in pseudo code for training the predictor 104 and for learning the acquisition function. According to one aspect, a plurality of image representations are recursively processed, wherein in one recursion a set of samples is selected from the plurality of image representations, labeled, and added to a plurality of input points comprising the labeled images, wherein the plurality of input points comprising the labeled images are processed for training the predictor 104, and wherein the set of samples is removed from the plurality of image representations, in particular before starting the next recursion. Not all steps need to be performed in each recursion, and the order of the steps may vary between recursions.

[0099] After the method starts, in step 202, an image representation x is selected from the plurality of image representations M. n In one aspect, the plurality of image representations M correspond to input points arranged in a permuted order, and a subset of input points is selected from the top of the permuted order, in particular by selecting every k-th input point, where k is a natural number greater than 1.

[0100] Then execute step 204.

[0101] In step 204 , a first representation y of at least one first label for the image is determined, for example, by prediction 102 as described above.

[0102] Then execute step 206.

[0103] In step 206, a second representation y of at least one second label is determined, for example by predictor 104, as described above. n , as the output of the deep Bayesian neural network f(·).

[0104] Then execute step 208.

[0105] In step 208, specifically based on the image representation x n , the first representation y and the second representation y n To determine the second representation yn uncertainty in the forecast.

[0106] Then step 210 is executed.

[0107] In step 210, depending on the image representation x n and / or the second representation y n The reward R is preferably determined based on the model and a comparison of the first measure of model fit with the second measure of model fit (e.g., based on R as described above). improv and optionally according to R div ) is determined. The first metric depends, for example, on the n At least one first parameter of the model determined before training the model, such as q old (θ old ) and depends on the use of x n At least one second parameter q of the model determined after training the model new (θ new ) to be determined.

[0108] Then execute step 212.

[0109] In step 212 , another representation of another image is selected depending on the uncertainty and on the reward R.

[0110] The acquisition function is determined as a function of the reward R and the probability state S. The acquisition function is learned as a function of the probability state S and based on the reward R, for example, by a probabilistic policy predictor, in particular a deep Bayesian neural network π(·).

[0111] Steps 202 to 212 may be recursively repeated until the tagging budget is exhausted, all batches of all sections of the input data are processed, or until another criterion is met.

[0112] Learning systems can generally be operated in a variety of appliances.

[0113] exist Figure 3 In one aspect depicted in , the learning system 100 is part of a control system 300 adapted to determine at least one control signal for an output 302 .

[0114] The control signal is particularly used for a video-based or image-based system 304. The control signal is determined as an output of the learning system 100 in response to an input received at an input 306, in particular an input of at least one image captured by a sensor 308 of the system 100. The sensor 308 is, for example, an image, video, radar, LiDAR, ultrasound or motion sensor.

[0115] In an example, the learning system 100 includes a processor 310 and a memory 312 for storing instructions that, when executed by the processor 310 , cause the processor to operate the learning system according to the method described above.

[0116] In examples, the video-based or image-based system can be a surveillance system, a customer monitoring system, a safety engineering system, a driver assistance system, a multifunctional comfort system, a safety-critical assistance system with intervention in vehicle control, a system for fusing driver information or driver assistance, a robotic system, a home or garden work robotic system, a networked or collaborative robotic system.

Claims

1. A computer-implemented method of operating a learning system for training a predictor based on randomly selected image representations and then training the predictor based on image representations selected based on rewards (R), characterized in that: Select an image representation from multiple image representations (x n ); determining a first representation (y) of at least one first label representing a class of the image by prediction; Determine, by the predictor, a second representation (y) of at least one second label representing the class of the image in the prediction n ); Through the predictor, based on the image representation (x n ), the first representation (y) and the second representation (y n ) to determine the second representation (y n )’s forecast uncertainty; Depends on the image representation (x n ) and / or the second representation (y n ) to determine the return (R); and selecting another image representation depending on the uncertainty and depending on a reward (R); wherein for training the predictor a subset of input points is selected from the plurality of image representations depending on an acquisition function, where the acquisition function is determined by the reward (R) and the probability state (S), where the probability state (S) is determined depending on the uncertainty, The acquisition function is learned by a probabilistic policy predictor, depending on the probabilistic state (S) and based on the reward (R).

2. The method according to claim 1, characterized in that The reward (R) is determined based on the model of the predictor and a comparison of a first measure of model fit with a second measure of model fit, where the first measure is determined based on: n ) to train the model before determining at least one first parameter of the model; and n ) to train the model and determine at least one second parameter of the model.

3. The method according to claim 1 or 2, characterized in that: Based on randomly selected image representations, we train the system to predict the second representation (y n )’s probability predictor; and then trained to predict the second representation (y based on the image representation selected depending on the reward (R) n ) is a probability predictor.

4. The method according to claim 1 or 2, characterized in that Input points are arranged in an arranged order depending on the acquisition function, and a subset of input points is selected from a top of the arranged order by selecting every k-th input point, where k is a natural number greater than 1.

5. The method according to claim 1 or 2, characterized in that The plurality of image representations are recursively processed, wherein a set of samples is selected from the plurality of image representations, labeled, and added to a plurality of input points including the labeled images, wherein the plurality of input points including the labeled images are processed for training of a predictor, and wherein the set of samples is removed from the plurality of image representations.

6. The method of claim 1, wherein the probabilistic policy predictor is a deep Bayesian neural network.

7. The method of claim 3, wherein the probabilistic predictor is a Gaussian process or a deep Bayesian neural network.

8. A system comprising units for carrying out the method of claims 1-7.

9. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that A memory having stored thereon the computer readable medium according to claim 9.