Difficulty Adaptation Training for Machine Learning Modules

By introducing difficulty functions and discriminant classifiers, the training data set is selected and expanded, and the generation of difficult samples is generated using the generative adversarial network, the training and evaluation problems of the machine learning module in an uncertain environment are solved, classification accuracy and self-evaluation capabilities are improved, and more secure automated decision-making is achieved.

CN112446413BActive Publication Date: 2025-07-11ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010884586.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-29
Filing Date
2020-08-28
Publication Date
2025-07-11
Estimated Expiration
2040-08-28

AI Technical Summary

Technical Problem

Existing machine learning modules are difficult to effectively train and evaluate their output uncertainty in the face of highly variable uncertainty, resulting in misclassification and operational risks in safety-critical applications.

Method used

By introducing difficulty functions and discriminant classifiers, the training data set is selected and expanded, and difficult samples are generated using Generative Adversarial Networks (GANs), and the training input data is optimized through the encoder-decoder architecture to generate training samples with high difficulty to improve the training effect of the machine learning module.

Benefits of technology

Improve the classification accuracy and self-evaluation capabilities of the machine learning module in uncertain environments, reduce the risk of misclassification, and achieve safer automated decision-making through sensor fusion and uncertainty calibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112446413B_ABST
    Figure CN112446413B_ABST
Patent Text Reader

Abstract

A method (100) for obtaining and / or augmenting a training data set (11*) for a machine learning module (1), comprising: • providing (110) a sample (11a) of training input data annotated with a label (13a) regarding a given problem; • obtaining (120) from the sample (11a) of training input data and the associated label (13a) a difficulty function (14) that provides a measure of the difficulty (14a) of evaluating the sample (11a) regarding the given problem; • obtaining (130) at least one candidate sample (15) of training input data and / or its representation (25) in a workspace (20); • calculating (140), by means of the difficulty function (14), a measure of the difficulty (14a) of evaluating this candidate sample (15) and / or its representation (25) regarding the given problem; and including (160) the candidate sample (15) in the training data set (11*) in response to this difficulty meeting a predetermined criterion (150).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to the training of machine learning modules for the purpose of such classification, which is particularly suitable for safety-critical applications with highly variable uncertainties. Background Art

[0002] For the automation of many safety-critical tasks, such as at least partially automated driving of vehicles in road traffic, the use of trainable machine learning modules is envisaged. For example, such a module can acquire physical measurement data from the environment around the vehicle and classify objects such as road markings, road signs, pedestrians or other vehicles. An added benefit of using such a module is that, based on training on a limited number of situations, correct operation of the vehicle can be reasonably expected even in situations that are not part of the training. In this sense, the training process can be regarded as similar to the training process for a human driver. A human driver only spends a few dozen hours driving during the training process, but then is expected to handle any unexpected situations that may occur during their lifetime.

[0003] For this and other safety-critical applications, it is crucial to know the uncertainty of the output of the machine learning module. For example, such uncertainty may be caused by poor quality of physical measurement data (such as poor visibility conditions), by the inherent difficulty of identifying certain objects in some situations (such as a car partially obscured by a truck), or even by deliberately manipulating road signs or cameras using "adversarial" patterns, which may deceive the machine learning module into making a misclassification.

[0004] US 2019 / 122 120 A1 discloses a method for augmenting a training data set for a generative adversarial network (GAN). For augmentation, samples generated by a generator and unlabeled samples from the training data set can be used.

[0005] US 2018 / 373 963 A1 discloses an image classification system that aggregates the outputs of two distinct classifiers. One classifier is a common instance classifier that is trained to identify and recognize common instances of objects that typically occur. The other classifier is a rare instance classifier that is trained to compute a rarity score that represents the likelihood that the input image is correctly classified by the common instance classifier. Summary of the Invention

[0006] The present invention is defined by the appended claims. Embodiments and examples that are not covered by the claims are presented to illustrate the claimed invention and facilitate understanding of the claimed invention.

[0007] Disclosure of the Invention

[0008] The present invention provides a method for obtaining and / or augmenting a training data set for a machine learning module that maps input data to output data that is meaningful with respect to a given problem to be solved. A prime example of such a problem is the classification of objects based on physical measurement data.

[0009] The term "machine learning module" shall specifically include: a module that embodies a function parameterized with adjustable parameters and ideally having a high generalization ability. When training a machine learning module, the parameters can be specifically adjusted such that when training input values are fed into the module, the associated ground truth labels are replicated as accurately as possible. Specifically, the machine learning model can include or be an artificial neural network.

[0010] During the process of this method, samples of training input data are provided. Each such sample is associated with a label. With respect to the problem to be solved, this label constitutes the "ground truth". That is, if the machine learning module maps a sample of the training input data to output data corresponding to the label of this sample (for example, it maps a photo of a stop sign to "stop sign"), then this is considered meaningful with respect to the given problem.

[0011] Based on the sample of the training input data and the associated label, a difficulty function is configured to map the sample of the training input data or its representation in the workspace to a measure of the difficulty of evaluating this sample with respect to the given problem. For example, the difficulty can be a measure of the uncertainty that occurs when classifying an object based on physical measurement data.

[0012] As mentioned above, there are many reasons why the difficulty may vary. For example, as every human driver knows, it is more difficult to recognize an object at night or in heavy rain at a distance. In a situation where the lighting is weak and the vehicle is moving very fast, motion blur in the camera picture may also hinder the recognition of the object. Even under the best conditions, some objects have an inherent possibility of being confused with each other. For example, in Germany, the green light in a set of traffic lights may take the form of an arrow pointing to the right, indicating that a right turn is allowed and unobstructed in the sense that no other person (such as a pedestrian) is ever allowed to cross the path of the vehicle making the right turn. However, there is also a similar sign with a very different meaning. There may be a fixed sign next to the red light with a green arrow painted on it. This indicates that a right turn is allowed even when the red light is on, but all vehicles doing so must first stop and also pay attention to the traffic at the intersection. That is, when the red light is red for one direction, other vehicles in the perpendicular direction have a green light, and if such a vehicle is approaching from left to right, a right turn is not possible even though there is a green arrow on the sign.

[0013] Introducing new traffic signs can also pose a risk of confusion. For example, in Germany, the recently introduced "Umweltzone" sign indicates an area where only vehicles meeting certain low emission standards are allowed to pass. It consists of a red circle with the black word "Umwelt" inside and the black word "Zone" below the circle. The layout of this sign is the same as that of the long-known "30 km / h zone" sign; the only difference is that "Umwelt" is inside the circle instead of the large number "30". Thus, the disadvantage of reusing the layout is that the two signs may be confused.

[0014] Obtain at least one candidate sample of the training input data and / or its representation in the workspace. By means of a difficulty function, calculate a measure of the difficulty of evaluating this candidate sample and / or its representation in the workspace with respect to a given problem. In response to this difficulty satisfying a predetermined criterion, include the candidate sample in the training data set.

[0015] Here, the workspace can be any space in which it is otherwise easier, faster, more accurate, or more advantageous to obtain or evaluate the difficulty function. For example, if the problem to be solved is classification, it may be advantageous to select a workspace in which the samples belong to the same class clusters and the clusters are separated from each other.

[0016] The beneficial effect of this method is twofold.

[0017] First, during the collection of physical measurement data for training, certain situations may occur more frequently than others. For example, if the data is collected by a camera mounted on a vehicle, traffic signs will appear with a frequency distribution corresponding to their usage in the area where the data is collected. Currently, this frequency distribution is dictated by the needs of human drivers and is not optimal in many ways for training a machine learning module to recognize objects. For example, stop signs and yield signs are designed to have distinctive shapes so that their meaning can be recognized even in the most adverse conditions, such as being completely covered by snow. The machine learning module may require very few samples of those under different imaging conditions and then can be trusted all the time to recognize these important signs. However, those signs are among the most frequently used signs, making them appear much more frequently in the dataset than necessary. On the other hand, there are many signs that do not appear as frequently but are still very important. Signs indicating railroad crossings and signs warning drivers not to drive off a dock into the water are examples of such signs. According to the method described above, if there is already enough training input data to recognize stop signs and yield signs, and the next training image with a stop sign or a yield sign is fed into the training, the difficulty function will consider the difficulty of this training image to be very low, and the redundant training image will not be included in the final training dataset. On the other hand, if the training has encountered very few dock warning signs so far, then the next image with such a sign will receive a very high difficulty level. Therefore, it is very likely to be accepted in the final training dataset.

[0018] That is, the composition of the training dataset can be better tailored to the actual needs to become better. New training data can be specifically selected to address the weaknesses in learning. At the same time, redundant information can be eliminated from the training dataset. Compared with what was previously possible, a smaller training dataset can be utilized and less training time can be used to achieve a given level of accuracy.

[0019] Second, the difficulty level generated by the difficulty function can be used to improve the self-assessment of the machine learning module regarding the uncertainty of its output. In safety-critical applications, the uncertainty of the output of the machine learning module can be at least as important as the output itself. For example, high uncertainty may indicate that the output should not be trusted at all. As will be discussed in more detail below, the precise quantification of uncertainty can be used to trigger many useful actions to address this uncertainty. For example, if visibility deteriorates, data from a radar sensor can be used to supplement data from a camera. In the worst-case scenario, the machine learning module may indicate that it cannot properly handle the situation and seek help from a human operator or remove at least part of the automated vehicle from the flowing traffic.

[0020] As discussed previously, in particularly advantageous embodiments, the problem to be solved includes mapping each record of the input data to a record of the output data that, for each of a set of multiple discrete classes, indicates the probability and / or confidence that the record of the input data belongs to the corresponding class. Then, the difficulty function can include training a discriminative classifier based on samples of the training input data and associated labels, which maps the samples of the training input data and / or their representations in the workspace to a classification indicating the probability and / or confidence that each of these samples of the training input data belongs to each of the discrete classes. Then, the difficulty determined by the difficulty function depends on this classification.

[0021] Here, the term "record" should not be understood to require that a "record" have a fundamentally different data type from a "sample". Rather, both are in the same space. The term "record" is the more general term, and a "sample" is a special kind of record that is part of a training data set or is being evaluated for inclusion in a training data set. In the field of machine learning, the term "sample" is meant to convey that a finite training data set is being used to sample an unknown ground truth data distribution.

[0022] By training a discriminative classifier, the machine learning module "learns how to learn" to some extent. Once the discriminative classifier is established, additional candidate samples of the training input data can be examined against a predefined criterion of difficulty that is associated with the problem to be solved by the machine learning module by virtue of the training of the classifier. These candidate samples can come from any source. For example, if new training data is obtained by driving a vehicle around and recording images using cameras mounted on the vehicle, then the entire image stream can be tested against the criterion of difficulty. Then only the images that meet the criterion can be selected for inclusion in the training data set. In this way, a first culling of a large amount of newly acquired physical measurement data (such as images) can be carried out in a well-motivated way using a quantitative criterion associated with the problem to be solved. This is different from what would be done in the case of delegating the first culling to a human. For example, when classifying images, features that make the classification of the image more difficult do not have to be so visually prominent in the image that they would attract a person's attention. It is even possible to deliberately hinder the classification of an image by the machine learning module by introducing unobvious modifications ("adversarial examples") that change the classification.

[0023] In particularly advantageous embodiments, the difficulty determined by the difficulty function is at least partially based on the ambiguity of the classification. For example, if the output of a discriminative classifier is a softmax score, the mass distribution of that score among the possible classes indicates ambiguity. If the mass of the softmax score is concentrated on a particular subset of classes, while all other classes receive only a small total mass, this indicates a particular ambiguity among the classes in the particular subset. To train a machine learning module to better resolve such ambiguities, the training data set can be deliberately enriched with samples that lie on the borderline between classes. The machine learning module can learn much more from samples of this kind than it can learn from only the clear examples of the corresponding classes. This situation is analogous to learning how to fly an airplane. The crucial part is not learning how to perform routine flights. Instead, the crucial part is handling a wide variety of emergency situations, where the correct action will restore control of the airplane, the wrong action will crash the airplane, and which action is the correct one is not immediately obvious.

[0024] Another metric that the difficulty function can use is the entropy of the classification. If the entropy of the softmax score is high, this is another special case of ambiguity, where the difference is that many classes may be involved. Moreover, it may not even be possible to discern which class the softmax score prefers.

[0025] In additional particularly advantageous embodiments, candidate samples are associated with labels corresponding to the classifications determined by a discriminative classifier. In this way, the uncertainty of the classification in the ultimately trained machine learning module can be determined with much greater precision, thus counteracting the tendency of the machine learning module to become "overly confident".

[0026] When labeling training data, this is typically done by human labor, which can make labeling the most expensive step in the workflow. For example, humans are given the task of classifying certain objects or features that appear in an image and annotating the image. Humans are good at classifying whether an object is present in an image, but the determination of difficulty is difficult to quantify and is very subjective. Therefore, in most previously used training data sets, the labels are "one-hot encoded" or "hard-coded", which accurately assigns a single class to each sample of the training input data, regardless of how easy or difficult the choice was to reach that label. In contrast, the new labels determined using a discriminative classifier are "soft" labels, in the sense that they can include values in a continuous interval (e.g., between 0 and 1). Since the classifications determined by the discriminative classifier are also used to determine difficulty, the new "soft" labels reflect the difficulty. This will reduce the confidence of the machine learning module in making classifications.

[0027] In a further particularly advantageous embodiment, the working space is selected such that the degree of correlation between the similarity of samples of the training input data with respect to a given problem and the distances between the representations of these samples in the working space is equal to or higher than the degree of correlation with the distances between these samples. For example, by selecting an appropriate working space, the representations of the samples of the training input data can form clusters corresponding to the available classes for classification.

[0028] This in turn makes the working space a particularly valuable tool for deliberately obtaining samples of training input data having certain properties with respect to the difficulty that will be encountered when classifying them using the machine learning module. In a further particularly advantageous embodiment, a search is performed within the working space for candidate representations of candidate samples of the training input data. The search is guided by a difficulty function which, in this example, operates within the working space. Candidate representations are obtained from the working space such that the difficulty determined by the difficulty function meets a predetermined criterion. This obtained candidate representation is then transformed into a candidate sample which can then be added to the training data set.

[0029] How precisely a candidate representation can be transformed into a candidate sample depends on the specific choice of the working space and on the specific method chosen for creating this working space. In a further particularly advantageous embodiment, an encoder for transforming these samples into representations in the working space and a decoder for reconstructing these samples based on the representations are co-trained based on the samples of the training input data. The training is performed with the aim that the result of the reconstruction best matches the corresponding samples from which they were ultimately derived. By training the encoder-decoder architecture in this way, a working space is automatically created based on the representations produced by the trained encoder from the samples of the training input data. In other words, based on the training input data, the data manifold is learned and a search for representations having the desired properties with respect to difficulty can be performed on this manifold. At the same time, the decoder allows the found candidate representation to be directly transformed into the sought-after candidate sample of the training input data.

[0030] In particular, the representation of a given sample of the training input data in the working space will likely have a smaller dimension than the sample itself, i.e., it is characterized by a smaller number of elements. The representation is then a compressed representation and the search space forms a certain degree of "information bottleneck" between the encoder and the decoder. Creating such an "information bottleneck" is one possible way of making the similarity of samples of the training input data with respect to a given problem more highly correlated with the distances between these representations in the working space. In other words, if samples of the training input data are similar with respect to a given problem, their representations in the working space tend to move together. If the samples are different with respect to a given problem, their representations in the working space tend to move apart.

[0031] For example, a generative adversarial network based on a variational autoencoder (VAE-GAN) can be used. This arrangement includes: an encoder E (z|x), which maps a sample x of the training input data from the data acquisition space X to a representation z in the working space Z. The arrangement further includes: a decoder G (x′|z), which maps the representation z back from the space Z to a sample x′ of the training input data in the data acquisition space X. The arrangement further includes: a discriminator D ω (X), which evaluates whether a sample from the data acquisition space X it is given is a true sample x of the training input data that has always resided in this space X, or whether it is the final result x′ of the transformation from the space X to the space Z and back via the VAE-GAN. The encoder, decoder, and discriminator are trained together. The combination of the encoder and decoder endeavors to produce better "pseudo" samples x', while the discriminator endeavors to improve its ability to distinguish "pseudo" samples x' from true samples x. After the training of the VAE-GAN is completed, x' is indistinguishable from x, so when a candidate representation z is found in the space Z as a result of the search, the inverse transformation sample x' in the space X can be used as a training sample for training a machine learning module just like any other true training sample.

[0032] A discriminative classifier C γ (y|z) that maps a sample z in the space Z to a class y can be trained after training the encoder-decoder architecture. However, it is advantageous to train the discriminative classifier and the encoder-decoder architecture simultaneously. This allows the encoder and decoder to accelerate the training process. Alternatively, the classifier can also take the form C γ (y|x'), i.e., it can operate on the reconstructed sample x'.

[0033] In a particularly advantageous embodiment, the execution of the search includes: obtaining a candidate representation by solving an optimization problem regarding a merit function in the working space based on the representation in the working space, where the merit function is at least partially based on a difficulty function.

[0034] In this way, if a suitable merit function is available, new samples of the training input data suitable for a specific purpose can be obtained relatively quickly. By virtue of the optimization, it has been ensured that this new sample will meet a predetermined criterion regarding the difficulty according to the difficulty function.

[0035] For example, considering a sample x associated with a class label y, this can be transformed by the encoder E (z|x) into a representation z. A new representation z' located on the class boundary between the current class y and another class y'≠y can then be obtained by solving the following optimization ("class target perturbation"):

[0036]

[0037] Then, with the help of the decoder G (x'|z), the obtained z′ is transformed into the sought-after new sample x′. This new sample x' exhibits attributes from at least two classes, and the soft label (i.e., the softmax output of the discriminative classifier C γ reflects this by distributing most of the mass across those two or more classes.

[0038] In another example, given a sample x with or without a label, this can be transformed by the encoder E (z|x) into the representation z again. The new representation with the maximum entropy can then be obtained by solving the following optimization:

[0039]

[0040] The obtained z′ can again be transformed by the decoder G (x'|z) into the sought-after new sample x′. This new sample x' exhibits features from possibly multiple classes and has a very high entropy.

[0041] In another particularly advantageous embodiment, the search is carried out by: drawing random representations from the workspace and evaluating whether these random representations meet a predetermined criterion. This may involve drawing many samples and checking them one by one, so it can be computationally intensive. However, it opens the way to determine new samples that meet the difficulty criterion for which no merit function is available. For example, in security-related applications, specific difficulty functions and difficulty criteria can be specified by regulatory constraints, so it is not possible to replace them with something more tractable using a merit function.

[0042] The present invention also provides a method for training a machine learning module. This method starts with obtaining and / or augmenting a training data set using the above method. After that, the trainable parameters of the machine learning module are optimized. Samples of the training input data are processed by the machine learning module to produce output data. The parameters of the machine learning module are optimized such that the output data best matches the corresponding labels associated with the samples of the training input data. In this way, the knowledge in the rich training data set (i.e., how to handle unclear and ambiguous situations) is transformed into the behavior of the machine learning module, which generalizes this knowledge to new situations that occur during the use of the machine learning module in its intended applications.

[0043] As discussed previously, evaluating physical measurement data acquired by sensors is an important application area of machine learning modules that have been acquired by at least one sensor. Therefore, the present invention also provides a method for evaluating physical measurement data.

[0044] During the process of this method, a machine learning module is provided. The machine learning module is trained using the above method. At least one sensor is used to obtain physical measurement data. The physical measurement data is provided as input data to the trained machine learning module. The trained machine learning module maps the input data to output data and an associated uncertainty.

[0045] Many efforts have been made in the art to improve the accuracy of output data. However, specifically for safety-critical applications, accurately knowing the uncertainty can be at least as important as the output data itself (such as object classification in an image). In such cases, it is most desirable that the uncertainty determined within the machine learning module actually matches the accuracy of the output data, that is, the self-assessment of the accuracy of the machine learning module regarding its own situation is correct. If this is the case, the uncertainty of the machine learning module is called "calibrated". Based on such calibrated uncertainty, appropriate actions can be taken in a technical system that relies on the output data of the machine learning module. The training dataset generated by the above method allows the machine learning module to achieve such calibration during training.

[0046] In a particularly advantageous embodiment, an actuation signal is calculated from the output data determined by the machine learning module according to a policy, and this actuation signal is provided to a vehicle, to a classification system, to a safety monitoring system, to a quality control system, and / or to a medical imaging system. The policy depends on the uncertainty associated with the output data.

[0047] For example, if the uncertainty is high, such a policy may be more conservative, and relying on the output data (which then turns out to be incorrect or inappropriate) may have adverse consequences. In the example of an at least partially automated vehicle, different policies may be appropriate depending on how certain the classification of an object on the road ahead of the vehicle is. If the object is classified as a cardboard box or other debris that any human driver would run over, and this classification is very certain, then it is appropriate for the vehicle to continue at the current speed and run over the object. A sudden brake could come as a complete surprise to other drivers and may cause a rear-end collision. If the object is something that any human driver would brake for (such as a hard object that could damage the vehicle), and this classification is very certain, then fully applying the brakes is appropriate because it is most likely to avoid damage. However, if the classification is uncertain, it is more appropriate to adopt a policy that is likely to avoid damage even in the case of a classification error. For example, the brakes can be applied with a milder deceleration, which allows other drivers to react, and at the same time, a free path around the object can be calculated to be prepared in case the result should not be to run over the object.

[0048] The determined uncertainty can also be proactively used to improve the further training of the machine learning module for a specific application. In a further particularly advantageous embodiment, in response to the uncertainty generated from the recording of at least one input data satisfying a predetermined criterion, the recording of the input data can be stored in a memory for later analysis and / or labeling. Alternatively or in combination, it can be determined that the recording of the input data corresponds to an extreme case regarding the problem to be solved by the machine learning module.

[0049] For example, when collecting images for the further training of a machine learning module by a camera mounted on a vehicle, it is desirable to specifically capture data from which the machine learning module can still learn something new. When driving the vehicle to capture data, it is only possible to a limited extent to deliberately cause such situations to occur in the data. Instead, one can only wait for such situations to happen. By monitoring the uncertainty, it is possible to discover online whether a new recording of the input data is a recording from which the future training of the machine learning module can benefit. Therefore, it is possible to immediately discard data that is not of interest for further training so as not to overwhelm the limited storage capacity on board the vehicle. For example, data found to be of interest can be labeled later by a human expert or sent to a labeling service. These resources are expensive, so it is advantageous to concentrate their use only on recordings of input data that are of interest.

[0050] In a further particularly advantageous embodiment, physical measurement data is acquired by means of a plurality of distinct sensors, for example the sensors can acquire measurement data by means of different physical modalities. For example, a camera sensing visible light can be combined with a radar, ultrasonic or LIDAR sensor.

[0051] The physical measurement data from each of the sensors is provided as input data to the trained machine learning module. For the data from each sensor, corresponding output data and uncertainty corresponding to each sensor are obtained from the machine learning module. The corresponding output data is fused into a final result at least partially based on the corresponding uncertainty.

[0052] In this way, different sensors can pool their capabilities in order to ultimately arrive at output data and actuation signals generated from such output data that are appropriate for the situations sensed by the sensors in a broader class of situations.

[0053] For example, if a camera image is used as the main information source for at least partially driving a vehicle automatically, situations may arise where the image is not truly conclusive and the uncertainty is high. There may be poor lighting, or an object may be partially occluded by other objects. In such situations during real-time driving, there is no time to analyze in detail what is wrong with the camera image and which other sensor can be used to correct this specific defect. By comparing uncertainties, regardless of what exactly makes the data from one sensor uncertain, it is possible to automatically determine which sensor provides the most useful information for evaluating the situation at hand.

[0054] Specifically, in the case of an attack in adversarial mode, this also provides automatic failover. Such an adversarial mode is crafted to cause a change in the input data, which in turn triggers a misclassification. It is not possible for the same adversarial mode to be able to change the input data captured by different physical contrast mechanisms in the same way to trigger exactly the same misclassification while remaining unobtrusive to humans. For example, an adversarial sticker fixed to a stop sign may cause a camera-based object detection to misclassify the stop sign as a "70 km / h speed limit" sign. However, this sticker will be completely invisible to a radar sensor that still detects the octagon made of metal.

[0055] The methods described above can be implemented, in whole or in part, by a computer. Accordingly, the present invention also relates to a computer program having machine-readable instructions that, when executed by one or more computers, cause the one or more computers to perform one or more of the methods described above. In this sense, an electronic control unit of a vehicle and other devices capable of executing programming instructions is also understood to be a computer.

[0056] The computer program can be embodied on or in any kind of non-transitory storage medium and / or implemented as a downloadable product. A downloadable product is a product that can be purchased and delivered immediately via a network connection, thus eliminating the need for physical delivery of a tangible medium.

[0057] Additionally, the computer program, non-transitory storage medium, and / or downloadable product can be provided to a computer.

[0058] In the following, additional advantageous options are described together with preferred embodiments of the present invention using the accompanying drawings. Description of the Drawings

[0059] The drawings illustrate:

[0060] Figure 1 An exemplary embodiment of a method 100 for obtaining and / or augmenting a training data set 11*;

[0061] Figure 2The relationship between the data acquisition space X and the working spaces Z, 20;

[0062] Figure 3 An exemplary smooth transition between the classes "6" and "0" of the MNIST image dataset;

[0063] Figure 4 An exemplary embodiment of a method 200 for training a machine learning module 1;

[0064] Figure 5 An exemplary embodiment of a method 300 for evaluating physical measurement data 3a;

[0065] Figure 6 An exemplary ambiguous fuzzy candidate sample 15 included in the training dataset 11*;

[0066] Figure 7 for the E-MNIST dataset ( Figure 7a ) and for the Fashion MNIST dataset ( Figure 7b ) and for the MNIST dataset ( Figure 7c ) exemplary improvement of the calibration confidence C of the measured control accuracy A. Detailed description of the preferred embodiments

[0067] Figure 1 is a flowchart of an exemplary embodiment of a method 100 for obtaining and / or augmenting a training dataset 11*. In step 11, a sample 11a of training input data associated with a label 13a is provided.

[0068] In step 115, the working space 20 can be selected such that the similarity of the samples 11a is more highly correlated with the distance between the representations 21a of these samples 11a in the working space 20 than with the distance between these samples 11a themselves. This is used to separate distinct classes from each other and to cluster the samples 11a into different classes. The relationship between the data acquisition space X in which the samples 11a are located and the working spaces 20, Z is explained in more detail in Figure 2 .

[0069] The working spaces 20, Z can be defined and, according to step 116, all the transformations 17, 18 to and from the data acquisition space X are established all at once. This is explained in more detail in Figure 2 .

[0070] In step 120, a difficulty function 14 is obtained from sample 11a and associated label 13a. For sample 11a of the training input data, and / or for representation 21a of such sample 11a in the workspace 20, this difficulty function 14 outputs a measure of the difficulty 14a of evaluating this sample 11a with respect to a given problem such as object classification. According to block 121, a discriminative classifier 16 can be trained, and the difficulty function 14 depends on the classification 16a output for sample 11a and / or for its representation 21a.

[0071] In step 130, at least one candidate sample 15 of the training input data is obtained for inclusion in the training data set 11*. This candidate sample 15 can be obtained directly in the data acquisition space X, or indirectly via the workspace 20. Inside box 130, an exemplary way of indirectly obtaining the candidate sample 15 via the workspace 20 is shown.

[0072] In step 131, a search for a candidate representation 25 is performed within the workspace 20 such that a criterion 150 for the difficulty 14a is satisfied. In step 132, this candidate representation 25 is transformed into a candidate sample 15 in the data acquisition space X where the original sample 11a of the training input data is located. In Figure 1 Two exemplary ways of performing the search 131 are shown.

[0073] According to block 131a, an optimization problem within the workspace 20 is solved with respect to a merit function. This merit function is at least partially based on the difficulty function 14, and can additionally be based on the criterion 150 for the difficulty 14a in order to ensure that the optimal guarantee criterion 150 of the merit function is satisfied.

[0074] In step 140, the difficulty function 14 is used to calculate a measure of the difficulty 14a of evaluating the candidate sample and / or its representation with respect to the problem at hand. That is, the difficulty function can operate in the data acquisition space X or in the workspace 20, and it can include a mixture of contributions from both spaces.

[0075] If the difficulty function 14 utilizes the discriminative classifier 16, then according to block 141, the classification 16a output by this discriminative classifier 16 can also be used to associate the candidate sample 15 with a label.

[0076] The difficulty 14a is checked against a predetermined criterion 150, which can for example include ambiguity and / or entropy in the classification 16a. For example, the criterion 150 can include a threshold. For example, the criterion 150 can also include: the selection of candidate samples 15 with the "top n" difficulty 14a from a larger set of candidate samples 15. There is complete freedom with respect to the criterion 150.

[0077] If the criterion 150 (true value 1) is satisfied, the candidate sample 15 is included in the training data set 11*.

[0078] Figure 2 Illustrated is the relationship and transformation between on the one hand the data acquisition space X in which the samples 11a of the training input data exist and on the other hand the working space Z, 20 in which the representations 21a of these samples 11a exist. The encoder 17 is trained to map the samples x, 11a to the representations 21a, and the decoder 18 is trained to map the representations 21a back to the reconstructed samples x', 11a' in X. Under the supervision of the discriminator 19, the encoder 17 and the decoder 18 are trained in a self-consistent manner, where the discriminator 19 measures the degree of match between the reconstructed samples x', 11a' and the corresponding original samples x, 11a, and is also trained. The training of the discriminative classifier 16 can be combined with this training such that the training can exchange synergistic effects and reduce the overall computation time.

[0079] As discussed previously, the search for new candidate samples 15 in X can be simplified to a search for new candidate representations z', 25 in Z. For example, this may be advantageous because there are suitable merit functions for optimization in Z. The found candidate representation 25 can be transformed into a candidate sample x', 15 in X.

[0080] Figure 3 Illustrated is the smooth transition between the various classes on the standard MNIST data set. The MNIST data set includes: images of handwritten digits from 0 to 9, which are correspondingly labeled with class labels from 0 to 9. The smooth transition is calculated in the working space Z, 20, which is created by training the encoder 17 from X to Z, training the decoder 18 from Z to X, and training the discriminator 19 to distinguish between the reconstructed samples 11a' and the real samples 11a. Then the sample 11a corresponding to picture (a) is transformed into a representation 21a, z in Z, and there, the optimization discussed above is transformed into a "perturbation for the class"

[0081]

[0082] Here, the higher the number c that determines the upper limit of the summation, the more the optimization on δ will bring the newly created representation z' towards the representation of the picture corresponding to the target class y'. To form Figure 3In FIGS. (b) to (h), new representations z' with an increased c value are created, and these representations z' are inversely transformed back into the images in X by means of the decoder 18. These images correspond to the candidate samples 15. Ranking of these candidate samples 15 using the difficulty function 14 will most likely reveal that for the image (d), it is most difficult to distinguish whether its representation is the digit "6" or the digit "0". Therefore, the corresponding sample 15 lies near the class boundary in Z between the classes "6" and "0", and this sample 15 will most likely be selected for inclusion in the training data set 11*.

[0083] Figure 4 is a flowchart of an exemplary embodiment of a method 200 for training a machine learning module 1. In step 210, the method 100 described in conjunction with Figure 1 is used to obtain and / or augment the training data set 11*. In step 220, the trainable parameters 12 of the machine learning module 1 are optimized such that the output data 13 of the machine learning module 1 best matches the label 13a associated with the sample 11a of the training input data. This training can be continued until a predetermined termination criterion is reached. At that time, the value of the trainable parameter 12 corresponds to the value of the fully trained machine learning module 1*.

[0084] Figure 5 is an exemplary method for evaluating physical measurement data 3a that has been acquired using at least one sensor 3. In step 310, a machine learning module 1 is provided. In step 320, the machine learning module 1 is trained using the method 200 described in conjunction with Figure 4 . In step 330, at least one sensor 3 is used to acquire physical measurement data 3a. In step 340, the physical measurement data 3a is provided as input data 11 to the trained machine learning module 1*. The machine learning module 1* maps this input data to output data 13 and an associated uncertainty 13b.

[0085] According to block 331, the physical measurement data 3a can be acquired by multiple sensors 3. According to block 341, then the physical measurement data 3a from all sensors can be provided to the trained machine learning module 1* such that the corresponding output data 13 and uncertainty 13b for each sensor 3 are obtained. According to block 351, these corresponding output data 13 can be fused into a final result at least partially based on the uncertainty 13b. At the output of step 350, whether the output data 13 and uncertainty 13b are based on the input data 11 from only one sensor 3 or from multiple sensors 3, they appear the same.

[0086] Figure 5Two exemplary uses of the uncertainty 13b are shown, and with improved training 200 under a training data set 11* improved according to the method 100, the uncertainty 13b can be determined much more accurately than would be possible according to the prior art.

[0087] In step 360, an actuation signal 4 is calculated from the output data according to a policy which in turn depends on the uncertainty 13b associated with the output data 13. In step 370, the actuation signal 4 is provided to the vehicle 50, to the classification system 60, to the safety monitoring system 70, to the quality control system 80, and / or to the medical imaging system 90.

[0088] The uncertainty 13b generated from at least one record of the input data 11 can also be checked against a predetermined criterion 380. If this criterion (truth value 1) is met, then in step 381, the record of the input data 11 is stored in a memory for later analysis and / or labeling. Alternatively or in combination, in step 382, it is determined that the record of the input data 11 corresponds to an extreme case of the problem at hand that requires special attention.

[0089] Figure 6 Three examples of fuzzy candidate samples 15 for inclusion in the training data set 11* are shown, which can be used to train the machine learning module 1 to classify handwritten digits from the MNIST data set. In addition to each of the samples 15, the weight distribution of the corresponding softmax scores for the ten classes "0" to "9" is also shown in shaded form, where the darker shading corresponds to a higher score.

[0090] Picture (a) has features from the digits "3", "7", and "9". Picture (b) has features from the digits "2", "7", and "8". Picture (c) has features from the digits "3" and "8". Even for humans, it is difficult to assign these pictures to a single class with high confidence. Therefore, the samples 15 shown in these pictures are good candidates for inclusion in the data set 11*.

[0091] In FIG. 7, the confidence C that the machine learning module 1 gives to its output data 13 is plotted against the accuracy A, which is measured using labeled validation data (which is not yet part of the training). Figures 7a to 7c is generated according to the same scheme but for different data sets. In Figures 7a to 7c each of them, curve a represents an ideal calibration state where the confidence C exactly matches the accuracy A. Curve b represents the training progress based on the corresponding original training data set when viewed from left to right. Curve c represents the training progress based on the augmented training data set 11*.

[0092] Figure 7a It was created on the E-MNIST dataset, which contains handwritten letters to be classified into 26 categories from 'A' to 'Z'. Figure 7b It was created on the Fashion MNIST dataset, which contains clothing images to be classified. Figure 7c It was created on the MNIST dataset.

[0093] For all three datasets, the following general trend emerged: Curve c was closer to the ideal curve a most of the time, although there were short segments where curve b was better.

Claims

1. A method (100) for training a machine learning module (1), comprising: Augmenting a training data set (11*) for a machine learning module (1) that maps input image data (11) to output data (13) that is meaningful with respect to image classification according to a usage method, the method comprising: · Providing (110) a sample (11a) of training input data associated with a label (13a) such that if the machine learning module (1) maps this sample (11a) of training input data to output data corresponding to the label (13a), this is considered meaningful with respect to image classification; · Obtaining (120) a difficulty function (14) from the sample (11a) of training input data and the associated label (13a), the difficulty function (14) being configured to map the sample (11a) of training input data or its representation (21a) in a workspace (20) to a measure of the difficulty (14a) of evaluating this sample (11a) with respect to image classification; · Obtaining (130) at least one candidate sample (15) of training input data and / or its representation (25) in the workspace (20); · Computing (140), by means of the difficulty function (14), a measure of the difficulty (14a) of evaluating this candidate sample (15) and / or its representation (25) with respect to image classification; and · Responsive to this difficulty satisfying a predetermined criterion (150), including (160) the candidate sample (15) in the training data set (11*), and further comprising: training the machine learning module (1) using the augmented training data set (11*), wherein image classification comprises: mapping each record of the input data (11) to a record of the output data (13) that, for each of a set of multiple discrete classes, indicates the probability and / or confidence that the record of the input data (11) belongs to the corresponding class, obtaining (120) the difficulty function (14) comprises: training (121) a discriminative classifier (16) based on the sample (11a) of training input data and the associated label (13a), which maps the sample (11a) of training input data and / or its representation (21a) in the workspace (20) to a classification (16a) indicating the probability and / or confidence that this sample (11a) of training input data belongs to each of the discrete classes, and wherein the difficulty (14a) determined by the difficulty function (14) depends on this classification (16a), the difficulty (14a) determined by the difficulty function (14) is at least partially based on the ambiguity of the classification (16a) and / or the entropy of the classification (16a).

2. The method (100) according to claim 1, further comprising: Associating (141) the candidate sample (15) with a label corresponding to the classification (16a) determined by the discriminative classifier (16).

3. The method (100) according to any one of claims 1 to 2, wherein Select the working space (20) such that the degree of correlation between the distance between the samples (11a) of the training input data with respect to image classification and their representations (21a) in the working space (20) is the same as or higher than the degree of correlation between the distance between these samples (11a).

4. The method (100) according to claim 3, further comprising: Perform (131) a search for candidate representations (25) of candidate samples (15) of the training input data within the working space (20) such that the difficulty (14a) determined by the difficulty function (14) satisfies a predetermined criterion (150), and transform (132) this candidate representation (25) into the candidate sample (15).

5. The method (100) according to claim 4, wherein, The performance of the search (131) includes: obtaining (131a) a candidate representation (25) by solving an optimization problem for a merit function in the working space (20) based on the representation (25) in the working space (20), wherein the merit function is at least partially based on the difficulty function (14).

6. The method (100) according to claim 4, wherein The performance of the search (131) includes: drawing (131b) random representations (25) from the working space (20) and evaluating whether these random representations (25) satisfy a predetermined criterion (150).

7. The method (100) according to any one of claims 1 to 2, further comprising: Based on the samples (11a) of the training input data, train (116) an encoder (17) for transforming these samples (11a) into representations (21a) in the working space (20) and a decoder (18) for reconstructing the samples (11a) from the representations (21a) such that the reconstructed result (11a') best matches the corresponding sample (11a).

8. A method (300) for evaluating physical measurement data (3a), comprising: · Providing (310) a machine learning module (1); · Training (320) the machine learning module (1) using the method according to claim 1; · Using at least one sensor (3) to obtain (330) physical measurement data (3a); and · Providing (340) the physical measurement data (3a) as input image data (11) to the trained machine learning module (1*) such that the trained machine learning module (1*) maps (350) the input data (11) to output data (13) and an associated uncertainty (13b), the output data (13) being meaningful with respect to image classification.

9. The method according to claim 8, further comprising: · Calculating (360) an actuation signal (4) from the output data (13) according to a strategy depending on the uncertainty (13b) associated with the output data (13); and · Providing (370) the actuation signal (4) to a vehicle (50), to a classification system (60), to a security monitoring system (70), to a quality control system (80) and / or to a medical imaging system (90).

10. The method (300) according to any one of claims 8 to 9, further comprising: In response to the uncertainty (13b) resulting from the recording of at least one input data (11) satisfying a predetermined criterion (380): · Storing (381) the recording of the input data (11) in a memory for later analysis and / or marking; and / or · Determine (382) that the record of the input data (11) corresponds to an extreme case of an image classification problem to be solved by the machine learning module (1).

11. The method (300) according to any one of claims 8 to 9, further comprising: · Obtain (331) physical measurement data (3a) by means of a plurality of distinct sensors (3); · Provide (341) the physical measurement data (3a) from each of the sensors (3) as input data (11) to the trained machine learning module (1*) in order to obtain corresponding output data (13) and uncertainty (13b) corresponding to each sensor (3); and · Fuse (351) the corresponding output data (13) into a final result, at least in part based on the corresponding uncertainty (13b).

12. A computer program product comprising machine-readable instructions which, when executed by one or more computers, cause the one or more computers to perform the method (100, 300) according to any one of claims 1 to 11.

13. A non-transitory storage medium and / or downloadable product having a computer program which, when executed by one or more computers, causes the one or more computers to perform the method (100, 300) according to any one of claims 1 to 11.

14. A computer having the non-transitory storage medium and / or downloadable product according to claim 13.

Citation Information

Patent Citations

  • Rare instance classifiers

    US20180373963A1

  • Self-training method and system for semi-supervised learning with generative adversarial networks

    US20190122120A1