Training data between observation modalities

By generating synthetic measurement data using generative adversarial networks and recurrent neural networks, the problem of insufficient training data under different physical observation modalities in vehicle environmental detection systems is solved, improving the training efficiency and recognition accuracy of the machine learning module and enhancing the system's adaptability.

CN112529208BActive Publication Date: 2025-12-05ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010978926.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-18
Filing Date
2020-09-17
Publication Date
2025-12-05
Estimated Expiration
2040-09-17

AI Technical Summary

Technical Problem

In the existing technology, the vehicle environment detection system lacks training data under different physical observation modalities, resulting in poor application performance of the machine learning module under different observation modalities. In particular, it is difficult to effectively train and identify environmental objects when the image quality is poor or the sensor is replaced.

Method used

By developing a method to convert real or simulated data from the first physical observation modality into synthetic data through a generator, the gaps between different observation modalities are filled in by synthetic measurement data. Generative Adversarial Networks (GANs) and CycleGANs are used to optimize the training data, thereby achieving cross-modal data conversion and generating synthetic measurement data consistent with the nominal signal for training machine learning modules.

Benefits of technology

This improves the training efficiency and recognition accuracy of the machine learning module under different physical observation modalities, reduces the dependence on single-modal data, and enhances the system's adaptability and recognition ability in multiple environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112529208B_ABST
    Figure CN112529208B_ABST
Patent Text Reader

Abstract

The invention relates to a method for training a generator comprising the steps of: feeding at least one actual signal to the generator, the actual signal comprising at least one observed real or simulated physical measurement data from the first region; transforming the actual signal by the generator into a transformed signal, the transformed signal representing a synthetic measurement data belonging to the second region; evaluating by means of a cost function to which extent the transformed signal agrees with one / more reference signal, wherein the at least one reference signal is constituted by real or simulated measurement data of the second physical observation modality Mod_B for the situation represented by the actual signal; optimizing trainable parameters characterizing the behavior of the generator with the goal to obtain a transformed signal which is better evaluated by the cost function. The invention relates to a method for operating a generator. The invention relates to a method with a complete effect chain.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to generating training data for systems that derive decisions for behavior of vehicles and other technical systems from physical observations of an area. BACKGROUND

[0002] For vehicles to be able to move at least partially automatically in street traffic, it is necessary to detect the environment of the vehicle and, if a collision with an object in the vehicle's environment is imminent, to introduce countermeasures. The creation and localization of a representation of the environment is also necessary for safe automated driving.

[0003] In order to be able to derive decisions about further behavior of the own vehicle from physical observations of the vehicle's environment, machine learning modules are often used. Similar to a human driver who typically drives less than 100 hours and over less than 1000 km until he obtains a driver's license, a machine learning module can also generalize knowledge acquired from a limited reserve of training data to many other situations that are not the subject of training.

[0004] Since there is no single physical mapping modality that provides qualitatively high-quality measurement data in all conceivable situations and enables unambiguous decisions for further behavior, it is advantageous to set up several mapping modalities at the same vehicle. SUMMARY

[0005] In the scope of the present invention, a method for training a generator for synthetic measurement data is developed. The generator therefor converts real or simulated physical measurement data relating to an observation of a first area with a first physical observation modality Mod_A into synthetic measurement data. The measurement data recorded in the first modality Mod_A belongs to a first space X defined by the first modality Mod_A, i.e. the measurement data "exists" in this first space X. These synthetic measurement data relate to an observation of a second area with a second physical observation modality Mod_B, which at least partially overlaps with the first area. The synthetic measurement data belongs to a second space Y defined by the second modality Mod_B, i.e. the synthetic measurement data "exists" in this second space Y. In this sense, the expression "exists in a particular space" is also used in the following.

[0006] In particular, such modalities that provide the following signals, respectively, are suitable as observation modalities, i.e. from the signals it can be inferred that an object present in the environment of the vehicle is present and / or of a type. The observation modalities can in particular, for example, be based on different physical contrast mechanisms and thus complement each other as follows: namely by one or more other modalities to fill the gaps that remain in the observation performed with only one modality.

[0007] This in turn leads to the fact that the areas observed with the modalities Mod_A and Mod_B will not generally coincide completely, since each modality has its strengths and weaknesses. In particular, the resolution and range of action of the modalities Mod_A and Mod_B will not be identical.

[0008] All measured data that are recorded in reality with the modalities Mod_A and Mod_B can be preprocessed in any way. For example, the optical flow can be calculated from video images. For example, "synthetic aperture radar" images can be generated from radar data.

[0009] The main use of the synthetic measured data to be generated with the generator is to reduce the total expenditure of resources for training the machine learning modules that process the physical measured data detected in the modalities Mod_A and Mod_B. A large part of this total expenditure is spent on obtaining the training data. Thus, for example, human labor is used to move a measuring vehicle equipped with video cameras and / or other sensors through the traffic and to detect a large number of traffic situations with sufficient variability. It occurs very often nowadays that in the situation there is already training data relating to the first observation modality Mod_A, but it is desired to train a machine learning module to process training data acquired in the second observation modality Mod_B or also in the fusion of the two observation modalities Mod_A and Mod_B. For example, when testing a camera-based system to recognize objects in the environment of the vehicle, it can turn out that in certain situations the image quality is too poor and the desire can arise to teach the system to recognize objects from radar data as well. Likewise, although so far only training data for a camera-based system has been detected, the desire can arise to train a system that works on the basis of radar data alone. So far, it has not been possible to capitalize on the large number of camera images already present in this situation. The training method for the generator opens up a way for this.

[0010] In the context of this training method, at least one actual signal is fed to the generator, which comprises at least one observed real or simulated physical measurement data from the first region. The behavior of the generator is characterized by at least one set of trainable parameters. If the actual signal is thus converted into a transformed signal by the generator in a next step, which represents the synthetic measurement data belonging to it and thus exists in the space of the physical measurement data recorded in the second observation modality Mod_B, this takes place in accordance with the trainable parameters.

[0011] Evaluation with cost function: To what extent the transformed signal coincides with one or more target signals. Here, at least one target signal is constituted by real or simulated measurement data of the second physical observation modality Mod_B, which relates to the situation represented by the actual signal. The trainable parameters characterizing the behavior of the generator are optimized, which have the goal of obtaining a transformed signal that is better evaluated by the cost function.

[0012] After this training has ended, the generator has thus learned the mapping from the space X of the physical measurement data acquired by observation in the first physical observation modality Mod_A on the one hand and the space Y of the simulated measurement data equivalent to the measurement data obtained in the second physical observation modality Mod_B on the other hand. A large stock of physical measurement data recorded in the first physical observation modality in the space X can thus be used for obtaining new training data, which exist in the space Y of the measurement data acquired in the modality Mod_B. The new training data in the space Y are thus only still to a lesser extent than hitherto represent a bottleneck for training the machine learning module, the goal of which is to no longer base the sought conclusions in the respective application context only on the measurement data acquired in the modality Mod_A, but also on the measurement data acquired in the modality Mod_B. Thus, for example, the recognition of objects with the machine learning module can no longer be based only on images, but alternatively or also in combination therewith, for example, on radar data.

[0013] The measurement data generated with the physical observation modalities Mod_A and Mod_B typically exist in very different spaces, whose dimensions also have completely different semantic or physical meanings, respectively. Thus, for example, images mostly exist as a set of pixels, which each state an intensity value and / or a color value for a specific location. Radar data, for example, exist as a set of radar spectra or as reflections, which can each be assigned a direction in the form of one or more angles, a distance, an intensity and, optionally, also a velocity. For example, the range of action of a radar is also significantly greater than that of a camera, while, on the other hand, the angular resolution is smaller than in the case of a camera. It is all the more surprising that a partial overlap of the regions observed with the two modalities is already sufficient for the generator to be able to learn a transformation between these very different spaces.

[0014] The transformed signals can be checked in various ways to what extent they agree with one or more of the target signals. For example, a cost function can evaluate to what extent the learning actual signals from a pre-given source are mapped by the generator to the pre-given target signals that fit them, respectively. However, this assignment by the generator is not a mandatory condition for the synthetic measurement data generated by the generator to be usable for further training of the machine learning module equivalently to the physical measurement data obtained by actual observation in the second physical observation modality Mod_B. Rather, it is sufficient, for example, if the synthetic measurement data cannot be distinguished from the measurement data obtained in modality Mod_B, which exists in the same space Y.

[0015] Thus, in one particularly advantageous configuration, the cost function comprises a GAN term, wherein the transformed signals the better value the GAN term takes on, the more the transformed signals cannot be distinguished from the pre-given amount of target signals according to the discriminator module. The discriminator module is additionally trained to distinguish the transformed signals from the target signals.

[0016] In this configuration, the generator G is embedded in a "conditional generative adversarial network", shortly "conditional GAN". The generator G obtains as input the actual signals x and, optionally, also a sample z drawn from a (e.g., normally distributed) multi-dimensional random variable, and tries to generate from each actual signal x a transformed signal y' that is as indistinguishable as possible from the assigned target signal y. In this context, "conditional" means that the generator G maps the input x (optionally together with the sample z) to an output that is related to the same scene The upper. Thus, not for example all actual signals x are mapped onto the same output y' which is most difficult to distinguish from a given amount of nominal signals y. The discriminator module D is only needed during training and is not used anymore later on when the training data is produced with the generator G or even later on when the training data is used for training the actual machine learning module for the set application.

[0017] The GAN term can for example take the form :

[0018] .

[0019] Here, E x,y denotes the expected value ("sample mean") with respect to the pairs x and y. Correspondingly, E x,z denotes the expected value with respect to the pairs x and z. The generator G strives to minimize while the discriminator D strives to maximize . The optimal generator G* is then the solution of the optimization problem:

[0020] .

[0021] Here, max D denotes the maximization over the parameters of the discriminator D. Correspondingly, argmin G denotes the minimization over the parameters of the generator G.

[0022] In another particularly advantageous configuration, the cost function additionally contains a similarity term, where the more similar the transformed signal is to the nominal signal according to a pre-given measure, the better value the similarity term takes. This also more strongly counteracts the possible tendency of the generator to simply optimize the indistinguishability according to numerical criteria and for this to seek "simple ways" which are not goal-oriented (zielführend) for the set application. If for example a certain type of object is particularly difficult to convert from the representation in the form of an optical image into the representation in the form of radar data, the generator cannot simply "cheat" (erschleichen) a better value by simply making this complex object disappear. This is counteracted by the similarity term. An example for a similarity term is

[0023] .

[0024] In this case, in principle also every measure different from the LI measure can be used.

[0025] In another advantageous configuration, the cost function additionally contains an application term which measures a desirable property of the transformed signal itself for the intended application. This application term is also referred to as "Perceptual Loss" L P (G). The application term is not limited to the fact that it must only depend on the final result of the transformation. Rather, the application term can also depend on intermediate results of the transformation, for example, if the generator comprises a multi-layer neural network. The intermediate results can also be extracted on the hidden layers between the input layer and the output layer.

[0026] The application term can measure, for example, whether the scene represented by the transformed signal is plausible in the sense of the respective application. Thereby, for example, transformed signals are discarded according to which a car is three times as high or as wide as usual or moves at almost sonic speed in the innerorts. Alternatively or in combination therewith, a comparison with a nominal signal on an abstract level can also be in the application term, for example. Thus, for example, a representation generated from the transformed signal by an autoencoder or a further KNN can be compared with a representation generated from the nominal signal by the same KNN.

[0027] With the similarity term and the application term, the optimization problem can be written overall, for example, as

[0028] .

[0029] Herein, and are hyperparameters which weight the different cost function terms.

[0030] In another particularly advantageous configuration, a training inverse generator module is used to transform the transformed signal back into a signal of the actual signal type. That is, the actual signal back into which the transformation takes place exists in the same space as the original actual signal. The cost function then additionally contains an inverse GAN term. The inverse GAN term takes on better values the more indistinguishable the signal back into which the transformation takes place is from the actual signal according to a further discriminator module.

[0031] The further discriminator module is trained to distinguish the signal back into which the transformation takes place from the actual signal. The cost function additionally also contains a consistency term. The consistency term is a measure for how identically the actual signal is regenerated when the transformation takes place by the generator module and the transformation back takes place by the inverse generator module.

[0032] In this type of training case, the architecture into which the generator is integrated is extended from the "conditional GAN" to the CycleGAN. The main advantage is that the rating signal no longer has to relate to the same scenario as the actual signal or the actual representation. A great advantage of the CycleGAN is that it can translate data between domains, which are characterized by unpaired quantities of examples, respectively.

[0033] This can make the training significantly simpler and cheaper, especially when the generator should translate the actual signal not only into the space Y of the measurement data recorded in the second physical observation modality Mod_B, but also into spaces belonging to other observation modalities. The more modalities should be considered, the more difficult it becomes to record the same scenario in all modalities. Even if something has changed at the recorded scenario, for example, the data stock recorded previously in the first physical observation modality Mod_A, for example, a video camera, is not suddenly "devalued" for the further training of the generator. Thus, for example, a street image can have changed as follows permanently: namely, after the original video camera recording, the street has been coated over a longer road section with a differently textured, low-noise asphalt, or the carriageway has been marked as an "environmental carriageway" in a visually conspicuous manner over a longer road section. The larger the area in which the data is recorded, the more difficult it becomes to document such changes.

[0034] Furthermore, even if the sensor used for the physical data recording in the second observation modality Mod_B is changed, it is still possible to continue using the data recorded in the first physical observation modality Mod_A completely. For example, a radar sensor used first for testing can prove to be unsuitable or too expensive later for series use in a vehicle. After the radar sensor has been replaced, it is sufficient to record new radar signals in the space Y with the new radar sensor. However, it is not necessary to record new images in the space X with the likewise installed video camera for this motivation.

[0035] As with the "conditional GAN", the CycleGAN learns a mapping G from the space X in which the actual signals exist to the space Y in which the generated synthetic measurement data y exists. In addition, a reverse mapping F from the space Y to the space X is learned. A first discriminator D x is learned, which attempts to distinguish between the generated data F(y) and the real actual signals x. A second discriminator D y is learned, which attempts to distinguish between the generated data G(x) and the real rating signals y. This can be expressed, for example, in a cost function term:

[0036] and

[0037] .

[0038] Here, z1 and z2 are samples of the random variables Z1 and Z2. The use of random variables Z1 and Z2 is optional.

[0039] Monitoring compliance with the consistency condition and Exemplary consistency terms that monitor compliance with the consistency condition are:

[0040]

[0041] The entire cost function for CycleGAN can then be written, for example, as:

[0042]

[0043] The cost function can also be extended with an application term L P that now depends on G as well as F: For example, this term can be added with a weight

[0044] Likewise, a similarity term can be added to the cost function for CycleGAN. Unlike in cGAN, there are now two terms for both generators G and F:

[0045] and

[0046]

[0047] For example, these terms can be added with a weight

[0048] In another particularly advantageous configuration, hyperparameters are optimized according to a predefined optimization criterion, which determine the relative weighting of the terms in the cost function to each other. These hyperparameters represent further degrees of freedom with which the generators can be adapted to a specific task. For example, a search space spanned by a plurality of hyperparameters can be searched in a predefined raster. This does not presuppose that the optimization criterion depends on the hyperparameters in a continuous manner.

[0049] As previously set out according to the formula, in another particularly advantageous configuration, at least one actual signal comprises not only real or simulated physical measurement data of the first physical observation modality Mod_A but also samples drawn from a random variable. For example, the samples can be added to the measurement data. The noise added in this way has a double effect: On the one hand, many other variants can be produced from a predefined stock of actual signals in order to increase the variability of the training. On the other hand, other features in the latent space can also be learned. ​​​​​

[0050] In another particularly advantageous configuration, the actual signal is selected which assigns at least one actual label to at least one component (Anteil) of real or simulated physical measurement data of the first physical observation modality Mod_A.

[0051] The label can assign an arbitrary statement from an arbitrary source to the measurement data. In the context of a classification task, the label may, for example, represent a class of an object to which the measurement data points. In the context of a regression task, the label may, for example, represent a regression value (Regressionswert) related to an object indicated by the measurement data. The regression value may, for example, be a distance, an extension or an orientation of an object. The data equipped with a label (“labeled”) is usually used in the context of supervised training of a machine learning module. In this context, the label often represents a statement which the machine learning module should derive from the labeled data (such as real or simulated physical measurement data of the first physical observation modality Mod_A) after the training has been completed.

[0052] For example, the ultimate goal pursued with the physical measurement data of the lifting modality Mod_A can be to classify objects with a machine learning module, wherein the measurement data indicates the presence of the objects. The label may, for example, then state for real or simulated images of a vehicle environment which objects, such as lane boundaries, traffic signs or other road users, are present in the scene reproduced in the image. The machine learning module can then be trained with a pre-given amount of images and the associated labels to also correctly identify the contained objects in images of unknown conditions.

[0053] As one possible implementation, the actual signal of Mod_A can be supplemented with further channels containing label information. For example, in the case of video images, labels can be added as a per-pixel semantic segmentation. For radar spectra (for example in range-velocity images), this would correspondingly be an additional channel per (multi-dimensional) FFT-bin. In the case of point-like radar data (reflections), labels can be added as an additional attribute per reflection point. Here, the labels can contain not only a class, but also one or more regression values (for example, distance, extension, orientation of an object) - one channel per regression value. This is advantageous in the configuration, in particular when the modality can estimate such parameters particularly precisely. Exemplarily, a radar can measure the distance of an object directly and transfer this to the pixel space of a video image. The labeled data can be used for other machine learning algorithms.

[0054] The assignment of labels to real or simulated physical measurement data in many cases requires manual work and thus can require a large part of the total cost of a system that classifies objects, for example, with the aid of a machine learning module. It has now been recognized that the training method described previously not only enables the generator to transform real or simulated physical measurement data of the first modality Mod_A in the space X into synthetic measurement data of the second modality Mod_B in the space Y. Rather, the assignment of labels to real or simulated physical measurement data contained in the actual signals can also be at least partially transferred together in a variety of ways into the space Y, so that the investment in labeling in the space X can continue to be used even when the measurement data of the second modality Mod_B in the space Y are processed later.

[0055] In another particularly advantageous configuration, therefore, when training the generator, a target signal is selected which assigns at least one target label to at least one component of the real or simulated physical measurement data of the second physical observation modality Mod_B. That is, the training method starts from the premise that not only the training data in the space X, but also the training data in the space Y, are labeled.

[0056] For example, the real or simulated physical measurement data of the second modality Mod_B can contain radar data, i.e. information characterizing radar reflections. The labels can then, for example, state which objects are present in a scene represented with radar reflections. If a machine learning module is trained with these training data in the space Y, this machine learning module can also determine for an unknown radar reflection constellation: which objects the radar reflections point to. This is a classification task in the broadest sense.

[0057] The actual labels contained in the actual signals are transformed by the generator into transformed labels that exist in the target label space. The cost function with which the generator is trained now contains a label term, which takes on better values the better the generated labels agree with the target labels.

[0058] The generator thus learns on the basis of labeled training data in the space X and labeled training data in the space Y in order to produce labeled synthetic measurement data of the second modality Mod_B in the space Y from labeled real or simulated measurement data of the first modality Mod_A in the space X.

[0059] The specific form of the label term in the cost function can depend, for example, on the task for which a machine learning module should be trained with the synthetic measurement data produced by the generator. For example, pixel-wise cross-entropy can be used for a classification task. For example, mean squared error can be used for a regression task.

[0060] The generator may, for example, comprise and / or be a KNN, an artificial neural network. The KNN has a large number of neurons and / or other processing units. The neurons and / or other processing units sum their respective inputs in a weighted manner according to trainable parameters of the generator and constitute their output by applying a non-linear activation function to the result of the weighted sum.

[0061] Here, it is particularly advantageous that the number of neurons and / or further processing units per layer monotonically decreases in the first sequence of layers and monotonically increases in the second sequence of layers. Thereby, a "bottleneck" is formed between the end of the first sequence of layers and the beginning of the second sequence of layers, in which there is an intermediate result having a significantly reduced dimensionality compared to the actual signal input. This "bottleneck" enables the KNN to learn and to compact the relevant features. Thereby, a better performance can be achieved and the computational effort is reduced.

[0062] In another particularly advantageous configuration, the KNN has at least one direct connection between a first layer from the first sequence of layers and a second layer from the second sequence of layers. In this way, specific information can be selectively directed past the "bottleneck", such that the information content in the transformed signal is generally increased. The direct connection is thus understood in particular as a connection that bypasses at least one layer from the first and / or second sequence of layers, which would otherwise have to be passed through.

[0063] When the generator has been trained first, a final state is embodied in a parameter set with parameters characterizing its behavior. In the case of a KNN, these parameters may, for example, include weights with which inputs delivered to a neuron or further processing unit are calculated for activating the neuron or the processing unit. The parameter set enables the generator to be arbitrarily reproduced without further training and is thus a product that can be sold independently.

[0064] The invention also relates to a method for operating a generator with which synthetic measurement data of a second observation modality Mod_B in a space Y can be generated from labeled real or simulated physical measurement data of a first observation modality Mod_A in a space X and can subsequently be labeled.

[0065] In the case of the method, at least one actual signal of real or simulated physical measurement data having a first physical observation modality Mod_A is transformed by the generator into at least one transformed signal. For this transformed signal, quantitative contributions are determined, wherein different components of the real or simulated physical measurement data of the first physical observation modality Mod_A make said quantitative contributions to the transformed signal. The components can be of arbitrary granularity (granular), down to individual observations, such as image recordings or parts thereof, performed with the first modality Mod_A.

[0066] For the different components, actual labels are determined separately. These actual labels can be, for example, identical to the actual labels exhibited when training the generator according to the previously described method. However, this is not a mandatory requirement. It is not even mandatory that actual labels are exhibited at all when training the generator. It is only important that at the point in time at which the components of the real or simulated physical measurement data and their quantitative contributions to the transformed signal are identified, actual labels are available for the components.

[0067] At least one label for the transformed signal is determined from the quantitative contributions in combination with the actual labels.

[0068] The logic behind this is that, since the regions observed in the two modalities Mod_A and Mod_B at least partially overlap, the same features of objects present in these regions make particularly important contributions to the observations made with the two modalities Mod_A and Mod_B, respectively. Therefore, if a particular feature stands out (hervorstechen) in the observations made with the first modality Mod_A on the one hand and in the observations made with the second modality Mod_B on the other hand, respectively, it can be assumed that these features originate from the same object. If a label is known for this object in the space X in which the observations with Mod_A were made, it can be transferred (übertragen) to the corresponding observations in space Y.

[0069] For example, a stop sign in a camera image stands out as an octagonal red surface with white lettering. If this camera image is transformed by the generator into radar data, these radar data will likewise be at least partially shaped by the reflections assignable to the octagonal object. This can be traced back to the space X in which the camera image was recorded, so that the label "stop sign" assigned there can be transferred to the octagonal object appearing in the synthetic radar image.

[0070] As previously described, the generator can in particular have and / or be a KNN with a large number of neurons and / or other processing units. The neurons and / or other processing units can in particular be wired such that they sum their respective inputs in a weighted manner according to trainable parameters of the generator and constitute their output by applying a non-linear activation function to the result of this weighted sum. In one particularly advantageous configuration, in which the transformed signal is taken as a starting point in the case of an architecture using a KNN, it can be determined to what extent components of real or simulated physical measurement data of the first physical observation modality Mod_A contribute to the rejection (Ausschlag) of at least one activation function. To this end, for example, backpropagation through the KNN can be used. In this way, a link between the quantitative contribution to the transformed signal and the existing label can be established.

[0071] Here, it is entirely possible that the transformed signal contains a plurality of outstanding features which can be reduced to measurement data with different labels (for example camera images showing different traffic signs). A label for the transformed signal which is representative of a class can then advantageously be determined in such a way that a majority vote is determined under actual labels which are likewise representative of a class. The label for the transformed signal which is related to a regression value can be determined from a combined function of actual labels which are likewise related to a regression value. For example, the combined function can constitute a mean or median value.

[0072] With a trained generator and actual labels for specific real or simulated physical measurement data of the mapping modality Mod_A in the space X as a starting point, it is possible with a further method provided by the present application to equip real physical measurement data obtained from a second physical observation modality Mod_B in the space Y with a label. Here, it is advantageous to use measurement data in the spaces X and Y which are related to the same point in time, i.e. are physically recorded simultaneously, for example. In this way, it can be ensured that the measurement data are related to the same scene without changes during this period.

[0073] In the case of this method, at least one actual signal with real or simulated physical measurement data of the first physical observation modality Mod_A is transformed into at least one transformed signal by means of the generator, for which at least one actual label is available.

[0074] At least one label for the transformed signal is provided. This label can be produced, for example, by the generator which has learned to label together with the transformed signal. However, the label can also be added retrospectively, for example using the previously described method for operating the generator.

[0075] The transformed signal is compared with at least one other signal, which comprises real or simulated physical measurement data of the second physical observation modality Mod_B. From the result of the comparison, at least one label for the other signal and / or a spatial offset between the two physical observation modalities Mod_A and Mod_B is determined in combination with the labels for the transformed signal.

[0076] Thus, if a signal recorded in real life with the second physical observation modality Mod_B, for example, is similar to a transformed signal obtained for an object with a known label, it can be concluded from this that the signal recorded in real life with modality Mod_B originates from precisely this object and can be assigned the known label. In this case, the machine learning module can be trained with real measurement data recorded with modality Mod_B, using labels that have already been adopted from modality Mod_A. The use of real measurement data instead of synthetic measurement data generated with a generator has the advantage that possible artefacts that can occur when generating synthetic measurement data do not influence the subsequent training of the machine learning module for the final application.

[0077] For example, let x be a real signal that exists in the space X of measurement data provided by modality Mod_A and whose actual label is known. The generator translates x into a transformed signal y', which exists in the space Y of measurement data provided by modality Mod_B. By at least one label being available for x, a label is then also available for y'. Let y now be a further signal recorded in real life with the real modality Mod_B. Depending on the implementation, it can be possible to transfer the label from y' directly to y. Here, additional steps are usually required for determining an offset between the synthetic y' and the real y representation. This can be determined, for example, by comparing the data y' and y, also over a plurality of measurements, for example in terms of a correlation measure. The label of y' can be transferred to y, taking into account the offset.

[0078] Because the sensors used according to the two modalities Mod_X and Mod_Y have different ranges of action, resolutions, etc., not all data in Y can usually be labelled. For example, a video camera has a much smaller range of action than a radar. Whereas the radar has a lower angular resolution, i.e. two objects close to one another cannot be separated. Shielding situations can also often be resolved by the radar, whereas a video sensor only detects objects in front. It is thus possible that the label is not transferred onto the entire space of Mod_B, but only onto a certain overlap region of the modalities.

[0079] This approach has the advantage that no explicit common coordinate representation or transformation between Mod_A and Mod_B is needed. The coordinate transformation is implicitly learned together while training. Such coordinate transformations are often not trivial; for example, a video camera measures in pixel space, while a radar measures in range and angle space. Additionally, a video camera cannot measure the radial relative velocity of objects, which is an important component of radar measurements.

[0080] The determination of the offset (Versatz) can also be used for calibrating the observation modalities Mod_A and Mod_B as such (für sich genommen), i.e. to eliminate the adverse effects of the spatial offset between these modalities Mod_A and Mod_B.

[0081] If, for example, a plurality of sensors are installed in a vehicle, the calibration of the sensors with respect to each other (external calibration) is usually not precisely known and must first be measured in a costly manner. The knowledge of the sensor calibration is indispensable in order to build an environmental model from the different sensors.

[0082] It is desirable to determine the external calibration automatically online, i.e. after the sensors have been installed.

[0083] In order to create the training data, the modalities Mod_A and Mod_B are installed in a test vehicle and precisely measured with respect to each other, so that the external calibration is known. The generator is trained as previously described to generate data Y_ of Mod_B from data X of Mod_A.

[0084] In this application, the same modalities Mod_A and Mod_B are installed in the vehicle, but the precise calibration of the sensors used is unknown. In order to distinguish them from the modalities used in the training, they are denoted Mod_A2 and Mod_B2. For example, a video or radar sensor can be installed slightly offset or slightly twisted. The recorded data x_2 of Mod_A2 is used as input for the trained generator in order to generate data y_2' of Mod_B2. The data y_2' corresponds to the calibration of the training data, which has been precisely measured since the generator learned this transformation together.

[0085] The offset between the synthetic data y_2' of Mod_B2 and the measured data y_2 can be determined via the method described above. If this offset is added to the measured calibration, the calibration of the now imprecisely installed modalities Mod_A2 and Mod_B2 with respect to each other is known.

[0086] The invention also provides a further method. This method comprises the entire effect chain from the trained generator up to the piloting technology system.

[0087] In the case of this method, the generator is trained as previously described. With the trained generator, at least one synthetic signal of the second observation modality Mod_B is generated from actual signals of real or simulated measurement data of the first observation modality Mod_A. The machine learning module is trained with the synthetic signals.

[0088] With at least one sensor, physical measurement data of the second observation modality Mod_B from the environment of the vehicle are recorded. The trained machine learning module is run in such a way that it obtains as input the physical measurement data provided by the sensor and maps this onto at least one class and / or at least one regression value. From the class and / or from the regression value, a maneuver signal is determined. With the maneuver signal, the vehicle is maneuvered.

[0089] The construction of the maneuver signal can in this case in particular include the check whether, in connection with the current or planned trajectory of the own vehicle, based on the result of the analysis, it should be a matter of concern that the trajectory of an object in the environment of the own vehicle intersects the current or planned trajectory of the own vehicle. If this is the case, the maneuver signal can in particular focus on altering the trajectory of the own vehicle so that it no longer intersects the trajectory of the identified object.

[0090] In particular, the method can be computer-implemented in whole or in part. The invention therefore also relates to a computer program having machine-readable instructions which, when implemented on one or more computers, cause the one or more computers to implement one of the described methods. In this sense, a control device for a vehicle and an embedded system for a technical device which are likewise capable of implementing machine-readable instructions can also be regarded as computers.

[0091] Likewise, the invention also relates to a download product and / or a machine-readable data carrier having a parameter set and / or having a computer program. The download product is a digital product which can be transmitted via a data network, i.e. can be downloaded by a user of a data network, which can for example be sold in an Online-Shop for immediate download.

[0092] Furthermore, a computer can be equipped with a parameter set, with a computer program, with a machine-readable data carrier or with a download product.

[0093] Further measures for improving the invention are depicted in more detail below together with the description of preferred embodiments of the invention according to the figures. BRIEF DESCRIPTION OF DRAWINGS

[0094] Figure 1 An embodiment of the training method 100 is shown;

[0095] Figure 2 An exemplary neural network 14 for use in the generator 1 is shown;

[0096] Figure 3 A diagram showing how a transformation of actual labels 11a to transformed labels 21a can be trained;

[0097] Figure 4 An embodiment of a method 200 for operating the generator 1 is shown;

[0098] Figure 5 A diagram showing Figure 4 the method shown in Fig. 2;

[0099] Figure 6 An embodiment of a method 300 for operating the generator 1 is shown;

[0100] Figure 7 A diagram showing Figure 6 the method shown in Fig. 3;

[0101] Figure 8 An embodiment of a method 400 with a full effect chain. DETAILED DESCRIPTION

[0102] Figure 1 is a flow chart of an embodiment of a training method 100. The method 100 starts with the fact that a region 10 can be observed in a first physical observation modality Mod_A. At the same time, a region 20 can be observed in a second physical observation modality Mod_B, which region 20 overlaps with the region 10. Physical measurement data 10a can be obtained by real or simulated observation of the region 10. Physical measurement data 20a can be obtained by real or simulated observation of the region 20.

[0103] In step 105, a generator 1 with an artificial neural network KNN 1 is selected. In step 110, an actual signal 11 with real or simulated measurement data 10a in the modality Mod_A is fed to the generator 1, which, according to block 111, can be equipped with one or more labels 11a. In addition to the real or simulated measurement data 10a in the modality Mod_A, the actual signal 11 can also contain, for example, metadata taken (erhoben) together with the measurement data 10a. Such metadata can for example include the measurement instrument used, parameters or settings of the measurement instrument, for example a camera or a radar device.

[0104] In step 120, the actual signal is transformed into a transformed signal 21 by the generator 1. As long as the tag 11a is present, it can be transformed into a transformed tag 21a by the generator 1. In step 130, it is checked according to a cost function 13 to what extent the transformed signal 21 coincides with at least one target signal 21'. In step 140, trainable parameters 1a characterizing the behavior of the generator 1 are optimized in such a way that the evaluation 130a by the cost function 13 presumably can have a better result for the then obtained transformed signal 21.

[0105] It is exemplarily depicted within block 13 how the evaluation 130a can be determined. According to block 131, at least one target signal 21' is selected, which is equipped with a target tag 21a'. The cost function 13 can thus contain a comparison between the transformed tag 21a and the target tag 21a'.

[0106] According to block 141, in addition to the generator 1, a discriminator module is simultaneously trained to distinguish the transformed signal 21 from the target signal 21' in order to thereby create an additional incentive for progress when training the generator 1.

[0107] The parameters 1a occurring at the end of the training determine the trained state 1* of the generator 1.

[0108] Figure 2 An exemplary KNN 14 that can be used in the generator 1 is schematically shown. In this example, the KNN 14 consists of seven layers 15a-15g, which respectively comprise neurons or other processing units 16. Here, the layers 15a-15c form a first layer sequence 17a, in which the number of neurons 16 of each layer 15a-15c monotonically decreases. The layers 15e-15g form a second layer sequence 17b, in which the number of neurons 16 of each layer 15e-15g monotonically increases. Between them, there is a layer 15d in which there is a maximally reduced representation of the actual signal 11. The KNN 14 additionally contains three direct connections 18a-18c between the layers 15a-15c from the first layer sequence 17a and the layers 15e-15g from the second layer sequence 17b, which in the example shown in Fig. 1 have the same number of neurons 16, respectively. Figure 3

[0109] Figure 3 It is schematically illustrated how the generator 1 can be trained to generate a new tag 21a in a space Y of observations with a modality Mod_B from a tag 11a present in a space X of observations with a modality Mod_A.

[0110] In Figure 3 ​In the example shown, the camera image exists in space X in a Cartesian coordinate system with coordinates a and b. Exemplarily, in Figure 3 In the environment 50a of the vehicle 50 (not shown in the image), pedestrians 51 and external vehicles 52 are shown as observations 10a. Information relating to pedestrians 51 or vehicles 52 constitutes their respective labels 11a.

[0111] If the same scene is observed using radar as Mod_B, then observation 20a is, for example, Figure 3 The radar spectrum depicted in the diagram, or alternatively or in combination with it, is, for example, radar reflection. The radar spectrum is represented by coordinates... (Angle) and The distance exists in space Y. The radar spectrum can be reassigned as respective nominal tags 21a' to information related to pedestrians 51 or vehicles 52.

[0112] According to step 120 of method 100, the actual tag 11a is transformed into a transformed tag 21a in space Y. As previously described, cost function 13 can then be used to examine the extent to which the transformed tag 21a corresponds to the nominal tag 21a'. Figure 3 In the instantaneous record shown, the consistency remains very poor. Cost function 13 then becomes the driving force for optimizing parameter 1a of the generator, with the objective of improving consistency.

[0113] Figure 4 This is a flowchart of an embodiment of method 200 for operating generators 1 and 1*. In step 205, generator 1 containing KNN 14 is selected. In step 210, at least one actual signal 11 is converted into a transformed signal 21 using generators 1 and 1*.

[0114] In step 220, quantitative contributions 22a-22c are determined, wherein different components 12a-12c of the real or analog physical measurement data 10a contained in the actual signal of modality Mod_A make the quantitative contributions to the transformed signal 21. In step 230, actual labels 12a*-12c* are determined for the different components 12a-12c. In step 240, at least one label 21* for the transformed signal 21 is determined from the contributions 22a-22c and the actual labels 12a*-12c*.

[0115] Within block 240, two possibilities of how the label 21* can be determined are exemplarily drawn. According to block 241, the multiple actual labels 12a*-12c* representing a class can be aggregated by majority voting with respect to the class. According to block 242, the multiple actual labels 12a*-12c* related to regression values can be aggregated by a combining function.

[0116] In Figure 5 the method 200 is explained in more detail. The spaces X and Y are the same as in Figure 3 and the same objects are present. However, differently to Figure 3 it is assumed that in space Y the ground truth labels 21' are not available for the observations 20a in the transformed signal 21.

[0117] For labeling the observations 20a in the transformed signal 21, quantitative contributions 22a-22c are determined, in which the components 12a-12c of the measurement data 10a in space X have contributed to the observations 20a. The components 12a-12c are regions (Gebiete) in the illustration chosen in Figure 5 According to step 221 of the method 200, the architecture of the KNN 14 is used, e.g. in the way of backpropagation through the KNN 14, in order to determine the contributions 22a-22c and the components 12a-12c.

[0118] As shown in Figure 5 , if these regions contain the labeled objects 51, 52 in X, these actual labels 12a*-12c* can be used in order to determine the label 21* for the transformed signal 21. In the situation shown in Figure 5 , each of the components 12a-12c comprises only a single labeled object 51 or 52, so that the respective label can be determined directly as the label 21* for the transformed signal 21.

[0119] Figure 6An embodiment of a method 300 for operating the generator 1 is shown. In step 310, an actual signal 11 of an observation 10a having a first modality Mod_A is transformed by the generator 1, 1* into a transformed signal 21, which is equipped with actual labels 11a. In step 315, at least one label 21a, 21* for the transformed signal 21 is determined, which can be done directly at the time of generating the transformed signal 21 by the labelling generator 1 (label 21a) or also, for example, retrospectively with the previously described method 200 (label 21*). In step 320, the transformed signal in the space Y of the modality Mod_B is compared to other signals 21b comprising real or simulated measurement data 20a of the modality Mod_B. Similar to the signal 11, the signals 21b can comprise, for example, metadata taken together with the measurement data 20a in addition to the measurement data 20a. In step 330, a label 21b* for the other signals 21b and / or a spatial offset between the observation modalities Mod_A and Mod_B is analysed from the label 21* for the transformed signal 21 in connection with the result 320a of the comparison 320.

[0120] Figure 7 The method shown in Figure 6 is illustrated in the diagram. For better intelligibility, the Figure 3 and 5 differences, here the space Y in which the transformed signal 21 exists is a Cartesian space in the coordinates a and b.

[0121] In the example shown in Figure 7 , the transformed signal 21, which is drawn as a rectangle and as a circle, corresponds quite well to the other signals 21b obtained by measurements in the space Y, for which a label 21* was determined respectively. Thereby, in steps 320 and 330 of the method 300, it can be concluded respectively that the other signals 21b respectively relate to the same object as the transformed synthetic signal 21. In correspondence therewith, as a label 21b* for these other signals 21b, the label 21* of the transformed signal 21 can be taken up again. At the same time, a spatial offset Δ between the observation modality Mod_A on the basis of which the transformed signal 21 was obtained and the observation modality Mod_B on the basis of which the other signals 21b were obtained can be determined in this way.

[0122] The label 21* of the transformed signal 21 is based on the original actual labels 11a from the space X. For the sake of clarity, these are not drawn in Figure 7 .

[0123] Figure 8An embodiment of the method 400 is shown which has the complete effect chain from the training of the generator 1 up to the steering of the vehicle 50.

[0124] In step 410, the generator 1 is trained with the previously described method 100 and thereby reaches its trained state 1*. In step 420, the actual signal 11 with real or simulated measurement data 10a of the modality Mod_A in the space X is transformed into the synthetic (transformed) signal 21 of the modality Mod_B in the space Y with the trained generator 1*. In step 430, the machine learning module 3 is trained on the basis of the synthetic signal 21 and thereby reaches its trained level (Stand) 3*. In parallel to this, in step 440, physical measurement data 20a of a second observed modality from the environment 50a of the vehicle 50 is recorded with at least one sensor 4.

[0125] The trained machine learning module 3* is run in step 450 in such a way that it obtains the measurement data 20a as input according to block 451 and maps said measurement data 20 onto at least one class 450a and / or onto at least one regression value 450b according to block 452. In step 460, the steering signal 360a is determined from the class 450a and / or from the regression value 450b. In step 470, the vehicle 50 is steered with the steering signal 460a.

Claims

1. A method (100) for training a generator (1) to convert real or simulated physical measurement data (10a) related to an observation of a first region (10) having a first physical observation modality Mod_A into synthetic measurement data (20a) related to an observation of a second region (20) having a second physical observation modality Mod_B, wherein the first region (10) and the second region (20) at least partially overlap, the method having the steps of: • feeding (110) the generator (1) with at least one actual signal (11) comprising real or simulated physical measurement data (10a) from at least one observation of the first region (10); • transforming (120) by the generator (1) the actual signal (11) into a transformed signal (21) representing synthetic measurement data (20a) belonging thereto; • evaluating (130) by means of a cost function (13) to which extent the transformed signal (21) coincides with one or more target signals (21'), wherein at least one target signal (21') consists of real or simulated measurement data of the second physical observation modality Mod_B for a situation represented by the actual signal (11); • optimizing (140) trainable parameters (1a) characterizing the behavior of the generator (1) with the goal to obtain a transformed signal (21) which is better evaluated by the cost function (13), wherein the measurement data produced in the physical observation modalities Mod_A and Mod_B exist in different spaces, the dimensions of which also have completely different semantics or physical meanings, respectively.

2. The method (100) according to claim 1, wherein • the cost function (13) comprises a GAN term, wherein the GAN term takes a better value the more indistinguishable the transformed signal (21) is from a pre-given amount of target signals (21') according to a discriminator module; and • the discriminator module is additionally trained (141) to distinguish transformed signals (21) from target signals (21').

3. The method (100) according to claim 2, wherein the cost function (13) additionally comprises a similarity term, wherein the similarity term takes a better value the more similar the transformed signal (21) is to the target signal (21') according to a pre-given measure.

4. The method (100) according to any one of claims 2 to 3, wherein the cost function (13) additionally comprises an applicability term, which measures a desirable property of the transformed signal (21) itself for an intended application.

5. The method (100) according to any one of claims 2 to 3, wherein • training (142) an inverse generator module to transform the transformed signal (21) back into a signal of the type of the actual signal (11), wherein the cost function (13) additionally contains an inverse GAN term, wherein the more indistinguishable the transformed back signal is from the actual signal (11) according to a further discriminator module, the better value the inverse GAN term takes; • training (143) the further discriminator module to distinguish the transformed back signal from the actual signal (11); and • the cost function (13) contains a consistency term which is a measure of how identically the actual signal (11) is regenerated when transformed through the generator (1) and transformed back through a further generator.

6. The method (100) according to any one of claims 2 to 3, wherein the hyperparameters (13a) determining the relative weighting of the terms in the cost function (13) from each other are optimized according to a pre-given optimization criterion.

7. The method (100) according to any one of claims 1 to 3, wherein at least one actual signal (11) comprises not only real or simulated physical measurement data (10a) of the first physical observation modality Mod_A but also samples drawn from random variables.

8. The method (100) according to any one of claims 1 to 3, wherein an actual signal (11) is selected (111) which assigns at least one actual label (11a) to at least one component of real or simulated physical measurement data (10a) of the first physical observation modality Mod_A.

9. The method (100) according to claim 8, wherein • at least one nominal signal (21') is selected (131) which assigns at least one nominal label (21a') to at least one component of real or simulated physical measurement data (20a) of the second physical observation modality Mod_B; • the actual label (11a) is transformed (120) by the generator (1) into a transformed label (21a) which exists in the space of the nominal label (21a'), and • the cost function (13) contains a label term which takes a better value the more the generated label (21a) coincides with the nominal label (21a').

10. The method (100) according to any one of claims 1 to 3, wherein • a generator (1) is selected (105) which comprises and / or is at least one artificial neural network, KNN (14); and • the KNN (14) has a number of neurons and / or other processing units (16) which sum their respective inputs in a weighted manner according to trainable parameters (1a) of the generator (1) and constitute their output by applying a non-linear activation function to the result of this weighted sum.

11. The method (100) of claim 10, wherein the KNN (14) is built layer by layer, and wherein the number of neurons and / or further processing units (16) per layer (15a-15g) monotonically decreases in a first sequence of layers (17a) and monotonically increases in a second sequence of layers (17b).

12. The method (100) of claim 11, wherein the KNN (14) has at least one direct connection (18a-18c) between a first layer (15a-15c) from the first sequence of layers (17a) and a second layer (15e-15g) from the second sequence of layers (17b).

13. A parameter group of parameters characterizing the behavior of a generator (1), obtained with the method (100) of any one of claims 1 to 12.

14. A method (200) for operating a generator (1), obtained with the method (100) of any one of claims 1 to 12, comprising the steps of: • transforming (210) at least one actual signal (11) of real or simulated physical measurement data of the first physical observation modality Mod_A with the generator (1) into at least one transformed signal (21); • determining (220) quantitative contributions (22a-22c) for the transformed signal (21), wherein different components (12a-12c) of real or simulated physical measurement data (10a) of the first physical observation modality Mod_A make the quantitative contributions to the transformed signal (21); • determining (230) actual labels (12a*-12c*) for the different components (12a-12c) of real or simulated physical measurement data (10a) of the first physical observation modality Mod_A, respectively; • determining (240) at least one label (21*) for the transformed signal (21) from the quantitative contributions (22a-22c) in combination with the actual labels (12a*-12c*).

15. The method (200) of claim 14, wherein • a generator (1) is selected (205) that has and / or is a KNN (14) with a number of neurons and / or other processing units (16) that sum their respective inputs in a weighted manner according to trainable parameters (1a) of the generator (1) and that constitute their output by applying a non-linear activation function to the result of this weighted sum; and • starting from the transformed signal (21) it is determined (221) in the context of the architecture of the KNN (14) to which extent components (12a-12c) of real or simulated physical measurement data (10a) of the first physical observation modality Mod_A contribute to rejecting at least one activation function.

16. The method (200) of any one of claims 14 to 15, wherein • determining (241) a label for the transformed signal (21) representing the class from a majority vote under the actual labels (12a*-12c*) representing the class, and / or • determining (242) a label for the transformed signal (21) representing the regression value from a combined function of the actual labels (12a*-12c*) representing the regression value.

17. A method (300) for operating a generator (1) obtained with the method (100) according to any one of claims 1 to 12, the method having the steps of: • transforming (310) with the generator (1) at least one actual signal (11) having real or simulated physical measurement data (10a) of the first physical observation modality Mod_A into at least one transformed signal (21) for which at least one actual label (11a) is available; • determining (315) at least one label (21*) for the transformed signal (21); • comparing (320) the transformed signal (21) with at least one other signal (21b) comprising real or simulated physical measurement data (20a) of the second physical observation modality Mod_B; • determining (330) from the label (21*) for the transformed signal (21) in combination with the result (320a) of the comparison (320) at least one label (21b*) for the other signal (21b) and / or a spatial offset (Delta) between the two physical observation modalities Mod_A and Mod_B.

18. A method (400) having the steps of: • training (410) a generator (1) with the method (100) according to any one of claims 1 to 12; • generating (420) with the trained generator (1*) at least one synthetic signal (21) of a second observation modality Mod_B from an actual signal (11) having real or simulated measurement data (10a) of a first observation modality Mod_A; • training (430) a machine learning module (3) with the synthetic signal (21); • recording (440) physical measurement data (20a) of the second observation modality Mod_B from an environment (50a) of a vehicle (50) with at least one sensor (4); • operating (450) the trained machine learning module (3*) in such a way that it takes (451) the physical measurement data (20a) provided by the sensor (4) as input and maps (452) it onto at least one class (450a) and / or at least one regression value (450b); • determining (460) a maneuvering signal (460a) from the class (450a) and / or from the regression value (450b); • maneuvering (470) the vehicle (50) with the maneuvering signal (460a).

19. A computer program product comprising machine readable instructions which, when implemented on one or more computers, cause the one or more computers to implement the method (100, 200, 300, 400) according to any one of claims 1 to 12 or 14 to 18.

20. A download product and / or a machine readable data carrier having the parameter set according to claim 13 and / or having a computer program contained in the computer program product according to claim 19.

21. A computer equipped with the parameter set according to claim 13, with a computer program contained in the computer program product according to claim 19 and / or with the download product and / or the machine readable data carrier according to claim 20.

Citation Information

Patent Citations

  • A face image age recognition method based on an improved ensemble learning strategy

    CN109726703A

  • "Matching Adversarial Networks"

    US20190147320A1