Method and apparatus for testing the robustness of artificial neural networks

Through the neural network system of the new generative model and the reverse training of the generator, encoder and discriminator, the deficiencies of image data enhancement and anomaly detection are solved, the robustness and effectiveness are improved, and the statistical distribution consistency of the training data set and the reliability of monitoring are ensured.

CN114008633BActive Publication Date: 2025-10-03ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080046927.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-28
Filing Date
2020-06-10
Publication Date
2025-10-03
Estimated Expiration
2040-06-10

AI Technical Summary

Technical Problem

Existing generative methods have shortcomings in image data enhancement and anomaly detection, making it difficult to achieve both robustness and effectiveness.

Method used

A new generative model is adopted, including a neural network system of generator, encoder and discriminator. The robustness of the system is improved through a reverse training mechanism. The generator is used to generate images, the encoder extracts potential features, and the discriminator identifies the authenticity of the image, thereby achieving enhancement of the training data set and anomaly detection.

Benefits of technology

It achieves simple enhancement of image data and robust anomaly detection, ensures the statistical distribution consistency of the training dataset, and improves the training convergence and monitoring reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114008633B_ABST
    Figure CN114008633B_ABST
Patent Text Reader

Abstract

A computer-implemented neural network system comprising: a first machine learning system, in particular a first neural network (GEN); a second machine learning system, in particular a second neural network (ENC); and a third machine learning system, in particular a third neural network (DIS), wherein the first machine learning system (GEN) is configured to determine a higher-dimensional constructed image (formula (A)) from predeterminable low-dimensional latent variables (z), wherein the second machine learning system (ENC) is configured to determine the latent variables (z) from the higher-dimensional constructed image (formula (A)), and wherein the third machine learning system (DIS) is configured to identify whether an image supplied to the third machine learning system (DIS) is a real image (x).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for testing the robustness of an artificial neural network, a method for training an artificial neural network, a method for operating an artificial neural network, a training device, a computer program, and a machine-readable storage medium. Background Art

[0002] “CVAE-GAN: Fine-Grained Image Generation through Asymmetric Training” (arXiv preprint arXiv: 1703.10155, 2017, Jianmin Bao, Dong Chen, Fang Wen, Houqiang Li, and Gang Hua) presents an overview of known generative methods like Variational Autoencoders and Generative Adversarial Networks. Summary of the Invention

[0003] Advantages of the present invention

[0004] In contrast, the invention having the features of independent claim 1 has the advantage that a novel generative model is provided which is suitable in an equally advantageous manner for enhancing image data and for anomaly detection.

[0005] Further aspects of the invention are the subject matter of the independent claims. Advantageous developments are the subject matter of the dependent claims.

[0006] Disclosure of the Invention

[0007] In a first aspect, the present invention relates to a computer-implemented neural network system comprising: a first machine learning system, in particular a first neural network, also referred to as a generator; a second machine learning system, in particular a second neural network, also referred to as an encoder; and a third machine learning system, in particular a third neural network, also referred to as a discriminator. The first machine learning system is configured to determine a higher-dimensional constructed image from predeterminable low-dimensional latent variables; the second machine learning system is configured to in turn determine latent variables from the higher-dimensional constructed image; and the third machine learning system is configured to discern whether an image supplied to the third machine learning system is a real image, i.e., an image recorded using a sensor. This inverse autoencoder of the network system has the advantage that correlations of latent features (such as, for example, the hair color of a detected pedestrian) can be particularly easily extracted, making augmentation of the training dataset particularly simple. At the same time, anomaly detection can be performed particularly robustly because the system can be trained in an adversarial manner.

[0008] In a parallel aspect, the present invention relates to a method for training a neural network system, wherein a first machine learning system (and in particular only the first machine learning system) is trained such that activations in a predeterminable feature map of the third machine learning system assume predominantly identical values ​​if a real image or an image of the real image reconstructed by a series connection of the second machine learning system and the first machine learning system is fed to a third machine learning system. This has been shown to lead to particularly good training convergence.

[0009] In a further development of this aspect, the first machine learning system can also be trained so that the third machine learning system is unlikely to recognize that images generated by the first machine learning system and fed to it are not real images. This results in particularly robust anomaly detection.

[0010] Alternatively or additionally, it can be provided that the second machine learning system (and in particular only the second machine learning system) is trained in such a way that the reconstruction of the latent variables determined by the cascade of the first and second machine learning systems is as identical as possible to these latent variables. It has been found that the convergence of the method is significantly improved if this reconstruction is selected so that only the parameters of the second machine learning system are trained, since otherwise the cost functions of the encoder and the generator are difficult to reconcile with each other.

[0011] In order to achieve the best possible improvement in the training results, an extended solution can be provided in which the third machine learning system is trained as follows: the third machine learning system recognizes as far as possible that the image generated by the first machine learning system and supplied to the third machine learning system is not a real image; and / or the third machine learning system is also trained as follows: the third machine learning system recognizes as far as possible that the real image supplied to the third machine learning system is a real image.

[0012] In other parallel aspects, the present invention relates to a method for monitoring the correct operation of a machine learning system, in particular a fourth neural network, wherein the machine learning system is used for classification and / or semantic segmentation of an input image (x) supplied to it, for example for identifying pedestrians and / or other traffic participants and / or street signs, wherein the monitoring unit comprises a first machine learning system and a second machine learning system of the neural network system, which are trained using one of the above-mentioned methods according to one of the claims, wherein the input image is supplied to the second machine learning system, the second machine learning system determines low-dimensional latent variables therefrom, and the first machine learning system determines a reconstruction of the input image from the latent variables, wherein it is determined based on the input image and the reconstructed input image whether the machine learning system is robust.

[0013] Such monitoring is particularly reliable if the machine learning system and the neural network system are trained using datasets comprising the same input images, since it is particularly simple to ensure that the statistical distributions of the training datasets are comparable (ie identical).

[0014] In another parallel aspect, the present invention relates to a method for generating an augmented training data set comprising input images, the augmented training data set being used for training a machine learning system which is set up for classification and / or semantic segmentation of the input images, wherein latent variables are determined from the input images respectively with the aid of a second machine learning system of a neural network system, in particular which has been trained using one of the above-mentioned methods, wherein the input images are classified according to determined feature representations (Merkmalsauspraegungen) of their image data, and wherein an augmented input image of the augmented training data set is determined from at least one of the input images according to an average value of the determined latent variables in at least two of the categories.

[0015] Using this method, it is possible to analyze the characteristics of the image in a latent variable space and to extract unbundled features, making it possible to make particularly targeted changes to the characteristics of the image in the described behavior.

[0016] If the image category is selected so that the input image is classified into it ( ) are consistent with respect to their representation in a predeterminable set of other features, then the unbundled features among these features are particularly clean.

[0017] In this case, it can be advantageously provided that the enhanced input image is determined based on the determined enhanced latent variables using a first machine learning system of a neural network system, in particular one that has been trained using the aforementioned training method. This allows for efficient generation of a modified image.

[0018] In order to modify predeterminable features of an existing image in a highly targeted manner, it can be provided that an enhanced latent variable is determined from predeterminable ones of the determined latent variables and from the difference in mean values. This changes the image feature corresponding to the predeterminable ones of the determined latent variables.

[0019] In order to obtain as many new feature representations as possible, it can be provided that the differences are weighted using predeterminable weighting factors In particular, this makes it possible to generate a large number of training images whose features vary strongly to varying degrees.

[0020] For example, it is possible to vary the visual properties of pedestrians in a large number of representations for street scenes and thus to provide particularly large training or test data sets that ensure a very high coverage with respect to this feature.

[0021] In particular, it can be provided that if the monitoring device has determined using one of the above-mentioned monitoring methods that the machine learning system is not robust, the generated augmented training data set is used to train the machine learning system.

[0022] In other aspects, the present invention relates to a computer program which is set up to carry out the above-mentioned method and to a machine-readable storage medium on which the computer program is stored. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Subsequently, the embodiments of the present invention will be described in more detail with reference to the accompanying drawings. In the drawings:

[0024] Figure 1 Schematically shows the structure of an embodiment of the present invention;

[0025] Figure 2 Schematically illustrates an embodiment for controlling an at least partially autonomous robot;

[0026] Figure 3 An embodiment for controlling a production system is schematically shown;

[0027] Figure 4 An embodiment for controlling an access system (Zugangssystem) is schematically shown;

[0028] Figure 5 Schematically illustrates an embodiment for controlling a monitoring system;

[0029] Figure 6 An embodiment for controlling a personal assistant is schematically shown;

[0030] Figure 7 Schematically illustrates an embodiment for controlling a medical imaging system;

[0031] Figure 8 A possible structure of a monitoring unit is shown;

[0032] Figure 9 A possible structure of a first training device 141 is shown;

[0033] Figure 10 A neural network system is shown;

[0034] Figure 11 A possible design of a second training device 140 is shown. DETAILED DESCRIPTION

[0035] Figure 1 An actuator 10 is shown, interacting with a control system 40 in its environment 20. At preferably regular time intervals, the environment 20 is detected by a sensor 30, particularly an imaging sensor such as a video sensor. This sensor 30 can also be provided by multiple sensors, for example, a stereo camera. Other imaging sensors are also conceivable, such as radar, ultrasound, or lidar. Thermal imaging cameras are also conceivable. A sensor signal S from the sensor 30 (or each sensor signal S in the case of multiple sensors) is transmitted to the control system 40. The control system 40 thus receives a series of sensor signals S. From these, the control system 40 determines a control signal A and transmits this control signal A to the actuator 10.

[0036] The control system 40 receives a series of sensor signals S from the sensor 30 in an optional receiving unit 50. The receiving unit 50 converts the series of sensor signals S into a series of input images x (alternatively, each sensor signal S can be directly used as an input image x). The input images x can be, for example, segments of the sensor signals S or further processed. The input images x include individual frames of a video recording. In other words, the input images x are determined based on the sensor signals S. The series of input images x are fed to a machine learning system, in this embodiment, to an artificial neural network 60.

[0037] The artificial neural network 60 is preferably configured by parameters To parameterize, the parameters are stored in the parameter memory P and provided by said parameter memory P.

[0038] Artificial neural network 60 determines output variables y from input image x. These output variables y can include, in particular, a classification and / or semantic segmentation of input image x. Output variables y are supplied to an optional deformation unit 80, which determines actuation signals A therefrom, which are supplied to actuator 10 for corresponding actuation of actuator 10. Output variables y include information about objects detected by sensor 30.

[0039] The control system 40 further includes a monitoring unit 61 for monitoring the operating mode of the artificial neural network 60. The input image x is also supplied to the monitoring unit 61. Based on this, the monitoring unit 61 determines a monitoring signal d, which is also supplied to the deformation unit 80. The control signal A is also determined based on the monitoring signal d.

[0040] Monitoring signal d indicates whether neural network 60 reliably determines output variable y. If monitoring signal d indicates unreliability, it can be provided, for example, that control signal A is determined according to a guaranteed operating mode (whereas control signal A is otherwise determined in normal operating mode). A guaranteed operating mode can, for example, include reducing the dynamics of actuator 10 or shutting down a function for controlling actuator 10.

[0041] Actuator 10 receives control signal A, is controlled accordingly, and performs a corresponding action. In this case, actuator 10 may include a control logic device (not necessarily structurally integrated) that determines a second control signal from control signal A and uses this second control signal to control actuator 10.

[0042] In other embodiments, the control system 40 includes the sensor 30 . In still other embodiments, the control system 40 also includes the actuator 10 as an alternative or in addition.

[0043] In other preferred embodiments, the control system 40 includes one or more processors 45 and at least one machine-readable storage medium 46, on which instructions are stored. When the instructions are executed on the processor 45, the instructions prompt the control system 40 to perform the method according to the present invention.

[0044] In an alternative embodiment, a display unit 10 a is provided as an alternative to the actuator 10 or in addition to the actuator 10 .

[0045] Figure 2 It is shown how the control system 40 can be used to control an at least partially autonomous robot, here an at least partially autonomous motor vehicle 100 .

[0046] Sensor 30 may be, for example, a video sensor which is preferably situated in motor vehicle 100 .

[0047] The artificial neural network 60 is designed to unambiguously identify objects from an input image x.

[0048] Actuators 10, preferably situated in motor vehicle 100, may be, for example, brakes, drives, or steering systems of motor vehicle 100. Actuation signal A may then be determined such that one or more actuators 10 are actuated so that motor vehicle 100 prevents, for example, a collision with an object unambiguously identified by artificial neural network 60, particularly when the object is a specific class, such as a pedestrian.

[0049] Alternatively, the at least partially autonomous robot may be another mobile robot (not shown), for example, a robot that moves by flying, floating, diving, or walking. The mobile robot may also be, for example, an at least partially autonomous lawn mower or an at least partially autonomous cleaning robot. In these cases, the control signal may also be determined so that the drive and / or steering of the mobile robot is controlled so that the at least partially autonomous robot, for example, prevents a collision with an object identified by artificial neural network 60.

[0050] Alternatively or additionally, the display unit 10a can be activated using the control signal A and, for example, display the determined safety zone. For example, in a motor vehicle 100 with non-automated steering, it is also possible to activate the display unit 10a using the control signal A so that the display unit 10a outputs a visual or acoustic warning signal when an imminent collision of the motor vehicle 100 with one of the clearly identified objects is determined.

[0051] Figure 3The following exemplary embodiment is shown, in which a control system 40 is used to control a production machine 11 of a production system 200 by controlling an actuator 10 that controls the production machine 11. The production machine 11 may be, for example, a machine for punching, sawing, drilling, and / or cutting. The sensor 30 may be, for example, an optical sensor that detects properties of the products 12a and 12b. The products 12a and 12b may be movable. It is possible to control the actuator 10 that controls the production machine 11 based on the detected assignment of the products 12a and 12b so that the production machine 11 performs the subsequent processing steps for the correct one of the products 12a and 12b. It is also possible to adapt the production machine 11 to the same production steps for processing the subsequent product by identifying the correct properties of the same one of the products 12a and 12b (i.e., if there is no incorrect assignment).

[0052] Figure 4 The following exemplary embodiment is shown: in this embodiment, a control system 40 is used to control an access system 300. Access system 300 may include a physical access control, such as a door 401. A video sensor 30 is configured to detect persons. The detected images can be interpreted using an object identification system 60. If multiple persons are detected simultaneously, the identities of the persons (i.e., objects) can be determined particularly reliably by assigning them to one another, for example, by analyzing their movements. The actuator 10 may be a lock that, in response to an actuation signal A, releases or does not release the access control, for example, by opening or not opening door 401. To this end, actuation signal A can be selected based on the interpretation of the object identification system 60, for example, based on the determined identity of the person. Instead of physical access control, a logical access control may also be provided.

[0053] Figure 5 The following embodiment is shown: In this embodiment, the control system 40 is used to control the monitoring system 400. Figure 5 The embodiment shown in FIG differs in that a display unit 10 a is provided instead of the actuator 10 and is controlled by the control system 40. For example, the artificial neural network 60 can reliably determine the identity of the objects recorded by the video sensor 30 in order to infer from this which objects are suspicious, for example, and then select the control signal A so that the display unit 10 a highlights these objects in color.

[0054] Figure 6An embodiment is shown in which the control system 40 is used to control a personal assistant 250. The sensor 30 is preferably an optical sensor which receives images of the user's 249 gestures.

[0055] Based on the signals from sensor 30, control system 40 determines an actuation signal A for personal assistant 250, for example, by performing gesture recognition using a neural network. The determined actuation signal A is then transmitted to personal assistant 250, and personal assistant 250 is controlled accordingly. Determined actuation signal A can, in particular, be selected so that it corresponds to a presumed desired actuation by user 249. The presumed desired actuation can be determined based on the gesture recognized by artificial neural network 60. Control system 40 can then select actuation signal A to be transmitted to personal assistant 250 based on the presumed desired actuation and / or can select actuation signal A to be transmitted to personal assistant 250 in accordance with the presumed desired actuation 250.

[0056] The corresponding control may include, for example, personal assistant 250 calling up information from a database and presenting it in a form acceptable to user 249 .

[0057] Instead of personal assistant 250 , a domestic appliance (not shown), in particular a washing machine, a stove, an oven, a microwave oven or a dishwasher, can also be provided to be controlled accordingly.

[0058] Figure 7 The following exemplary embodiment is shown, in which a control system 40 is used to control a medical imaging system 500, such as an MRT system, an X-ray machine, or an ultrasound system. Sensor 30 may be, for example, an imaging sensor, and display unit 10a is controlled by control system 40. For example, neural network 60 may determine whether an area recorded by the imaging sensor is of interest, and control signal A is then selected such that this area is highlighted in color on display unit 10a.

[0059] Figure 8 A possible structure of a monitoring unit 61 is shown. An input image x is fed to an encoder ENC, which determines a so-called latent variable z from it. The latent variable z has fewer dimensions than the input image x. The latent variable z is fed to a generator GEN, which generates a reconstructed image from it. In this embodiment, the encoder ENC and the generator GEN are respectively given by convolutional neural networks (English: "convolutional neural network"). Input image x and reconstructed image is fed to the discriminator DIS. The discriminator DIS has been trained to produce as good a quantity as possible: this quantity characterizes whether the image fed to the discriminator DIS is a real image or whether it has been generated by the generator GEN. This is discussed below with Figure 10 The generator GEN is also a convolutional neural network.

[0060] No. Layer (where is a pre-given number) to derive when the input image x or the reconstructed image is fed to the generator GEN The feature maps (English: "feature maps") use or These feature maps are fed to an evaluator BE where the reconstruction error is e.g. In an alternative embodiment (not shown), it is also possible to select the reconstruction error as .

[0061] Next, the outlier A(x) may be determined as the proportion (Anteil) of input images of a reference dataset (e.g., a training dataset with which the discriminator DIS and / or the generator GEN and / or the encoder ENC have been trained) whose reconstruction error is less than the determined reconstruction error If the abnormal value A(x) is greater than a predefinable threshold value, the monitoring signal d is set to the value d=1, which signals that the output variable y is potentially unreliable. Otherwise, the monitoring signal d is set to the value d=0, which signals that the determination of the output variable y is classified as reliable.

[0062] Figure 9 A possible structure of a first training device 141 for training the monitoring unit 51 is shown. This is parameterized with parameters θ, which are provided by a parameter memory P. The parameters θ comprise the generator parameters θ which parameterize the generator GEN. , Encoder parameters for parameterizing the encoder ENC and the discriminator parameters that parameterize the discriminator DIS .

[0063] The training device 141 comprises a provider 71 which provides an input image e from a training data set. The input image e is fed to the monitoring unit 61 to be trained, which determines the output quantity a from it. The output quantity a and the input image e are fed to an evaluator 74 which determines the output quantity a from it as described in the example above. Figure 10As described in connection with this, a new parameter θ′ is determined, which is transferred to the parameter memory P and replaces the parameter θ there.

[0064] The method performed by the training device 141 may be stored on a machine-readable storage medium 146 in the form of a computer program, and may be executed by the processor 145 .

[0065] Figure 10 The diagram illustrates the interaction of the generator GEN, encoder ENC and discriminator DIS during training. The arrangement of the generator GEN, encoder ENC and discriminator DIS shown here is also referred to as a neural network system in this literature.

[0066] First, the discriminator DIS is trained. The subsequent steps for training the discriminator DIS can be repeated, for example, DIS times, where n DIS is a predeterminable integer.

[0067] First, a batch of real input images x is provided. The real input images x are said to have a (usually unknown) probability distribution of These input images These are real images provided, for example, from a database. The entire set of input images is also called a training dataset.

[0068] In addition, a set of latent variables z is used as are provided, they have been randomly selected from the probability distribution The probability distribution is obtained. In this case, this is, for example, the (multidimensional) standard normal distribution.

[0069] In addition, a batch of random variables as are provided, they have been randomly selected from the probability distribution The probability distribution In this case, for example, it is a uniform distribution over the interval [0; 1].

[0070] The latent variable z is fed to the generator GEN and given the constructed input image ,that is

[0071] .

[0072] The input image x and the constructed input image Between, using random variables To interpolate, that is

[0073] .

[0074] With a predeterminable gradient factor λ, which can be chosen to be λ=10, for example, the discriminator cost function

[0075]

[0076] The new discriminator parameters are determined. It can be determined from

[0077] ,

[0078] Here, “Adam” stands for gradient descent. Thus, the training of the discriminator DIS is completed.

[0079] Next, we train the generator GEN and encoder ENC. Here, the real input image is also provided , and provide randomly selected latent variables . Re-determine:

[0080] .

[0081] From this, determine the latent variables for reconstruction , which is to construct the image: is sent to the encoder ENC, that is

[0082] .

[0083] Likewise, as in Figure 8 As illustrated in , with the help of encoder ENC and generator GEN, we try to reconstruct the input image x, that is,

[0084] .

[0085] Now, the generator cost function , the reconstruction cost function of the input image x and the reconstruction cost function of the latent variable z Determined to be

[0086] .

[0087] Then, the new generator parameters and the new encoder parameters Determined to be

[0088] .

[0089] New generator parameters , new encoder parameters and the new discriminator parameters Then replace the generator parameters , encoder parameters and the discriminator parameters .

[0090] At this point, the convergence of the parameters θ can be checked and, if necessary, the training of the discriminator DIS and / or the generator GEN and encoder ENC can be repeated until convergence is achieved. The method then ends.

[0091] Figure 11 An exemplary second training device 140 for training a neural network 60 is shown. The training device 140 includes a provider 72 that provides an input image x and a desired output quantity ys (e.g., a desired classification). The input image x is fed to the artificial neural network 60 to be trained, which determines the output quantity y therefrom. The output quantity y and the desired output quantity ys are fed to a comparator 75, which determines a new parameter based on the correspondence between the respective output quantity y and the desired output quantity ys. , the new parameter are transferred to the parameter memory P and replaced there with the parameter .

[0092] The method performed by the training system 140 may be stored on a machine-readable storage medium 148 in the form of a computer program and may be executed by the processor 147 .

[0093] A data set comprising an input image x and the associated desired output quantity ys may be enhanced or generated (e.g. by a provider 72) as follows. First, a data set comprising an input image x is provided. The input image is classified according to predefined representations of features (e.g., "A" and "B"). For example, vehicles can be classified according to the features "headlights on" or "headlights off," or identified cars can be classified according to the types "sedan" or "station wagon." For example, in the case of pedestrian recognition, different representations of the feature "hair color" are also possible. Depending on the representation of this feature, the input image is divided into two sets, namely I A ={i| With representations "A"} and I B ={i| With representation "B"}. Advantageously, these sets are also homogenized as follows: For a predeterminable set of further features, preferably all further features, the input image have the same form X, that is,

[0094] {i| Has the form X}

[0095] {i| Has the form X}.

[0096] With the help of encoder ENC, for the input image Each input image determines the latent variable .

[0097] Then, determine the mean value of the latent variable on the set, that is,

[0098] .

[0099] Next, the difference of the mean values ​​is formed, i.e.

[0100] .

[0101] Now, for the A The image is formed with a predeterminable scaling factor The new latent variable of , the scaling factor can take a value between 0 and 1, that is,

[0102] .

[0103] Correspondingly, for the B The new latent variable can be formed as

[0104] .

[0105] From this, a new image can be generated using the following formula

[0106] .

[0107] Of course, the entire image does not have to be classified. It is possible to use a detection algorithm to classify image segments as objects, for example, which are then cut out and new image segments are generated (corresponding to the new image segments). ), and the removed image segment is inserted into the associated image instead. In this way, for example, it is possible to selectively adapt the hair color of a detected pedestrian in an image.

[0108] Apart from the classification of the features that thus change between the representation forms "A" and "B", the associated desired output variable ys can be adopted unchanged. In this way, an enhanced data set can be generated and the neural network 60 can be trained thereby. The method thus ends.

[0109] The term "computer" includes any device for processing predefinable calculation rules. These calculation rules can exist in the form of software, hardware, or a hybrid form consisting of software and hardware.

Claims

1. A computer-implemented neural network system comprising: a first machine learning system comprising a first neural network (GEN), a second machine learning system comprising a second neural network (ENC), and a third machine learning system comprising a third neural network (DIS), wherein the first machine learning system is configured to determine a higher-dimensional structure image from a predeterminable low-dimensional latent variable (z) The second machine learning system is configured to construct an image from the higher dimensional The reconstruction of the low-dimensional latent variable (z) is determined in And wherein the third machine learning system is configured to receive the higher dimensional structure image And determine whether the image fed to the third machine learning system is a real image (x).

2. A method for training a neural network system comprising: a first machine learning system comprising a first neural network (GEN), a second machine learning system comprising a second neural network (ENC), and a third machine learning system comprising a third neural network (DIS), wherein the first machine learning system is configured to determine a higher-dimensional structure image from a predeterminable low-dimensional latent variable (z) The second machine learning system is configured to construct an image from the higher dimensional The reconstruction of the low-dimensional latent variable (z) is determined in And wherein the third machine learning system is configured to receive the higher dimensional structure image and discern whether the image fed to the third machine learning system is a real image (x), In the method: The first machine learning system is trained as follows: If a real image (x) or an image of the real image (x) reconstructed by the concatenation of the second machine learning system and the first machine learning system is fed to the third machine learning system Then the activations (DIS l ) take the same value as much as possible, and The second machine learning system is trained as follows: The reconstruction of the low-dimensional latent variable (z) determined by the concatenation of the first machine learning system and the second machine learning system Be as identical as possible to the low-dimensional latent variable (z).

3. The method according to claim 2, wherein: The first machine learning system is also trained such that the third machine learning system is unlikely to recognize images generated by the first machine learning system and fed to the third machine learning system. Not a real image.

4. The method according to claim 2, wherein: The third machine learning system is trained as follows: the third machine learning system is able to recognize, as far as possible, images generated by the first machine learning system and fed to the third machine learning system. Not a real image.

5. The method according to claim 4, wherein The third machine learning system is also trained as follows: the third machine learning system recognizes as far as possible that a real image (x) supplied to the third machine learning system is a real image.

6. A training device (141) comprising a memory and a processor, wherein a computer program is stored on the memory and is configured to carry out the method according to any one of claims 2 to 5 when the computer program is run on the processor.

7. A method for monitoring the correct operation of a machine learning system (60) by means of a monitoring unit (61), the machine learning system (60) being configured to classify and / or semantically segment an input image (x) supplied to the machine learning system (60), wherein the monitoring unit (61) comprises a neural network system according to claim 1, wherein the input image (x) is supplied to the second machine learning system, the second machine learning system determining a low-dimensional latent variable (z) therefrom, and the first machine learning system determining a reconstruction of the input image from the low-dimensional latent variable (z). According to the input image (x) and the reconstructed input image To determine whether the machine learning system (60) is robust.

8. The method according to claim 7, wherein: The robustness of the machine learning system (60) is determined based on the following aspects: when the input image (x) is supplied to the third machine learning system, the activation (DIS l ) which value (DIS l (x)); and when the reconstructed input image is fed to the third machine learning system When the activation (DIS l ) which value to take 9. The method according to claim 7 or 8, wherein: The machine learning system (60) and the neural network system have been trained using a data set comprising the same input image (x).

10. The method according to any one of claims 7 to 8, wherein A control signal (A) for controlling an actuator (10) is selected based on whether the machine learning system (60) has been determined to be robust, and the control signal (A) is provided based on an output signal (y) of the machine learning system (60).

11. A monitoring unit (61) comprising a memory and a processor, wherein a computer program is stored on the memory and is configured to carry out the method according to claim 7 when the computer program is run on the processor.

12. A method for generating a (i) ), the enhanced training data set being used for training a machine learning system (60) configured for classification and / or semantic segmentation of an input image (x), wherein a second machine learning system trained by the method according to claim 2 using a neural network system according to claim 1 is used to obtain a classification image from the input image (x). (i) ) to determine the low-dimensional latent variables (z (i) ), where the input image (x (i) ) are classified according to the determined characteristic representations of their image data, and wherein from the input image (x (i) ), according to the low-dimensional latent variables (z (i) ) Determine the augmented input image of the augmented training dataset 13. The method according to claim 12, wherein: By means of the first machine learning system of the neural network system according to claim 1, the enhanced low-dimensional latent variables are determined To determine the enhanced input image The first machine learning system has been trained using a method according to any one of claims 2 to 5.

14. The method according to claim 13, wherein: From the determined low-dimensional latent variables (z (i) ) and from the mean value The difference (v A-B ), determine the enhanced potential variables This changes the image and the determined low-dimensional latent variable (z (i) ) corresponds to the features of the pre-given latent variables.

15. The method according to claim 14, wherein The difference (v A-B ) for weighting.

16. The method according to any one of claims 12 to 15, wherein If monitoring using the method according to any one of claims 7 to 10 has shown that the machine learning system (60) is not robust, the machine learning system (60) is trained using the generated augmented training data set.

17. A training device (140) comprising a memory and a processor, wherein a computer program is stored on the memory and is configured to carry out the method according to claim 12 when the computer program is run on the processor. 18 . A computer program product comprising a computer program which is configured to carry out the method according to claim 1 , when the computer program is run on a processor.

19. A machine-readable storage medium (46, 146, 148) having a computer program stored thereon, which is configured to carry out the method according to any one of claims 1 to 5 or 7 to 10 or 12 to 16 when the computer program is run on a processor.

Citation Information

Patent Citations

  • Convolutional neural network and user habitual behavior analysis combination-based AR system gesture identification method

    CN108334814A

  • Adversarial and Dual Inverse Deep Learning Networks for Medical Image Analysis

    US20180225823A1