Method and apparatus for training a machine learning system
By combining the CVAE-GAN method with the monitoring unit, the robustness problem of the machine learning system in terms of feature changes is solved, training data with high coverage is generated, and the training effect of the system is improved.
Patent Information
- Application Number
- CN202080046427.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2020-06-10
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2040-06-10
AI Technical Summary
Existing machine learning systems are not robust enough in terms of feature changes, making it difficult to effectively generate training data with diverse features, resulting in poor training results.
The CVAE-GAN method is used to modify image features in the latent variable space through autoencoders and generative adversarial networks to generate enhanced data records, and the monitoring unit and neural network system are used to ensure the robustness of the system, including the combined training of the first, second and third machine learning systems.
It realizes the robustness monitoring and training of machine learning systems, generates training data with high coverage, and improves the robustness of the system in terms of feature changes and the training effect.
Smart Images

Figure CN113994349B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, a training device, a computer program and a machine-readable storage medium for training a machine learning system. Background Art
[0002] “CVAE-GAN: Fine-Grained Image Generation through Asymmetry Training” (arXiv preprint arXiv: 1703.10155, 2017, Jianmin Bao, Dong Chen, Fang Wen, Houqiang Li, and Gang Hua) provides an overview of known generative methods such as variational autoencoders and generative adversarial networks. Summary of the Invention
[0003] Advantages of the invention
[0004] The invention having the features of independent claim 1 has the advantage that a particularly good enhanced data record can be provided. This is possible because the characteristics of the image can be analyzed particularly well in a latent variable space (English: "latent space") and unbound features can be extracted, making it possible to modify the characteristics of the image in a particularly targeted manner in the described manner.
[0005] Further aspects of the invention are the subject matter of the independent claims. Advantageous developments are the subject matter of the dependent claims.
[0006] Summary of the Invention
[0007] In a first aspect, the present invention relates to a computer-implemented method for generating an enhanced data record comprising an input image for training a machine learning system, wherein the machine learning system is configured for classification and / or semantic segmentation of the input image, the machine learning system comprising a first machine learning system, in particular a first neural network, and a second machine learning system, in particular a second neural network, wherein the first machine learning system is configured as a decoder of an autoencoder and the second machine learning system is configured as an encoder of an autoencoder, wherein latent variables are determined from the input images with the aid of the encoders, wherein the input images are classified according to determined feature representations of their image data, and wherein enhanced input images of the enhanced data record are determined from at least one of the input images in at least two classes according to average values of the determined latent variables, wherein the image classes are selected such that the input images belonging to the image classes are consistent with respect to their representations in a predeterminable set of further features.
[0008] In this case, it can then be advantageously provided that the enhanced input image is determined by means of the decoder as a function of the determined enhancement latent variable. This allows for efficient generation of modified images.
[0009] In order to modify predeterminable features of an existing image in a highly targeted manner, it can be provided that the enhancement latent variable is determined from the difference between a predeterminable latent variable of the determined latent variables and the mean value. The features of the image corresponding to the predeterminable latent variable of the determined latent variables are thereby changed.
[0010] In order to obtain as many new feature representations as possible, it can be provided that the weighting factors The differences are weighted. In particular, this makes it possible to generate a large number of training images whose features vary to varying degrees.
[0011] For example, it is possible to vary the visual properties of pedestrians in a large number of representations for a street scene and thus to provide a particularly large training or test data set that ensures a very high coverage with respect to this feature.
[0012] In one refinement, provision can be made to use the generated enhanced data record to test, in particular, whether the trained machine learning system is robust, and in relation thereto, to continue training if, in particular, only if, the test indicates that the machine learning system is not robust. This makes it possible to particularly reliably test whether the machine learning system is robust with respect to changing characteristics.
[0013] Alternatively or additionally, provision can be made for the machine learning system to be trained with the generated augmented data records when, in particular only when, monitoring has shown that the machine learning system is not robust.
[0014] In an improved solution of this aspect, it can be provided that the machine learning system is monitored by means of a monitoring unit, the monitoring unit comprising a first machine learning system and a second machine learning system, wherein the input image is fed to the second machine learning system, the second machine learning system determines low-dimensional latent variables therefrom, and the first machine learning system determines a reconstruction of the input image from the low-dimensional latent variables, wherein whether the machine learning system is robust is determined based on the input image and the reconstructed input image.
[0015] In an improvement of this aspect, it can be provided that the monitoring unit also includes a third machine learning system of the neural network system, wherein the neural network system includes the first machine learning system, the second machine learning system and a third machine learning system, in particular a third neural network, wherein the first machine learning system is constructed to determine the constructed higher-dimensional image from predeterminable low-dimensional latent variables, wherein the second machine learning system is constructed to determine the latent variables from the constructed higher-dimensional image, and wherein the third machine learning system is constructed to distinguish whether the image supplied to the third machine learning system is a real image, wherein whether the machine learning system is robust is determined based on the following items: when the input image is supplied to the third machine learning system, which value an activation in the predeterminable feature map of the third machine learning system takes and when the reconstructed input image is supplied to the third machine learning system, which value the one activation in the predeterminable feature map of the third machine learning system takes.
[0016] In this case, it can be provided that the first machine learning system is trained such that when a real image or an image of a real image reconstructed in series by the second machine learning system and the first machine learning system is fed to the third machine learning system, an activation in a predefinable feature map in the feature map of the third machine learning system assumes the same value. It has been shown that this leads to particularly good training convergence.
[0017] A refinement of this aspect can provide that the first machine learning system is also trained so that the third machine learning system is unlikely to recognize that an image generated by the first machine learning system and fed to the third machine learning system is not a real image. This ensures particularly robust anomaly detection.
[0018] Alternatively or additionally, provision can be made for the second machine learning system, and in particular only the second machine learning system, to be trained in such a way that the reconstruction of the latent variable, determined in series by the first and second machine learning systems, is as close to the latent variable as possible. It has been found that the convergence of the method is significantly improved if this reconstruction is selected such that only the parameters of the second machine learning system are trained, since otherwise it is difficult to bring the cost functions of the encoder and the generator into harmony with one another.
[0019] In order to achieve the best possible improvement in the training results, in an improved scheme it can be provided that the third machine learning system is trained as follows, namely, the third machine learning system recognizes as far as possible that the image generated by the first machine learning system and supplied to the third machine learning system is not a real image and / or the third machine learning system is also trained as follows, namely, the third machine learning system recognizes as far as possible that the real image supplied to the third machine learning system is a real image.
[0020] In order to achieve the best possible improvement in the training results, in an improved scheme it can be provided that the third machine learning system is trained as follows, namely, the third machine learning system recognizes as far as possible that the image generated by the first machine learning system and supplied to the third machine learning system is not a real image and / or the third machine learning system is also trained as follows, namely, the third machine learning system recognizes as far as possible that the real image supplied to the third machine learning system is a real image.
[0021] Since it is particularly easy to ensure that the statistical distribution of the training data records is comparable (ie identical), monitoring is particularly reliable if the machine learning system and the neural network system are trained using data records comprising the same input images.
[0022] In other aspects, the present invention relates to a computer program configured to perform the above-described method and to a machine-readable storage medium having the computer program stored thereon. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. In the drawings:
[0024] Figure 1 Schematically illustrates the structure of an embodiment of the present invention;
[0025] Figure 2 Schematically illustrates an embodiment for controlling an at least partially autonomous robot;
[0026] Figure 3 Schematically illustrates an embodiment for controlling a production system;
[0027] Figure 4 Schematically illustrates an embodiment of a system for controlling access;
[0028] Figure 5 Schematically illustrates an embodiment for controlling a monitoring system;
[0029] Figure 6 Schematically illustrates an embodiment for controlling a personal assistant;
[0030] Figure 7 Schematically illustrates an embodiment for controlling a medical imaging system;
[0031] Figure 8 shows a possible structure of a monitoring unit;
[0032] Figure 9 A possible structure of the first training device 141 is shown;
[0033] Figure 10 A neural network system is shown;
[0034] Figure 11 A possible structure of the second training device 140 is shown. DETAILED DESCRIPTION
[0035] Figure 1 An actuator 10 is shown interacting with a control system 40 in its environment 20. At preferably regular time intervals, the environment 20 is detected by a sensor 30, particularly an imaging sensor, such as a video sensor. This sensor can also be provided by multiple sensors, such as a stereo camera. Other imaging sensors are also conceivable, such as radar, ultrasound, or lidar. Thermal imaging cameras are also conceivable. The sensor signal S of the sensor 30—or, in the case of multiple sensors, each sensor signal S—is transmitted to the control system 40. The control system 40 thus receives a sequence of sensor signals S. From these, the control system 40 determines a control signal A, which is transmitted to the actuator 10.
[0036] The control system 40 receives a sequence of sensor signals S from the sensor 30 in an optional receiving unit 50, which converts the sequence of sensor signals S into a sequence of input images x. (Alternatively, the sensor signals S can also be directly used as input images x.) For example, the input images x can be segments of the sensor signals S or further processed. The input images x include individual frames of a video recording. In other words, the input images x are determined based on the sensor signals S. The sequence of input images x is fed to a machine learning system, in this embodiment, to an artificial neural network 60.
[0037] The artificial neural network 60 is preferably configured by parameters is parameterized, the parameters are stored in the parameter memory P and provided by said parameter memory.
[0038] Artificial neural network 60 determines output variables y from input image x. These output variables y can, in particular, include a classification and / or semantic segmentation of input image x. Output variables y are supplied to an optional transformation unit 80, which determines actuation signals A therefrom, which are supplied to actuator 10 in order to actuate actuator 10 accordingly. Output variables y include information about the object detected by sensor 30.
[0039] The control system 40 further comprises a monitoring unit 61 for monitoring the operating mode of the artificial neural network 60. The input image x is likewise supplied to the monitoring unit 61. From this, the monitoring unit determines a monitoring signal d, which is likewise supplied to the transformation unit 80. The control signal A is also determined as a function of the monitoring signal d.
[0040] Monitoring signal d indicates whether neural network 60 reliably determines output variable y. If monitoring signal d indicates unreliability, it can be provided, for example, that control signal A is determined according to a safety operating mode (whereas otherwise it is determined in normal operating mode). The safety operating mode can, for example, include reducing the dynamics of actuator 10 or shutting down functionality for controlling actuator 10.
[0041] Actuator 10 receives control signal A, is controlled accordingly, and performs a corresponding action. In this case, actuator 10 may include a control logic circuit (not necessarily structurally integrated) that determines a second control signal from control signal A and then controls actuator 10 using the second control signal.
[0042] In another embodiment, the control system 40 includes the sensor 30. In yet another embodiment, the control system 40 may further include the actuator 10, alternatively or additionally.
[0043] In another preferred embodiment, the control system 40 includes a single or multiple processors 45 and at least one machine-readable storage medium 46 on which instructions are stored, which, when executed on the processor 45, cause the control system 40 to perform the method according to the present invention.
[0044] In an alternative embodiment, a display unit 10 a is provided instead of or in addition to the actuator 10 .
[0045] Figure 2 It is shown how the control system 40 can be used to control an at least partially autonomous robot, here an at least partially autonomous motor vehicle 100 .
[0046] Sensor 30 may be, for example, a video sensor which is preferably situated in motor vehicle 100 .
[0047] The artificial neural network 60 is designed to reliably identify objects from an input image x.
[0048] The actuator 10 preferably situated in the motor vehicle 100 may be, for example, a brake, a drive, or a steering system of the motor vehicle 100. The actuation signal A may then be determined such that the actuator or actuators 10 are actuated so that the motor vehicle 100 avoids a collision with an object reliably identified by the artificial neural network 60, in particular if the object is a specific class, such as a pedestrian.
[0049] Alternatively, the at least partially autonomous robot may be another mobile robot (not shown), for example, one that moves forward by flying, swimming, diving, or walking. The mobile robot may also be, for example, an at least partially autonomous lawn mower or an at least partially autonomous cleaning robot. Even in these cases, the control signal A may be determined so that the drive and / or steering device of the mobile robot are controlled so that the at least partially autonomous robot avoids a collision with an object identified by the artificial neural network 60, for example.
[0050] Alternatively or additionally, the display unit 10a can be actuated using the actuation signal A and, for example, display the determined safety zone. For example, in the case of a motor vehicle 100 with a non-automatic steering system, it is possible to actuate the display unit 10a using the actuation signal A so that it outputs a visual or acoustic warning signal when an imminent collision of the motor vehicle 10 with one of the reliably identified objects is determined.
[0051] Figure 3 An exemplary embodiment is shown in which a control system 40 is used to control a production machine 11 of a production system 200 by controlling an actuator 10 that controls the production machine 11. The production machine 11 can be, for example, a machine for punching, sawing, drilling and / or cutting.
[0052] The sensor 30 can be, for example, an optical sensor that detects the properties of the production products 12a and 12b. These production products 12a and 12b may be movable. The actuator 10 controlling the production machine 11 may be controlled based on the detected distribution of the production products 12a and 12b so that the production machine 11 performs the subsequent processing steps for the correct one of the production products 12a and 12b. It is also possible that, by identifying the correct properties of the same one of the production products 12a and 12b (i.e., if there is no incorrect distribution), the production machine 11 adapts the same production steps for processing the subsequent product.
[0053] Figure 4 An exemplary embodiment is shown in which a control system 40 is used to control an access system 300. Access system 300 may include physical access controls, such as a door 401. A video sensor 30 is configured to detect persons. The detected images may be interpreted using an object identification system 60. If multiple persons are detected simultaneously, their identities can be determined particularly reliably, for example, by assigning the persons (i.e., objects) to one another, for example, by analyzing their movements. Actuator 10 may be a lock that, depending on an actuation signal A, releases or does not release access controls, for example, by opening or not opening door 401. To this end, actuation signal A may be selected based on the interpretation of the object identification system 60, for example, based on the determined identity of the person. Instead of physical access control, a logical access control may also be provided.
[0054] Figure 5 An embodiment is shown in which a control system 40 is used to control a monitoring system 400. Figure 5 The embodiment shown in FIG differs in that a display unit 10 a is provided instead of the actuator 10, which is activated by the control system 40. For example, the identity of the objects recorded by the video sensor 30 can be reliably determined by the artificial neural network 60 in order to infer which ones are suspicious, for example, and the activation signal A can then be selected so that the display unit 10 a displays these objects in a highlighted color manner.
[0055] Figure 6 An embodiment is shown in which the control system 40 is used to control the personal assistant 250. The sensor 30 is preferably an optical sensor that receives an image of the user's 249 gesture.
[0056] Based on the signals from sensor 30, control system 40 determines an actuation signal A for personal assistant 250, for example, by a neural network performing gesture recognition. This determined actuation signal A is then transmitted to personal assistant 250, and the personal assistant is controlled accordingly. This determined actuation signal A can, in particular, be selected so that it corresponds to a presumed desired actuation by user 249. This presumed desired actuation can be determined based on a gesture recognized by artificial neural network 60. Control system 40 can then select actuation signal A for transmission to personal assistant 250 based on the presumed desired actuation and / or select actuation signal A for transmission to personal assistant 250 corresponding to the presumed desired actuation 250.
[0057] This corresponding control may include, for example, personal assistant 250 calling up information from a database and reproducing it in an acceptable manner for user 249 .
[0058] Instead of personal assistant 250 , a domestic appliance (not shown), in particular a washing machine, a stove, an oven, a microwave oven or a dishwasher, can also be provided to be controlled accordingly.
[0059] Figure 7 An embodiment is shown in which a control system 40 is used to control a medical imaging system 500, such as an MRT, X-ray, or ultrasound device. Sensor 30 can be provided, for example, by an imaging sensor, and display unit 10a is controlled by control system 40. For example, neural network 60 can determine whether an area recorded by the imaging sensor is conspicuous and then select control signal A so that the display unit 10a displays this area in a highlighted color.
[0060] Figure 8 A possible structure of a monitoring unit 61 is shown. An input image x is fed to an encoder ENC, which determines a so-called latent variable z from it. The latent variable z has a smaller dimension than the input image x. This latent variable z is fed to a generator GEN, which generates a reconstructed image from it. In this embodiment, the encoder ENC and the generator GEN are respectively given by convolutional neural networks. Input image x and reconstructed image is fed to the discriminator DIS. The discriminator DIS has been trained to generate parameters as good as possible that characterize whether the image fed to the discriminator DIS is a real image or whether it was generated by the generator GEN. This is combined with Figure 10 The generator GEN is also a convolutional neural network.
[0061] When the input image x or the reconstructed image is fed to the generator GEN. Layer (where is a pre-given number) feature maps (English: "feature maps") are used or These feature maps are fed to the estimator BE, where the reconstruction error is In an alternative embodiment (not shown), it is also possible to select the reconstruction error as .
[0062] The outliers can then be Determine the components of the input image of the reference data record (e.g. the training data record with which the discriminator DIS and / or the generator GEN and / or the encoder ENC were trained) for which the reconstruction error is less than the determined reconstruction error E x If the outlier value A(x) is greater than a predefinable threshold value, the monitoring signal d is set to the value d=1, which signals that the output variable y may not be reliably determined. Otherwise, the monitoring signal d is set to the value d=0, which signals that the determination of the output variable y is classified as reliable.
[0063] Figure 9 A possible structure of a first training device 141 for training the monitoring unit 51 is shown. The first training device is parameterized with parameters θ, which are provided by a parameter memory P. The parameters θ include the generator parameters of the parameterized generator GEN , encoder parameters of parameterized encoder ENC and the discriminator parameters of the parameterized discriminator DIS .
[0064] The training device 141 comprises a provider 71 which provides an input image e from a training data record. The input image e is fed to the monitoring unit 61 to be trained, which determines the output variable a therefrom. The output variable a and the input image e are fed to an evaluator 74 which, for example, combines Figure 10 Determine new parameters from this as described , the new parameters are transferred to the parameter memory P and replace the parameters θ there.
[0065] The method performed by training device 141 may be implemented as a computer program stored on machine-readable storage medium 146 and executed by processor 145 .
[0066] Figure 10The diagram illustrates the interaction of the generator GEN, the encoder ENC and the discriminator DIS during training. The arrangement of the generator GEN, the encoder ENC and the discriminator DIS shown here is also referred to herein as a neural network system.
[0067] First, the discriminator DIS is trained. The following steps for training the discriminator DIS can be repeated, for example times, among which is a predeterminable integer.
[0068] First, a batch of real input images x is provided. These real input images are said to have a (generally unknown) probability distribution of These input images are real images provided, for example, from a database. The entire set of these input images is also called a training data record.
[0069] In addition, a set of latent variables z is obtained from the probability distribution Randomly extracted from In this case, the probability distribution An example is the (multidimensional) standard normal distribution.
[0070] In addition, a set of random variables are drawn from the probability distribution Randomly extracted from In this case, the probability distribution For example, it is a uniform distribution on the interval [0;1].
[0071] The latent variable z is fed to the generator GEN and given the constructed input image , that is,
[0072] .
[0073] Use the random variable ϵ between the input image x and the constructed input image Interpolate between
[0074] .
[0075] The discriminator cost function is then determined using a predeterminable gradient coefficient λ
[0076] ,
[0077] The gradient coefficient can be chosen to be, for example, λ = 10. The new discriminator parameters It can be determined from
[0078] ,
[0079] Here, “Adam” represents the gradient descent method. Thus, the training of the discriminator DIS is completed.
[0080] Subsequently, the generator GEN and encoder ENC are trained. Here, the real input image is also used. and randomly selected latent variables Provided. Re-determined
[0081] . Determine the reconstruction latent variables , which is done by reconstructing the image Sent to the encoder ENC, that is
[0082] .
[0083] like Figure 8 As illustrated in the figure, we also try to reconstruct the input image x with the help of the encoder ENC and the generator GEN, that is,
[0084] .
[0085] Now, the generator cost function , the reconstruction cost function of the input image x and the reconstruction cost function of the latent variable z Determined to be
[0086] .
[0087] Then, the new generator parameters and the new encoder parameters Determined to be
[0088] .
[0089] New generator parameters , new encoder parameters and the new discriminator parameters Then replace the generator parameters , encoder parameters and the discriminator parameters .
[0090] At this point, the convergence of the parameters θ can be checked and, if necessary, the training of the discriminator DIS and / or the generator GEN and encoder ENC can be repeated until convergence is achieved. The method ends here.
[0091] Figure 11An exemplary second training device 140 for training a neural network 60 is shown. The training device 140 comprises a provider 72 which provides an input image x and a desired output variable ys, for example a desired classification. The input image x is fed to the artificial neural network 60 to be trained, which determines the output variable y therefrom. The output variable y and the desired output variable ys are fed to a comparator 75 which determines a new parameter from these based on the consistency of the respective output variable y and the desired output variable ys. , the new parameters are transferred to the parameter memory P and replace the parameters there .
[0092] The methods performed by training system 140 may be implemented as computer programs stored on machine-readable storage medium 148 and executed by processor 147 .
[0093] The data record comprising the input image x and the associated desired output variable ys can be enhanced or generated (for example by the provider 72) as follows. First, a data record comprising the input image x is provided. These input images are classified according to predeterminable representations of features (designated "A" and "B" by way of example), e.g., vehicles can be classified according to the feature "headlights on" or "headlights off," or identified cars can be classified according to the type "limousine" or "station wagon." For example, different representations of the feature "hair color" are also possible for identified pedestrians. Depending on which representation this feature has, the input images are sorted. Divide into two sets, namely and Advantageously, these sets are also homogenized by the fact that for a predeterminable set of further features, preferably all other features, the input image have the same form X, that is,
[0094]
[0095] With the help of encoder ENC, for the input image Each of the latent variables in .
[0096] Then the average value of the latent variable is determined on the set, that is,
[0097] .
[0098] The difference in means is then formed, i.e.
[0099] .
[0100] Now for theA The image is scaled using a predeterminable scaling factor Determine a new latent variable, the scaling factor can take a value between 0 and 1, i.e.
[0101]
[0102] Accordingly, for the B The new latent variables can be constructed as
[0103] .
[0104] This can be done with the help of
[0105]
[0106] Generate new images .
[0107] Of course, it is not necessary to classify the entire image. For example, it is possible to classify image segments as objects using a detection algorithm, then cut out these image segments and, if necessary, generate new image segments (corresponding to the new image segments). ) and inserted into the associated image at the location of the cut-out image segment. In this way, it is possible, for example, to selectively adapt the hair color of a detected pedestrian in an image having the pedestrian.
[0108] Apart from the classification of the features that thus change between the representation forms "A" and "B", the associated target output variable ys can be taken over unchanged. Thus, an enhanced data record can be generated and the neural network 60 can be trained. The method ends here.
[0109] The term "computer" includes any device for executing predefinable calculation rules. These calculation rules can exist in the form of software or hardware or in a hybrid form consisting of software and hardware.
Claims
1. A method for training a machine learning system (60), the method comprising generating an input image ( ) is used to train the machine learning system (60) using the enhanced data record of the input image (x), the machine learning system being configured to classify and / or semantically segment the input image (x), the machine learning system comprising a first machine learning system and a second machine learning system, the first machine learning system being configured as a decoder of an autoencoder, the second machine learning system being configured as an encoder of an autoencoder, wherein the encoders are used to respectively extract the input image (x) from the decoder. ) to determine the latent variables ( ), where for the input image ( ) is classified according to the determined characteristic representation of its image data, and wherein at least two of the categories are selected from the input image ( ) in at least one input image according to the determined latent variable ( ) of the average value ( 、 ) determines the enhanced input image of the enhanced data record ( ), where the image category is selected so that the input image ( ) with respect to its manifestation being consistent in a predeterminable set of other characteristics, wherein the enhanced input image is determined according to the determined enhancement latent variable by means of the decoder, and, The enhanced latent variable is thereby determined from a difference between a predefinable latent variable of the determined latent variables and the mean value.
2. The method of claim 1 , wherein the first machine learning system is a first neural network.
3. The method of claim 1 , wherein the second machine learning system is a second neural network.
4. The method according to claim 1 , wherein the decoder is used to determine the enhanced latent variable ( ) determines the enhanced input image ( ).
5. The method according to claim 4, wherein the latent variables ( ) in the predeterminable latent variables and the mean value ( 、 ) difference ( ) to determine the enhanced latent variable ( ).
6. The method according to claim 5 , wherein a predeterminable weighting factor ( ) for the difference ( ) for weighting.
7. A method according to any one of claims 1 to 3, wherein the generated enhanced data record is used to check whether the machine learning system (60) is robust and, in connection with this, further training is carried out when the check has shown that the machine learning system (60) is not robust.
8. The method according to any one of claims 1 to 3, wherein the machine learning system (60) is trained using the generated augmented data records when monitoring has concluded that the machine learning system (60) is not robust.
9. The method according to claim 8, wherein the monitoring of the machine learning system (60) is performed by means of a monitoring unit (61), the monitoring unit comprising a first machine learning system and a second machine learning system, wherein the input image (x) is fed to the second machine learning system, the second machine learning system determines low-dimensional latent variables (z) therefrom, and the first machine learning system determines a reconstruction of the input image (x) from the low-dimensional latent variables (z). ), where the input image (x) and the reconstructed input image ( ) determines whether the machine learning system (60) is robust.
10. The method according to claim 9, wherein the monitoring unit (61) further comprises a third machine learning system of a neural network system, The neural network system comprises the first machine learning system, the second machine learning system and the third machine learning system, wherein the first machine learning system is configured to determine the constructed higher-dimensional image ( , wherein the second machine learning system is constructed to again generate a model from the constructed higher-dimensional image ( ), and wherein the third machine learning system is configured to distinguish whether an image fed to the third machine learning system is a real image (x), Wherein whether the machine learning system (60) is robust is determined according to the following: when an input image (x) is fed to a third machine learning system, an activation ( ) takes which value ( ) and when the reconstructed input image ( ) is supplied to the third machine learning system, the one activation ( ) takes which value ( ).
11. The method of claim 10, wherein the third machine learning system is a third neural network.
12. The method according to claim 10, wherein the first machine learning system is trained as follows: when a real image (x) or an image of the real image reconstructed by the second machine learning system and the first machine learning system is presented to the user, the user is trained. ) is supplied to the third machine learning system, an activation ( ) take the same value.
13. The method according to claim 12, wherein the first machine learning system is also trained so that the third machine learning system does not recognize, as far as possible, images generated by the first machine learning system and fed to the third machine learning system ( ) is not a real image.
14. The method according to claim 12 or 13, wherein the second machine learning system is trained such that the reconstruction ( ) is as equal to the latent variable (z) as possible.
15. The method according to any one of claims 12 to 13, wherein the third machine learning system is trained as follows, that is, the third machine learning system recognizes as much as possible: the image generated by the first machine learning system and fed to the third machine learning system ( ) is not a real image. 16 . The method according to claim 15 , wherein the third machine learning system is also trained such that the third machine learning system recognizes as closely as possible that a real image (x) supplied to the third machine learning system is a real image.
17. The method according to claim 15, wherein the machine learning system (60) and the neural network system have been trained using data records comprising the same input image (x). 18 . A training device ( 140 , 141 ) configured to carry out the method according to claim 1 . 19 . A computer program product comprising a computer program configured to carry out the method according to claim 1 .
20. A machine-readable storage medium (146, 148) on which the computer program according to claim 19 is stored.
Citation Information
Patent Citations
Visual search target decoding method based on image generation model
CN107516113A
Method for detecting an anomalous image among a first dataset of images using an adversarial autoencoder
CN109741292A