Devices used for robust classification and regression of time series data.
By employing adversarial training methods, the worst-case training time series is determined and superimposed on a noise signal to optimize the parameters of the machine learning system. This addresses the problem of decreased prediction accuracy under noise interference and enhances the system's robustness in noisy environments.
Patent Information
- Application Number
- CN202180086134.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-21
- Filing Date
- 2021-12-09
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-12-09
AI Technical Summary
When machine learning systems process time series of sensor signals, noise interference leads to a decrease in prediction accuracy, and existing technologies make it difficult to train the system to be robust to noise.
An adversarial training method is adopted to optimize the parameters of the machine learning system by determining the worst possible training time series and superimposing it with a noisy signal, making it more robust to noise.
It improves the prediction accuracy of machine learning systems in noisy environments and enhances their robustness against noise.
Smart Images

Figure CN116670669B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computer-implemented machine learning system, a training device for training the machine learning system, a computer program, and a machine-readable storage medium. Background Technology
[0002] EP 19174931.6 discloses a method for robustly training machine learning systems relative to adversarial examples.
[0003] Advantages of the invention
[0004] Sensor recordings are typically subjected to varying degrees of strong noise, which is reflected in the sensor signals determined by the sensors. When such sensor signals are processed automatically using machine learning systems, this noise is a typical source of interference and can significantly degrade the prediction accuracy of the machine learning system. Noise can have a particularly strong negative impact on prediction accuracy when processing time series of sensor signals.
[0005] Therefore, it is desirable to train a machine learning system for processing time series data to be robust to noise. An advantage of a machine learning system having the features of claim 1 is that it becomes more robust to noise due to its construction. Surprisingly, the inventors were able to determine that adversarial training methods can also be used to train a machine learning system to be robust to noise. Summary of the Invention
[0006] In a first aspect, the present invention relates to a computer-implemented machine learning system (60), wherein the machine learning system is configured to determine an output signal based on a time series of an input signal of a technical system, the output signal representing a classification and / or regression result of at least one first operating state and / or at least one first operating variable of the technical system, wherein training of the machine learning system includes the following steps:
[0007] a. Determine a first training time series of an input signal and a desired training output signal corresponding to the first training time series from multiple training time series, wherein the desired training output signal characterizes the desired classification and / or desired regression result of the first training time series;
[0008] b. Determine the worst possible training time series, wherein the worst possible training time series represents the superposition of the first training time series and the determined first noise signal;
[0009] c. Using the machine learning system, determine the training output signal based on the worst-case training time series;
[0010] d. Adapt at least one parameter of the machine learning system according to the gradient of the loss value, wherein the loss value characterizes the deviation between the desired output signal and the determined training output signal.
[0011] Each input signal in the time series can preferably characterize a second operating state and / or a second operating variable of the technical system at a predefined time point. The input signals can be recorded, in particular, by means of sensors, especially the sensors of the technical system. In this case, the first operating state or the first operating variable can particularly characterize the temperature and / or pressure and / or voltage and / or force and / or speed and / or rotational speed and / or torque of the technical system.
[0012] The machine learning system can therefore also be understood as a virtual sensor, which can derive a first operating state or first operating variable from multiple second operating states or second operating variables.
[0013] Training a machine learning system can be understood as supervised training. The first training time series used for training may preferably include input signals, each input signal representing a second operating state and / or a second operating variable of the technical system, or a technical system with the same or similar structure, or a simulation of the second operating state and / or the second operating variable at a predefined time point. In other words, the training time series among multiple training time series may be based on the input signals of the technical system itself. Alternatively or additionally, training time series of input signals from similar technical systems may be recorded, wherein the similar technical system may, for example, be a prototype or initial development of the technical system. The input signals of the training time series can also be determined from other technical systems, such as other technical systems from the same one or more production sequences. The input signals of the training time series can also be determined based on simulations of the technical system.
[0014] Typically, the input signal of the first training time series is similar to the input signal of the time series; in particular, the input signal of the training time series should represent the same second running variable as the input signal of the time series.
[0015] For training purposes, training time series can be provided from a database, which includes multiple training time series. The machine learning system can preferably execute step ad iteratively. Preferably, multiple training time series can also be used in each iteration to determine the loss value, i.e., training can be performed using a batch of training time series.
[0016] The output signal may include classification and / or regression results. Regression results should be understood in this context as the outcome of regression. Therefore, a machine learning system can be viewed as a classifier and / or a regressor. A regressor can be understood as a device that predicts at least one true value with respect to at least one true value.
[0017] The time series and the training time series are preferably represented as column vectors, where each dimension of the vector represents a measurement value of the time series or the training time series at a specific time point.
[0018] The worst-case training time series can be understood as the training time series that occurs when a first training time series is superimposed with a noise signal, such that the distance between the training output of the machine learning system for the superimposed training time series and the training output determined by the machine learning system for the first training time series becomes as large as possible. Specifically, the noise can be constrained with appropriate boundary conditions so that the worst-case training time series is not a negligible result of the superposition. In the described invention, the noise signal is specifically constrained to correspond to a desired noise signal. The desired noise signal can be understood, in particular, based on multiple training time series. In this sense, the method can be understood as a form of adversarial training, wherein the adversarial training is advantageously constrained to noise characterizing the training time series. The inventors have found that adversarial training in this way unexpectedly and advantageously leads to a more robust machine learning system to noise.
[0019] Preferably, in step b, the first noise signal is determined by optimization, such that the distance between the second output signal and the desired output signal is increased, wherein the second output signal is determined by the machine learning system based on the superposition of the training time series and the first noise signal.
[0020] The noise signal can exist, in particular, as a vector, where the vector has the same dimension as the vector form of the first training time series. Thus, the superposition can be, for example, the sum of the vector of the first training time series and the vector of the noise signal. Optimization here can be understood as mathematical optimization under boundary conditions. Specifically, the expected noise signal can be introduced as a boundary condition into the method.
[0021] Therefore, in a preferred design of the machine learning system, the first noise signal is determined in step b based on the expected noise values of the plurality of training time series, wherein the expected noise values characterize the average noise intensity of the training time series.
[0022] Specifically, the expected noise value can be the average distance between one of the plurality of training time series and the corresponding denoised training time series.
[0023] In a preferred design of the machine learning system, it can be based on the formula
[0024]
[0025] Determine the expected noise value, where n is the number of training time series in the multiple training time series, and z i It is for training time series x i The training time series for denoising, ||·||2 is the Euclidean norm.
[0026] This can be understood as first denoising the training time series, then determining the distance between the original training time series and the denoised training time series. The average distance between all or at least some of the training time series can then be interpreted as the expected noise. Therefore, the expected noise can be understood as a scalar value.
[0027] Preferably, it can be based on the formula
[0028]
[0029] Determine the training time series for denoising, where It is a pseudo-inverse covariance matrix.
[0030] In this case, the pseudo-inverse covariance matrix can be determined through the following steps:
[0031] e. Determine the second covariance matrix, wherein the second covariance matrix is the plurality of training time series (x i The covariance matrix of ).
[0032] f. Determine a predefined set of maximum eigenvalues and corresponding eigenvectors of the second covariance matrix;
[0033] g. Determine the pseudo-inverse covariance matrix according to the following formula.
[0034]
[0035] Where λ i It is the i-th eigenvalue among multiple maximum eigenvalues, and k is the number of maximum eigenvalues among the predefined multiple maximum eigenvalues.
[0036] The pseudo-inverse covariance matrix can be understood as part of the noise model. As mentioned above, the pseudo-inverse covariance matrix can be used to analyze the first training time series x. i Denoising, thereby determining the denoising training time series z. i Then the distance between the first training time series and the denoised training time series can be understood as the noise value of the first training time series.
[0037] Therefore, the multiple maximal eigenvalues include a predefined number of eigenvalues, of which only the maximal eigenvalue of the covariance matrix is included in the multiple maximal eigenvalues.
[0038] In this case, the eigenvector can be understood as a column vector.
[0039] In a preferred design of a machine learning system, the first noise signal can be determined based on the provided adversarial perturbation, wherein the provided adversarial perturbation is limited according to the expected noise value.
[0040] Adversarial perturbation can be understood as a perturbation that generates adversarial examples by superimposing the corresponding training time series with the perturbation.
[0041] In a preferred design of the machine learning system, the adversarial perturbation is constrained to ensure that the noise value of the adversarial perturbation is no greater than the expected noise value. The adversarial perturbation can preferably be provided according to the following steps:
[0042] h. Provide the first counter-perturbation;
[0043] i. Determine a second adversarial perturbation, wherein the second adversarial perturbation is stronger than the first adversarial perturbation;
[0044] j. If the distance between the second adversarial perturbation and the first adversarial perturbation is less than or equal to a predefined threshold, then the second adversarial perturbation is provided as an adversarial perturbation;
[0045] k. Otherwise, if the noise value of the second adversarial perturbation is less than or equal to the expected noise value, then step i is performed, wherein the second adversarial perturbation is used as the first adversarial perturbation when step i is performed;
[0046] l. Otherwise, determine the planned disturbance and perform step j, wherein the planned disturbance is used as a second adversarial disturbance when performing step j, and wherein the planned disturbance is determined by optimization such that the distance between the planned disturbance and the second adversarial disturbance is as small as possible, and the noise value of the planned disturbance is equal to the expected noise value.
[0047] The first adversarial perturbation can be randomly determined or contain at least one predefined value. Since the adversarial perturbation preferably exists in vector form, the first adversarial perturbation in step h can be, for example, a zero vector or a random vector.
[0048] If the distance between the second training output signal determined with respect to the training time series superimposed with the second adversarial perturbation and the expected training output signal of the training time series is greater than the distance between the first training output signal determined with respect to the training time series superimposed with the first adversarial perturbation and the expected training output signal of the training time series.
[0049] According to the formula
[0050]
[0051] Determine the noise value of the adversarial disturbance, where δ is the adversarial disturbance.
[0052] Preferably, in step i, the formula can be used.
[0053] δ2=δ1+α·C k ·g
[0054] Determine the second adversarial perturbation, where δ1 is the first adversarial perturbation, α is a predefined increment value, and C k is the first covariance matrix, and g is the gradient.
[0055] This expression can be understood as an adaptation of projected gradient descent, where the gradient is adapted to fit the noise model. The inventors were able to determine that the noise signal determined in this way is significantly closer to the real noise signal than the noise signal determined by normal projected gradient descent. Due to this improved noise signal, machine learning systems can become significantly more robust to expected noise.
[0056] According to the formula
[0057]
[0058] Determine the gradient g, where L is the loss function and t i It is the expected training output signal, f(x), relating to the training time series. i +δ1) is the result of the machine learning system when the training time series superimposed with the first adversarial perturbation δ1 is passed to the machine learning system.
[0059] According to the formula
[0060]
[0061] Determine the first covariance matrix.
[0062] According to the formula
[0063]
[0064] Identify the adversarial disturbances in the plan.
[0065] Furthermore, it is possible that the output signal represents the regression of at least one first operating state and / or at least one first operating variable of the system, wherein the loss value represents the squared Euclidean distance between the determined training output and the expected training output.
[0066] Specifically, the technical system may be an injection device for an internal combustion engine, and each input signal of the time series represents at least one pressure value or average pressure value of the injection device (e.g., a common rail diesel engine), and the output signal represents the amount of fuel injected, wherein each output signal of the training time series also represents at least one pressure value or average pressure value of a simulated internal combustion engine or an internal combustion engine with the same or similar structure, and the desired training output signal represents the amount of fuel injected.
[0067] Alternatively, the technical system may be a manufacturing machine that manufactures at least one workpiece, wherein each input signal of the time series represents the force and / or torque of the manufacturing machine, and the output signal represents a classification of whether the workpiece is correctly manufactured. Furthermore, each input signal of the training time series represents a simulated force and / or torque of the manufacturing machine or a manufacturing machine of the same or similar structure, and the desired training output signal is a classification of whether the workpiece is correctly manufactured.
[0068] In another aspect, the present invention relates to a training apparatus configured to train the machine learning system corresponding to steps a to d. Attached Figure Description
[0069] Embodiments of the present invention will now be explained in more detail with reference to the accompanying drawings. In the drawings:
[0070] Figure 1 The training system used to train the classifier is illustrated schematically.
[0071] Figure 2 The structure of a control system that uses a classifier to control the actuator is schematically shown;
[0072] Figure 3 An embodiment for controlling a manufacturing system is illustrated schematically;
[0073] Figure 4 An embodiment for controlling the injection system is illustrated schematically. Detailed Implementation
[0074] Figure 1An embodiment of a training system (140) for training a machine learning system (60) using a training dataset (T) is shown. Preferably, the machine learning system (60) includes a neural network. The training dataset (T) includes multiple training time series (x_t) of input signals from sensors of the technical system. i ), where the training time series (x) i ) is used to train the machine learning system (60), wherein the training dataset (T) is also used for each training time series (x) i This includes the expected training output signal (t) i The expected training output signal corresponds to the training time series (x). i And characterizes the training time series (x) i The classification and / or regression results of the training time series (x). i It is preferable to exist in the form of a vector, where each dimension represents the training time series (x). i (The point in time)
[0075] For training, the training data unit (150) accesses a computer-implemented database (St2), where the database (St2) makes the training dataset (T) available. The training data unit (150) first obtains data from multiple training time series (x... i The first covariance matrix is determined in the training data unit (150). To this end, the training time series (x) is first determined. i The empirical covariance matrix is then determined. The k largest eigenvalues and their associated eigenvectors are then identified, according to the formula...
[0076]
[0077] Determine the first covariance matrix C k , where λ i Belongs to the k largest eigenvalues, v i It is a columnar form belonging to λ i The eigenvectors of , where k is a predefined value. Additionally, according to the formula...
[0078]
[0079] Determine the pseudo-inverse covariance matrix Furthermore, according to the formula
[0080]
[0081] Determine the expected noise value Δ, where n is the training time series (x) in the training dataset (T). i The quantity of ).
[0082] Then, the training data unit (150) preferably randomly determines at least one first training time series (x) from the training dataset (T). i ) and corresponding to the training time series (x) i The expected training output signal (t) i Then, the training data unit (150) determines the worst possible training time series (x) based on the machine learning system (60) according to the following steps. i ′):
[0083] m. provides a first adversarial perturbation δ1, wherein the first adversarial perturbation is selected with respect to the first training time series (x). i Zero vectors of the same dimension;
[0084] n. According to the formula
[0085]
[0086] Determine the gradient g, where f(x) i +δ1) is the superposition output of the machine learning system (60) with respect to the first training time series;
[0087] o. According to the formula
[0088] δ2=δ1+α·C k ·g
[0089] Determine the second adversarial perturbation, where α is a predefined increment;
[0090] p. If the Euclidean distance between the second adversarial perturbation and the first adversarial perturbation is less than or equal to a predefined threshold, then the second adversarial perturbation is provided as an adversarial perturbation δ.
[0091] q. Otherwise, if the noise value of the second adversarial perturbation
[0092]
[0093] If the noise value is less than or equal to the expected noise value Δ, then proceed to step n, where the second adversarial perturbation is used as the first adversarial perturbation when proceeding to step n;
[0094] r. Otherwise, according to the formula
[0095]
[0096] Determine the planned disturbance and execute step p, wherein the planned disturbance is used as a second adversarial disturbance when executing step p.
[0097] Then, based on the provided adversarial perturbation, according to the formula...
[0098] x′i =x i +δ
[0099] Determine the worst possible training time series (x) i ′).
[0100] Then, the machine learning system (60) transmits the worst-case training time series (x) i ′), and the machine learning system trains the time series (x) on the worst-case scenario. i ′) Determine the training output signal (y i ).
[0101] The desired training output signal (t) i ) and the determined training output signal (y i ) is transferred to the change unit (180).
[0102] Then, the unit (180) is changed based on the desired training output signal (t). i ) and the determined output signal (y) i To determine new parameters (Φ′) for the machine learning system (60), the unit (180) is modified by using a loss function to adjust the desired training output signal (t). i ) and the determined training output signal (y i The comparison is performed. The loss function determines the representation of the determined training output signal (y). i ) and the expected training output signal (t) i The first loss value is determined by how far the difference is. In this embodiment, the negative log-likelihood function is chosen as the loss function. In alternative embodiments, other loss functions may also be considered.
[0103] The changing unit (180) determines new parameters (Φ′) based on the first loss value. In this embodiment, this is accomplished by means of gradient descent, preferably stochastic gradient descent, Adam, or AdamW.
[0104] The determined new parameters (Φ′) are stored in the model parameter memory (St1). Preferably, the determined new parameters (Φ′) are provided as parameters (Φ) to the classifier (60).
[0105] In a further preferred embodiment, the described training iteratively repeats a predefined number of iteration steps or iteratively repeats until a first loss value is below a predefined threshold. Alternatively or additionally, it may be envisioned that training terminates when the average first loss value associated with the test or validation dataset is below a predefined threshold. In at least one of the said iterations, new parameters (Φ′) determined in previous iterations are used as parameters (Φ) of the classifier (60).
[0106] Furthermore, the training system (140) may include at least one processor (145) and at least one machine-readable storage medium (146) containing instructions that, when executed by the processor (145), cause the training system (140) to perform a training method according to one aspect of the invention.
[0107] Figure 2 A control system (40) is shown, which controls the actuator (10) of the technology system by means of a machine learning system (60), wherein the machine learning system (60) has been trained by means of a training device (140). A second operating variable or a second operating state is detected by a sensor (30) at time intervals according to a preferred rule. The input signal (S) detected by the sensor (30) is transmitted to the control system (40). The control system (40) thus receives the sequence of input signals (S). The control system (40) determines from this the control signal (A) to be transmitted to the actuator (10).
[0108] The control system (40) receives a sequence of input signals (S) from the sensor (30) in a receiving unit (50), which converts the sequence of input signals (S) into a time series (x). This can be done, for example, by sorting the input signals (S) of a predefined number of last recorded input signals (S). In other words, the time series (x) is determined based on the input signals (S). The sequence of input signals (x) is then fed into a machine learning system (60).
[0109] The machine learning system (60) determines an output signal (y) from the time series (x). The output signal (y) is fed to an optional shaping unit (80), which determines from it a control signal (A) to be sent to the actuator (10) to control the actuator (10) accordingly.
[0110] The actuator (10) receives the control signal (A), is controlled accordingly, and performs the corresponding action. In this case, the actuator (10) may include (not necessarily structurally integrated) control logic that determines a second control signal from the control signal (A) and then uses the second control signal to control the actuator (10).
[0111] In a further embodiment, the control system (40) includes a sensor (30). In a further embodiment, the control system (40) alternatively or additionally includes an actuator (10).
[0112] In a further preferred embodiment, the control system (40) includes at least one processor (45) and at least one machine-readable storage medium (46) on which instructions are stored, which, when executed on at least one processor (45), cause the control system (40) to perform the method according to the invention.
[0113] In an alternative implementation, a display unit (10a) is provided as an alternative to or supplement to the actuator (10).
[0114] Figure 3 One embodiment is shown in which a control system (40) manipulates a manufacturing machine (11) of a manufacturing system (200) by manipulating an actuator (10) that controls the manufacturing machine (11). The manufacturing machine (11) may be, for example, a welding machine.
[0115] The sensor (30) can preferably be a sensor (30) that determines the voltage of the welding apparatus of the manufacturing machine (11). The machine learning system (60) can be specifically trained to classify whether the welding process is successful based on the time series (x) of the voltage. In the event of an unsuccessful welding process, the actuator (10) can automatically reject the corresponding workpiece.
[0116] In an alternative embodiment, the manufacturing machine (11) may also use pressure to join the two workpieces. In this case, the sensor (30) may be a pressure sensor and the machine learning system (60) may determine whether the joining is correct.
[0117] Figure 4 An embodiment for controlling an injector (40) of an internal combustion engine is shown. In this embodiment, the sensor (30) is a pressure sensor that determines the pressure of the injection system (10) that supplies fuel to the injector (40). A machine learning system (60) can be configured, in particular, to accurately determine the amount of fuel injected based on a time series (x) of pressure values.
[0118] Then, the actuator (10) can be manipulated in the future injection process based on the determined injection amount, so that excessive or insufficient fuel injection is compensated accordingly.
[0119] In an alternative implementation, as an alternative or supplement to the control unit (40), at least one additional device (10a) is operated by means of a control signal (A). The device (10a) may, for example, be a pump of the common rail system to which the injector (20) belongs. Alternatively or additionally, the device may be a control device for an internal combustion engine. Alternatively or additionally, the device (10a) may be a display unit by means of which the amount of fuel determined by the machine learning system (60) can be displayed to a person (e.g., a driver or mechanic).
[0120] The term "computer" includes any device used to process pre-given computational rules. These computational rules can exist in software, hardware, or a hybrid of both.
[0121] Generally, a plurality can be understood as indexed, meaning that each element in the plurality is assigned a unique index, preferably by assigning consecutive integers to the elements contained in the plurality. Preferably, when the plurality includes N elements, where N is the number of elements in the plurality, these elements are assigned integers from 1 to N.
Claims
1. A computer-implemented machine learning system (60), wherein the machine learning system (60) is configured to determine an output signal (y) based on a time series (x) of an input signal of a technical system, the output signal representing a classification and / or regression result of at least one first operating state and / or at least one first operating variable of the technical system, wherein training the machine learning system (60) comprises the following steps: a. From multiple training time series (x i The first training time series (x) of the input signal is determined in ) i ) and the first training time series (x) i The expected training output signal (t) corresponds to i ), wherein the desired training output signal (t) i ) represents the first training time series (x) i The expected classification and / or expected regression results; b. Determine the worst-case training time series (x i '), wherein the worst possible training time series (x) i ') represents the first training time series (x i The worst possible training time series is understood as the training time series that occurs when the first training time series is superimposed on the noise signal such that the distance between the training output of the machine learning system for the training time series superimposed in this way and the training output determined for the first training time series becomes as large as possible. c. Using the machine learning system (60) based on the worst-case training time series (x) i ') Determine the training output signal (y) i ); d. Adapt at least one parameter of the machine learning system (60) according to the gradient of the loss value, wherein the loss value characterizes the desired training output signal (t). i ) and the determined training output signal (y i () deviation.
2. The machine learning system (60) according to claim 1, wherein in step b, the first noise signal is determined by optimization such that the distance between the second output signal and the desired output signal is increased, wherein the second output signal is determined by the machine learning system (60) based on the first training time series (x i The result is determined by superimposing the signal with the first noise signal.
3. The machine learning system (60) according to claim 1, wherein in step b, the plurality of training time series (x) are used as the basis for the machine learning system. i The expected noise value is used to determine the first noise signal, wherein the expected noise value characterizes the training time series (x). i The average noise intensity.
4. The machine learning system (60) according to claim 3, wherein the expected noise value is the plurality of training time series (x i A training time series (x) i The average distance between the training time series and the corresponding denoised training time series.
5. The machine learning system (60) according to claim 4, wherein according to the formula Determine the expected noise value, where n is the plurality of training time series (x i Training time series (x) i The quantity of z i It is for training time series x i The training time series for denoising, ||·||2 is the Euclidean norm.
6. The machine learning system (60) according to claim 5, wherein according to the formula Determine the training time series for the denoising, wherein It is a pseudo-inverse covariance matrix.
7. The machine learning system (60) of claim 6, wherein the pseudo-inverse covariance matrix is determined by the following steps: e. Determine the second covariance matrix, wherein the second covariance matrix is the plurality of training time series (x i The covariance matrix of ). f. Determine a predefined set of maximum eigenvalues and corresponding eigenvectors of the second covariance matrix; g. Determine the pseudo-inverse covariance matrix according to the following formula. Where λ i It is the i-th eigenvalue among multiple maximum eigenvalues, and k is the number of maximum eigenvalues among the predefined multiple maximum eigenvalues.
8. The machine learning system (60) according to any one of claims 3 to 7, wherein the first noise signal is determined based on the provided adversarial perturbation, wherein the provided adversarial perturbation is limited according to the expected noise value.
9. The machine learning system (60) of claim 8, wherein the adversarial perturbation is limited to such that the noise value of the adversarial perturbation is not greater than the expected noise value.
10. The machine learning system (60) according to claim 9, wherein according to the formula Determine the noise value of the adversarial disturbance, where δ is the adversarial disturbance.
11. The machine learning system (60) of claim 8, wherein the adversarial perturbation is provided according to the following steps: h. Provide the first counter-perturbation; i. Determine a second adversarial perturbation, wherein the second adversarial perturbation is relative to the first training time series (x i It is stronger than the first adversarial perturbation; j. If the distance between the second adversarial perturbation and the first adversarial perturbation is less than or equal to a predefined threshold, then the second adversarial perturbation is provided as an adversarial perturbation; k. Otherwise, if the noise value of the second adversarial perturbation is less than or equal to the expected noise value, then step i is performed, wherein the second adversarial perturbation is used as the first adversarial perturbation when step i is performed; 1. Otherwise, determine the planned perturbation and perform step j, wherein the planned perturbation is used as a second adversarial perturbation when performing step j, and wherein the planned perturbation is determined by optimization such that the distance between the planned perturbation and the second adversarial perturbation is as small as possible, and the noise value of the planned perturbation is equal to the expected noise value.
12. The machine learning system (60) according to claim 11, wherein the first adversarial perturbation is randomly determined in step h.
13. The machine learning system (60) of claim 11, wherein the first adversarial perturbation in step h comprises at least one predefined value.
14. The machine learning system (60) according to claim 11, wherein in step i, the formula is used... δ2=δ1+α·C k ·g The second adversarial perturbation is determined, where δ1 is the first adversarial perturbation, α is a predefined increment value, and C k is the first covariance matrix, and g is the gradient.
15. The machine learning system (60) according to claim 14, wherein according to the formula Determine the gradient g, where L is the loss function and t i It is about the first training time series (x) i The expected training output signal (t) i ), f(x) i +δ1) is the first training time series (x) superimposed on the first adversarial perturbation δ1, which is transmitted to the machine learning system (60). i The results of the machine learning system (60) at that time.
16. The machine learning system (60) according to claim 14, wherein according to the formula Determine the first covariance matrix.
17. The machine learning system (60) according to claim 11, wherein in step 1, the formula is used... Identify the adversarial perturbations in the proposed plan.
18. The machine learning system (60) according to any one of claims 1 to 7, wherein each input signal characterizes the temperature and / or pressure and / or voltage and / or force and / or speed and / or rotational speed and / or torque of the technical system.
19. The machine learning system (60) of claim 18, wherein each input signal is recorded by at least one sensor (30).
20. The machine learning system (60) according to any one of claims 1 to 7, wherein each input signal of the time series (x) characterizes a second operating state and / or a second operating variable of the technical system at a predefined time point, and the first training time series (x) i Each input signal represents a second operating state and / or a second operating variable of the technical system or a technical system with the same or similar structure, or a simulation of the second operating state and / or the second operating variable at a predefined time point.
21. The machine learning system (60) according to any one of claims 1 to 7, wherein the output signal (y) characterizes at least one first operating state and / or at least one first operating variable of the technical system, wherein the loss value characterizes the determined training output (y). i ) and the expected training output (t) i The square Euclidean distance between them.
22. The machine learning system (60) of claim 21, wherein the system is an injection device for an internal combustion engine and each input signal of the time series (x) represents at least one pressure value or average pressure value of the injection device, and the output signal (y) represents the amount of fuel injected, wherein the first training time series (x) i Each output signal of the internal combustion engine also characterizes at least one pressure value or average pressure value of the internal combustion engine or an internal combustion engine with the same or similar structure, or a simulated internal combustion engine, and the desired training output signal (y) i This characterizes the amount of fuel injected.
23. The machine learning system (60) according to any one of claims 1 to 7, wherein the technical system is a manufacturing machine for manufacturing at least one workpiece, wherein each input signal of the time series (x) characterizes the force and / or torque of the manufacturing machine, and the output signal (y) characterizes a classification of whether the workpiece is correctly manufactured, wherein the first training time series (x) i Each input signal also characterizes the simulated force and / or torque of the manufacturing machine or a manufacturing machine of the same or similar structure, and the desired training output signal (y) i () is a classification of whether a workpiece has been correctly manufactured.
24. The machine learning system (60) according to any one of claims 1 to 7, wherein the machine learning system (60) determines the output signal (y) by means of a neural network.
25. The machine learning system (60) according to claim 24, wherein the neural network is a recurrent neural network.
26. The machine learning system (60) according to claim 24, wherein the machine learning system (60) is a convolutional neural network (CNN).
27. The machine learning system (60) of claim 24, wherein the neural network is a transformer.
28. The machine learning system (60) of claim 24, wherein the neural network is a multilayer perceptron.
29. A training device configured to train a machine learning system (60), wherein the machine learning system (60) is configured to determine an output signal (y) based on a time series (x) of an input signal of a technical system, the output signal representing a classification and / or regression result of at least one first operating state and / or at least one first operating variable of the technical system, wherein training the machine learning system (60) comprises the following steps: a. From multiple training time series (x i The first training time series (x) of the input signal is determined in ) i ) and the first training time series (x) i The expected training output signal (t) corresponds to i ), wherein the desired training output signal (t) i ) represents the first training time series (x) i The expected classification and / or expected regression results; b. Determine the worst-case training time series (x i '), wherein the worst possible training time series (x) i ') represents the first training time series (x i The worst possible training time series is understood as the training time series that occurs when the first training time series is superimposed on the noise signal such that the distance between the training output of the machine learning system for the training time series superimposed in this way and the training output determined for the first training time series becomes as large as possible. c. Using the machine learning system (60) based on the worst-case training time series (x) i ') Determine the training output signal (y) i ); d. Adapt at least one parameter of the machine learning system (60) according to the gradient of the loss value, wherein the loss value characterizes the desired training output signal (t). i ) and the determined training output signal (y i () deviation.
30. A computer program product having a computer program, said computer program being configured to perform the following steps a to d when the computer program is executed by a processor (45, 145): a. From multiple training time series (x i The first training time series (x) of the input signal is determined in ) i ) and the first training time series (x) i The expected training output signal (t) corresponds to i ), wherein the desired training output signal (t) i ) represents the first training time series (x) i The expected classification and / or expected regression results; b. Determine the worst-case training time series (x i '), wherein the worst possible training time series (x) i ') represents the first training time series (x i The worst possible training time series is understood as the training time series that occurs when the first training time series is superimposed on the noise signal such that the distance between the training output of the machine learning system for the training time series superimposed in this way and the training output determined for the first training time series becomes as large as possible. c. Using the machine learning system (60) based on the worst-case training time series (x) i ') Determine the training output signal (y) i ); d. Adapt at least one parameter of the machine learning system (60) according to the gradient of the loss value, wherein the loss value characterizes the desired training output signal (t). i ) and the determined training output signal (y i () deviation.
31. A machine-readable storage medium (46, 146) having a computer program stored thereon, the computer program being configured to perform the following steps a to d when executed by a processor (45, 145): a. From multiple training time series (x i The first training time series (x) of the input signal is determined in ) i ) and the first training time series (x) i The expected training output signal (t) corresponds to i ), wherein the desired training output signal (t) i ) represents the first training time series (x) i The expected classification and / or expected regression results; b. Determine the worst-case training time series (x i '), wherein the worst possible training time series (x) i ') represents the first training time series (x i The worst possible training time series is understood as the training time series that occurs when the first training time series is superimposed on the noise signal such that the distance between the training output of the machine learning system for the training time series superimposed in this way and the training output determined for the first training time series becomes as large as possible. c. Using the machine learning system (60) based on the worst-case training time series (x) i ') Determine the training output signal (y) i ); d. Adapt at least one parameter of the machine learning system (60) according to the gradient of the loss value, wherein the loss value characterizes the desired training output signal (t). i ) and the determined training output signal (y i () deviation.
Citation Information
Patent Citations
Classification robust against multiple perturbation types
US20200364616A1