METHOD FOR SECURELY TRAINING A DYNAMIC MODEL

DE502019013276D1Active Publication Date: 2025-05-22ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE502019013276
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-09-05
Filing Date
2019-08-12
Publication Date
2025-05-22
Estimated Expiration
2039-08-12

AI Technical Summary

Technical Problem

Existing methods for active learning in time series models of physical systems lack a comprehensive approach to combine dynamic exploration, active exploration, and safety considerations, potentially leading to system damage during dynamic stimulation.

Method used

The proposed procedure integrates dynamic exploration, active exploration, and safe exploration by using a Gaussian process with a non-linear exogenous structure to generate input and output curves, while incorporating a safety criterion to ensure the system is not damaged during exploration.

Benefits of technology

This approach allows for efficient information gathering in a safe manner, maximizing information gain while preventing system damage, thus enabling precise modeling of dynamic systems without excessive resource consumption or risk.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method with safety condition for active learning for modeling dynamic systems with the aid of time series based on Gaussian processes, a system trained with this method, a computer program comprising instructions configured to carry out the method when executed on a computer, a machine-readable storage medium on which the computer program is stored, and a computer configured to carry out the method. State of the art

[0002] The publication by authors Mark Schillinger et al: "Safe Active Learning and Safe Bayesian Optimization for Tuning a PI-Controller", IFAC-PAPERSONLINE, Vol. 50, No. 1, 1 July 2017 (2017-07-01), pages 5967-5972, DE ISSN: 2405-8963, DOI: 10.1016 / j.ifacol.2017.08.1258 discloses a method for safe learning using Bayesian optimization.

[0003] Safe exploration in active learning is known from "Safe Exploration for Active Learning with Gaussian Processes" by J. Schreiter, D. Nguyen-Tuong, M. Eberts, B. Bischoff, H. Markert, and M. Toussaint (ECML / PKDD, Vol. 9286, 2015). Specifically, point-based data is collected in a static state.

[0004] Active learning involves sequential data labeling to learn an unknown function. Data points are selected sequentially for labeling in such a way that the availability of information required to approximate the unknown function is maximized. The overall goal is to create an accurate model without providing more information than necessary. This makes the model more efficient, as potentially costly measurements can be avoided.

[0005] Active learning is widely used for data classification, e.g., for image labeling. For active learning in time series models representing physical systems, the data must be generated in such a way that relevant dynamic processes can be captured.

[0006] This means that the physical system must be excited by dynamically moving the input domain using input curves in such a way that the collected data—i.e., input and output curves—contain as much information about the dynamics as possible. Examples of input curves that can be used include sine, ramp, and step functions, as well as white noise. However, safety requirements must also be observed when exciting the physical systems. The excitation must not damage the physical system while the input domain is being dynamically explored.

[0007] Therefore, it is important to identify areas where dynamic stimulation can be carried out safely. Advantages of the invention

[0008] The method having the features of independent claim 1 has the advantage over the prior art that it combines dynamic exploration, active exploration and safe exploration.

[0009] Dynamic exploration refers to the acquisition of information under changing conditions of the system being measured. Active exploration aims to acquire information as quickly as possible, collecting information sequentially in such a way that a large amount of information can be gathered in a short period of time. In other words, the information gain from each measurement is maximized. Finally, safe exploration ensures that the system being measured is not damaged as much as possible.

[0010] With the method according to the invention, these three types of exploration can be combined.

[0011] The measures listed in the dependent claims enable advantageous further developments and improvements of the method specified in the independent claim. Disclosure of the invention

[0012] The present invention discloses an active learning environment with dynamic exploration (active learning) for time series models based on Gaussian processes, which takes into account the aspect of security by deriving an appropriate criterion for the dynamic exploration of the input domain.

[0013] Active learning is useful in a number of applications, such as simulations and forecasting. The general goal of learning processes is to create a model that describes reality. To do this, a real process, a real system, or a real object is measured, also referred to as the target, in the sense that information about the target is recorded. The created model of reality can then be used instead of the target in a simulation or forecasting. The advantage of this approach is that it saves time by avoiding having to repeat the process, which usually consumes resources, and by preventing the object or system from being exposed to the process being simulated, which could result in it being consumed, damaged, or altered.

[0014] It is advantageous if the model describes reality as accurately as possible. The present invention is particularly advantageous in that active learning can be used while taking safety conditions into account. These safety conditions are intended to ensure that the target to be detected is negatively / critically influenced as little as possible, e.g., in the sense that the object or system is damaged.

[0015] The invention uses a Gaussian process with a time series structure, e.g., with a nonlinear exogenous structure or a nonlinear autoregressive exogenous structure. Through dynamic exploration of the input domain, appropriate input and output curves, or output measurements, are generated for the time series model. The output measurements, i.e., the data labels, serve as information for the time series model. The input curve is parameterized into successive curve ranges, e.g., successive sections of ramp or step functions, which are determined stepwise using an exploratory approach, given safety requirements and previous observations.

[0016] The respective subsequent section is determined taking into account the previous observations in such a way that the information gain with regard to a criterion regarding the model is maximized.

[0017] A Gaussian process with nonlinear exogenous structures and a suitable exploration criterion is used as a time series model. At the same time, another Gaussian process model is used to predict safe input ranges with respect to the given safety requirements. The segments of the input curve are determined by solving an optimization problem with constraints to consider the safety prediction.

[0018] Exemplary applications of the invention are, for example, test benches for internal combustion engines, in which processes in the machines are to be simulated. Parameters to be recorded here include, for example, pressure values, exhaust gas values, consumption values, performance values, etc. Another application is, for example, the learning of dynamic models for robot controllers, in which a dynamic model is to be learned that maps joint positions to joint moments of the robot, which can then be used to control the robot. This model can be actively learned by exploring the joint area, but this should be done safely so that the movement limits of the joints are not exceeded, which could damage the robot. Another application is, for example, the learning of a dynamic model that serves as a replacement for a physical sensor.The data for learning this model can be actively generated and measured through exploration on the physical system. Safe exploration is essential, as a measurement in an unsafe region could damage the physical system. Another application is, for example, learning the behavior of a chemical reaction, where safety requirements can affect parameters such as temperature, pressure, acidity, or similar. Short description of the drawings

[0019] Embodiments of the invention are illustrated in the drawing and explained in more detail in the following description. It shows: Figure 1 the sequence 100 of the method for securely training a computer-aided model; Figure 2 the process 200 of the method for safely training a computer-aided model. Embodiments of the invention

[0020] The approximation to an unknown function f : X ⊂ ℝ d → Y ⊂ ℝ is to be achieved. In the case of time series models, such as the well-known nonlinear exogenous (NX) model, the input domain consists of discretized values, the so-called manipulated variables.

[0021] With xk for the time k applies: xk = ( uk , u k- 1 , ... , u k- d̃ + 1 ) , where u k k , u k ∈ Π ⊂ ℝ d ¯ represents the discretized control curve. d the dimension of the input domain Π of the system, d̃ the dimension of the NX structure and d = d · d̃ the dimension of X.

[0022] The elements UK are measured by the physical system and do not need to be equidistant. For simplicity, equidistance is assumed for the purposes of this notation. In general, the control curves are continuous signals and can be controlled explicitly.

[0023] In the model’s learning environment, data is collected in the form of n consecutive curve sections D n f = τ i ρ i i = 1 n , observed, with the input curve the is a matrix and consists of m input points of dimension d, ie τ i = x 1 i , … , x m i ∈ ℝ d × m . The output curve p i contains m corresponding output measurements, ie ρ i = y 1 i , … , y m i ∈ ℝ m .

[0024] The next curve section to be entered into the physical system as excitation t n +1 should now be determined in such a way that the information gain D n + 1 f regarding the modeling of f increased, but taking safety conditions into account.

[0025] To approximate the function f, a Gaussian process (abbreviated to GP) is used, which is defined by its mean function µ ( x ) and its covariance function k ( xi ,xj ), ie f ( xi ) ~ ( µ ( xi ), k ( xi , xj )) .

[0026] Assuming noisy observations of the input and output curves, the joint distribution according to the Gaussian process is given as p ( P n | Tn ) = ( P n |0, Kn + s 2< I ), where P n ϵ ℝ n ⋅ m is a vector connecting output curves, and T n ϵ ℝ n ⋅ m × d is a matrix containing input curves. The covariance matrix is ​​given by K n ϵ ℝ n ⋅ m × n ⋅ m For illustration, a Gaussian kernel is used as the covariance function, ie k x i x j = σ f 2 exp − 1 2 x i − x j T Λ f 2 x i − x j , which is θ f = σ f 2 Λ f 2 Furthermore, a zero vector 0 ∈ ℝ n ⋅ m as an average, a nm dimensional identity matrix I and s 2< is assumed as the output noise variance.

[0027] Given the joint distribution, the predicted distribution p ρ ∗ τ ∗ D n f for a new curve section t * be expressed as p ρ ∗ τ ∗ D n f = N ρ ∗ μ τ ∗ , ∑ τ ∗ , where μ τ ∗ = k τ ∗ T n T K n + σ 2 I − 1 P n , ∑ τ ∗ = k ∗ ∗ τ ∗ τ ∗ − k τ ∗ T n T K n + σ 2 I − 1 k τ ∗ T n , where k ∗ ∗ ϵ ℝ m × m a matrix with k ij ∗ ∗ = x i x j Furthermore, the matrix contains k ϵ ℝ m × n ⋅ m Key evaluations regarding t * the previous n Input curves. Since the covariance matrix is ​​completely filled, the input points x correlate completely with a curve segment as well as across different curves, using the correlations to plan the next curve. Since the matrix Kn + s 2< I possibly a high number of dimensions n m , inverting them can be time-consuming, so GP approximation techniques can be used.

[0028] The security state of the system is affected by an unknown function g described, with g : X ⊂ ℝ d → Z ⊂ ℝ that each input point x a security value z which serves as a security indicator. The values z are determined using information from the system and are set up in such a way that for all values ​​of z which are greater than or equal to zero, the corresponding input point x is considered safe.

[0029] Such safety values ​​z depend on the respective system and, as explained above, can represent system-dependent values ​​for safe or unsafe pressure values, exhaust gas values, consumption values, power values, joint position values, movement limits, sensor values, temperature values, acidity values ​​or similar.

[0030] The values ​​of z are generally continuous and indicate the distance of a given point x from the unknown safety boundary in the input range. Therefore, the given function gor an estimate thereof the safety level for a curve t A curve is classified as safe if the probability that its safety value z is greater than zero is sufficiently large, ie ∫ z1,...,z,≥0 p (z 1 , ... , zm | t ) dz 1 , ... , zm > 1 - α, where α ∈ [0,1] represents the threshold for t is uncertain. With the given data D n g = τ i ζ i i = 1 n , where ζ i = z 1 i , … , z m i ∈ ℝ m , a GP can be used to perform the function g to approximate. The forecast distribution p ζ * τ * , D n g for a given curve section t * is calculated as p ζ * τ * , D n g = N ζ * μ g τ * , ∑ g τ * , where µ g ( t *) and Σ g ( t* ) are the corresponding mean and covariance values. The quantities µ g and Σ g are calculated as shown in equations 2 and 3 respectively, but with Z n ∈ ℝ n ⋅ m as a target vector that contains all ζ i By using a GP to approximate g the safety condition x ( t ) for a curve t be calculated as follows: ξ τ = ∫ z 1 , … , z m ≥ 0 N ζ μ g τ , ∑ g τ dz 1 , … , z m > 1 − α . In general, the calculation of x ( t ) is difficult to solve analytically and therefore some approximation can be used, such as a Monte Carlo simulation or expectation propagation.

[0031] For the efficient selection of an optimal t The curve must be parameterized appropriately. One option is to perform the parameterization in the input area. The curve can be parameterized, for example, as ramp or step functions.

[0032] For a curve parameterization with a forecast distribution according to equation A and safety conditions according to equation E, the next curve section tn +1 ( or *) can be obtained by solving the following optimization problem with constraints: η * = argmax η ∈ ∏ J ∑ η so dass ξ η > 1 − α , where or ∈ Π the curve parameterization and represents an optimality criterion.

[0033] According to equation F, the predictive variance Σ from equation A is used for exploration. This is a covariance matrix determined by the optimality criterion is mapped to a real number, as shown in equation F. For Different optimality criteria can be used depending on the system. For example, the determinant, i.e., equivalent to maximizing the volume of the prediction confidence ellipsoid of the multinormal distribution, the trace, i.e., equivalent to maximizing the average prediction variance, or the maximum eigenvalue, i.e., equivalent to maximizing the largest axis of the prediction confidence ellipsoid. Other optimality criteria are also conceivable.

[0034] Referring to Figure 1 In step 120, an initialization is carried out by n 0 safe initial curves are executed. Regression and safety processes, Gaussian processes, are also created. The initial curves are located in a small safe area, in which the exploration begins. This small safe area is selected in advance based on prior knowledge of the system. The initial curves are determined to D n f , g = τ i ρ i ζ i i = 1 n with n = n 0 .

[0035] Subsequently, in step 160, a new curve section is calculated according to equations F and G t n +1 determined by or is optimized.

[0036] Subsequently, in step 170, the specific curve section t n +1 is used as input, and in this range r n +1 and ζ n +1 measured on the physical system.

[0037] Then, in step 150, the regression and security processes are updated. The regression model f is calculated according to equation A using D n f = τ i ρ i i = 1 n updated, and the security model g is calculated according to equation D using D n g = τ i ζ i i = 1 n updated.

[0038] Steps 150 to 170 are executed N times. In addition to a predetermined number of runs, automatic termination is also conceivable after a termination condition is reached. This could be based, for example, on training error (the error measure for model prediction and system response) or on additional potential information gain (if the optimality criterion becomes too small).

[0039] Subsequently, in step 190, the regression and security model are output.

[0040] Referring to Figure 2 , an implementation 200 of the method is explained. In step 210, a security threshold is set. α a value between 0 and 1 is selected. An initialization is then carried out in step 220 by n 0 safe initial curves, D 0 f , g = τ i ρ i ζ i i = 1 n with n = n 0 ,This also involves creating regression and safety processes, Gaussian processes. The initial curves are located in a small safe region, where the exploration begins. This small safe region is selected in advance based on prior knowledge of the system.

[0041] Subsequently, the part of the method comprising steps 240 to 280 is executed N times, where k is the run variable, ie indicates the current run. As in Figure 1 In addition to a predetermined number of runs, automatic termination upon reaching a termination condition is also conceivable. This could be based, for example, on training error (error measure on model prediction and system response) or on additional potential information gain (if the optimality criterion becomes too small).

[0042] First, in step 240, the regression model f is calculated according to equation A using D k − 1 f = τ i ρ i i = 1 n updated. In step 250, the security model g according to equation D using D k − 1 g = τ i ζ i i = 1 n updated. When performing steps 240 to 280 for the first time, steps 240 and 250 can be omitted.

[0043] Subsequently, in step 260, a new curve section is calculated according to equations F and G t n +1 determined by or is optimized.

[0044] Then, in step 270, the specific curve section t n +1 is used as input, and in this range r n +1 and g n +1 measured on the physical system.

[0045] Subsequently, in step 280, the input and output curves processed in the previous steps are D k − 1 f or D k − 1 g added.

[0046] After completing the repetitions of steps 240 to 280, step 290 follows, in which the regression and security models are updated and output.

[0047] The incremental updating of the GP models for new data, i.e., steps 150, 240, and 250, respectively, can be performed efficiently, e.g., by updating the rank of the matrix (rank-one update). Although an NX structure in combination with the GP model for time series modeling is shown here as an example, the general nonlinear auto-regressive exogenous case can also be used, i.e., GP with NARX input structure, where xk = ( yk , y k -1 , ... , yk - q , uk , uk - 1 , ... ., you are welcome ) . For optimization and planning purposes for the next curve section, the forecast mean value of p ρ τ , D n f as a replacement for y k The input excitation of the system is still controlled via the manipulated variable UK carried out.

Claims

1. Computer-implemented method (200) for safe, dynamic exploration of an input value range of a physical system for training a time series-based model of the physical system, comprising the steps of: defining (210) a safety threshold value, a; initializing (220) by implementing safe initial curves as input values on the physical system, creating an initial regression model and an initial safety model, wherein the regression model and the safety model are each Gaussian processes, wherein the safety model is designed to output a safety value, z, which characterizes a distance between a given curve point, x, and an unknown safety limit and indicates whether the given curve point, x, is safe or unsafe in regard to damage to the physical system; repeatedly carrying out the steps of updating (240) the regression model; updating (250) the safety model; determining (260) a new curve section, T, which comprises a plurality of curve points, x1, ..., xm, as input values of the physical system, wherein the new curve section, T, is determined by optimization taking into account a safety condition, wherein the safety condition is that, on the basis of the safety values, z1, ..., zm, which were ascertained by means of the safety model for the plurality of curve points of the curve section, the curve section, T, is classified as safe if the probability ∫z1,...,zm≥0p(z1, ... ,zm|τ)dz1, ... ,zm is greater than 1 - a, wherein the safety threshold value, a, indicates a threshold value in respect of the new curve section being unsafe; implementing (270) the determined new curve section on the physical system and measuring output values; including (280) the output values in the regression model and in the safety model until a predefinable number N of passes have been carried out; and updating and outputting (290) the regression model and the safety model.

2. Method according to Claim 1, wherein the new curve section is determined (260) in such a way that the information gain is maximized while satisfying the safety criteria of the safety model.

3. Method according to either of Claims 1 and 2, wherein the new curve section is determined (260) using a covariance matrix.

4. Method according to any of Claims 1 to 3, wherein the system is a test bed for internal combustion engines, a robot controller, a physical sensor or a chemical reaction.

5. Method according to any of Claims 1 to 4, wherein the safety model comprises safety values of the system, specifically at least one out of pressure values, exhaust gas values, consumption values, performance values, joint position values, movement limits, sensor values, temperature values or acidity values.

6. Computer program comprising instructions which, when the program is executed by a computer, cause the latter to carry out the method according to any of Claims 1 to 5.

7. Machine-readable storage medium on which the computer program according to Claim 6 is stored.

8. Device comprising means for carrying out the method according to any of Claims 1 to 5.