Method of and system for training target machine learning model for target system

By using a multi-task Gaussian process to jointly model safety values of a target and auxiliary system, the method addresses the limitations of local exploration in existing safe learning methods, achieving efficient and global safety region exploration.

JP2025078059APending Publication Date: 2025-05-19ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024192892
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-11-01
Publication Date
2025-05-19

AI Technical Summary

Technical Problem

Existing safe learning methods rely on accurate safety modeling and tend to perform local exploration, missing safe but disconnected regions and requiring labor-intensive domain expert input for data collection.

Method used

The method employs a multi-task Gaussian process to jointly model safety values of a target system and an auxiliary system, leveraging transferable auxiliary knowledge to accelerate learning and enable global exploration of safety regions.

Benefits of technology

This approach reduces data consumption, allows for global exploration of disparate safety regions, and improves the efficiency and safety of the learning process, especially in early iterations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025078059000001_ABST
    Figure 2025078059000001_ABST
Patent Text Reader

Abstract

To provide a sequential learning method for training a target machine learning model for use in a target system in a process and a machine to be engineered.SOLUTION: The present invention is directed to a system 100 for training a target machine learning model for use in a target system in a process and machine to be subjected to engineering. A multi-task Gaussian process is mounted with a joint model of a target safety value 332 of a target system 330 and an auxiliary safety value 312 of an auxiliary system 310. A new state is selected for the target system 330, and at that time the target safety value 332 is predicted by the multi-task Gaussian process. Use of the safety value acquired from the auxiliary system allows the system to more satisfactorily model the safety value regarding the target system.SELECTED DRAWING: Figure 3a
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The subject matter disclosed herein relates to a method, a system, and a computer-readable medium for training a target machine learning model for a target system in an engineered process and machine.

Background Art

[0002] Background Despite the remarkable success of machine learning, data access remains a complex task. One notable approach involves the design of experiments. Specifically, active learning (AL) and Bayesian optimization (BO) employ sequential data selection procedures. These methods start with a limited data set, repeatedly calculate an acquisition function, select new data based on the acquisition score, obtain observations from an oracle, and update their beliefs. This process continues until the learning objective is satisfied or the acquisition budget is exhausted. In many cases, these learning algorithms use a Gaussian process as an alternative model for calculating the acquisition function.

[0003] Safety in exploration is important in various different fields. For example, in medical simulation devices, particularly implantable devices such as spinal cord stimulation devices, safety is a concern. This is discussed in "Effect of epidural stimulation of the lumbosacral spinal cord on voluntary movement, standing, and assisted stepping after motor complete paraplegia: a case study." by Harkema et al., which is incorporated herein by reference.

[0004] Safety is also important in robotics. Studies such as “Safe controller optimization for quadrotors with Gaussian processes” by Berkenkamp et al. and “GoSafe: Globally Optimal Safe Robot Learning” by Baumann et al., both incorporated herein by reference, address this problem.

[0005] One approach to implementing safe learning involves using an additional Gaussian process (GP) to model safety constraints. This process starts with a set of pre-determined safe observations. A safe set is defined to limit exploration to regions that exhibit a high level of safety confidence. As learning progresses, the safe set expands, thereby increasing the area available for exploration.

[0006] Safe learning algorithms rely on an accurate safety model. This approach requires a well-calibrated model of the safety values that should be available before exploration begins, which is often difficult. Another limitation of safe learning algorithms is that they tend to be oriented towards local exploration. Gaussian processes are inherently smooth, and uncertainty increases as they move beyond the boundaries of the currently identified safe set. As a result, regions that are actually safe but disconnected from the current safe set are misclassified as unsafe and remain unexplored. Since domain experts are required to supply safe data from multiple safe regions, the deployment of safe learning algorithms is more labor-intensive.

Prior Art Documents

Non-Patent Documents

[0007]

Non-Patent Document 1

[0008] Summary Sequential learning methods such as active learning and Bayesian optimization select the most beneficial data for learning about a task. In many medical or engineering applications, data selection is constrained by safety conditions that are a priori unknown. Safe learning methods use Gaussian processes (GPs) to model safety probabilities and perform data selection within regions with high safety confidence. However, accurate safety modeling requires prior knowledge or consumes data. Furthermore, safety confidence is concentrated around given observations, which leads to local exploration. [Means for Solving the Problems]

[0009] The inventors recognized that in experiments where safety is important, transferable auxiliary knowledge can often be utilized. In one embodiment, safe sequential transfer learning is used to accelerate the learning of safety values.

[0010] Some embodiments relate to a method for training a target machine learning model for a target system in engineered processes and machines. A multi-task Gaussian process implements a joint model of the safety values of the target system and the safety values of an auxiliary system. A new state is selected for the target system, and the target safety value is predicted by the multi-task Gaussian process.

[0011] By using the safety values obtained from the auxiliary system, the system can better model the safety values regarding the target system. This approach has been empirically demonstrated to learn tasks with less data consumption and globally explore multiple disparate safety regions under the guidance of auxiliary knowledge.

[0012] In one embodiment, only the part of the joint model related to the target system is updated when new data becomes available therefor. By pre-computing the auxiliary components, the additional computational load introduced by incorporating the auxiliary data is reduced.

[0013] The training method described herein may be applicable in a wide range of practical applications. Some such practical applications include an engine and a vehicle or robotic device configured for at least partial autonomous movement. Numerous other applications are described herein.

[0014] One embodiment of the method may be implemented on a computer as a computer-implemented method, or may be implemented in dedicated hardware, or in a combination of both. Executable code for one embodiment of the method may be stored in a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product includes non-transitory program code stored on a computer-readable medium for implementing one embodiment of the method when the aforementioned program product is executed on a computer.

[0015] In one embodiment, the computer program includes computer program code configured to perform all or part of the steps of one embodiment of the method when the computer program is executed on a computer. Advantageously, the computer program is embodied on a computer-readable medium.

[0016] Another aspect of the subject matter disclosed herein is a method of making a computer program available for download. This aspect is used when the computer program is uploaded and when the computer program is available for download.

[0017] Further details, aspects, and embodiments will be described by way of example only with reference to the drawings. The elements in the drawings are illustrated for purposes of simplicity and clarity and are not necessarily drawn to scale. In the drawings, elements corresponding to those already described may have the same reference numerals.

Brief Description of the Drawings

[0018]

Fig. 1a

Fig. 1b

Fig. 2a

Fig. 2b

Fig. 3a

Fig. 3b

Fig. 3c

Fig. 3d

Fig. 3e

Fig. 3f

Fig. 4a

Fig. 4b

Fig. 4c

Fig. 4d

Fig. 4e

Fig. 4f

Fig. 5a

Fig. 5b

Fig. 5c

Fig. 5d

Fig. 5e

Fig. 5f

Fig. 6

Fig. 7a

Fig. 7b

Best Mode for Carrying Out the Invention

[0019] Description of Embodiments The subject matter disclosed herein may be embodied in many different forms, but this disclosure should be considered as illustrative of the principles of the subject matter disclosed herein and is not intended to be limited to the specific embodiments shown and described in the drawings and detailed in this specification, under the understanding that one or more specific embodiments are shown in the drawings and are detailed herein.

[0020] In the following, for the sake of understanding, the elements of the operating embodiments will be described. However, it will be apparent that each element is configured to perform the functions described as being performed by each element.

[0021] Furthermore, the subject matter disclosed herein is not limited only to the embodiments, but also includes all other combinations of the features described herein or the features described in a plurality of mutually different dependent claims.

[0022] FIG. 1a schematically shows an example of an embodiment of conventional safe sequential learning.

[0023] The first image shows a schematic representation of the input domain. There are two safe regions in the input domain, and one of these two safe regions is labeled with reference numeral 101. A plurality of initial points that are known to be safe for the objective function are shown. The initial points or initial target states are schematically shown as small circles 111.

[0024] To learn the objective function in the input domain, the system can sequentially attempt new measurements, which are shown in a second schematic image. The selected target state, selected for the new measurement, is indicated by a small square 112.

[0025] Ultimately, a good representation of the objective function is learned, but only within the safe region 101. The other safe region is never discovered.

[0026] FIG. 1b schematically shows an example of an embodiment of safe sequential transfer learning according to an embodiment.

[0027] As in the case of FIG. 1a, the first image shows a schematic representation of the same input domain. There are two safe regions in the input domain. As in the case of FIG. 1a, a plurality of initial points that are known to be safe for the objective function are shown. The initial points or initial target states are schematically shown as small circles 111.

[0028] What is different from FIG. 1a is that a plurality of auxiliary points are shown such that the safety values are known in the auxiliary system. The auxiliary states are shown as small triangles 113. In this example, these points are extracted only from the safe region regarding the auxiliary system, but in one embodiment, these points may include points from both the safe region and the non-safe region regarding the auxiliary system. Typically, the safe points regarding the auxiliary system are not necessarily the safe points regarding the target system. In fact, some of the known auxiliary points are outside the range of the safe region as shown in the figure. However, the safety value regarding the auxiliary system and the safety value regarding the target system are related to some extent, for example, correlated. Therefore, it is a fact that the knowledge of the auxiliary safety values regarding a plurality of points is beneficial for the safety value of the target system. The knowledge of the safe region regarding the auxiliary system enables transfer learning of the safe region.

[0029] The second figure of FIG. 1b shows that the system can sequentially attempt new measurements to learn the objective function in the input domain, and these new measurements are shown in the second schematic image. The selected target state, selected for the new measurement, is indicated by a small square 112.

[0030] The exploration learns new values of the objective function, but also learns safety values regarding the points. With the knowledge of the safety values regarding the auxiliary system, it becomes possible to learn the safety values regarding the target system significantly faster. As a result, in Figure 2, as shown by the fact that the explored points spread more rapidly, the exploration becomes more aggressive and, in particular, it becomes possible to jump over the gap between the two safety regions. Finally, in Figure 3, both safety regions are explored despite being disjoint.

[0031] Finally, in both safety regions, good representations of the objective function are learned.

[0032] Since safe learning is always initialized using some prior knowledge, it is reasonable to assume that multiple correlated experiments have been conducted and that their results can be utilized as part of the prior knowledge. The assumption that auxiliary data can be utilized is usually satisfied in real-world experiments. Specific applications exist everywhere, including simulations of the real world, continuous production, and multi-fidelity modeling.

[0033] There are two advantages. That is, 1) the exploration and expansion of the safety regions are significantly accelerated, and 2) the auxiliary task can provide guidance regarding the safety regions that are disconnected from the initial target data, thereby assisting global exploration. Furthermore, it is observed that the queries are safer than the conventional approach, especially in the early iterations. Refer to the tables and the drawings of the false positives of the experiments described while referring to Figures 4a to 5f.

[0034] From a modeling perspective, transfer learning can be achieved by considering both the auxiliary task and the target task using a multi-output GP model.

[0035] Conventional calculations of GP have a three-dimensional time complexity due to the inversion of the Gram matrix. Therefore, introducing potentially large amounts of auxiliary data will result in a significant computational load, and in real experiments, the calculation time often becomes a bottleneck. In one embodiment, the multi-output GP is modularized so that components related to auxiliary calculations can be pre-computed and fixed. This reduces the computational complexity of the multi-output GP while maintaining its advantages. Since the goal is to learn only with respect to the target task of the present invention, this is not a problem.

[0036] FIG. 2a schematically shows an example of an embodiment of a system 200 for training a target machine learning model for a target system.

[0037] A target system 210, a training device 220, and a controller device 230 are shown, and these may be considered as part of the system 200.

[0038] The target system 210 is in engineered processes and machines. For example, the target system 210 may be an engine, a vehicle, or a robotic device, particularly a vehicle or robotic device configured for at least partial autonomous movement. The target system 210 enables configuration to one of a plurality of possible states. A state may include one or more input parameters that control the target system. A state may include one or more sensor values that at least partially define the state of the target system.

[0039] During operation, the target system enables measurement of at least one target safety value and physical quantities of the target system. The operation of the target system is defined as safe when at least one target safety value is within the safety region. The safety region may be set by an expert user and / or may be determined experimentally. The physical quantities are useful for monitoring or controlling the target system. There may be an overlap between the safety value and the physical quantity, or rather, the safety value and the physical quantity may be the same.

[0040] For example, in one embodiment, the target system 210 includes an engine. The state of the target system 210 may include various different parameters, for example, the following list: throttle position, fuel injection rate, ignition timing, intake air temperature, gear ratio, engine load, engine temperature, exhaust back pressure, variable valve timing, oxygen sensor feedback, air-fuel ratio, exhaust gas recirculation (EGR), idling speed, oil pressure, and may include one or more of them. Some of these parameters, for example, the throttle position, may be controlled by the user, or, for example, the fuel injection rate may be controlled by the controller. Some of these parameters are measurable but not directly changeable.

[0041] For example, the safety value regarding the engine may be the engine temperature, for example, preferably a temperature below the maximum temperature, for example, an engine temperature of 110 °C, or the oil pressure, for example, an oil pressure above the minimum value.

[0042] For safety considerations, soft safety constraints and hard safety constraints can be distinguished. Input states that generate outputs violating soft constraints may potentially compromise the system, but the consequences are not severe. Therefore, it is possible to continue measurements using the same machine. However, if hard constraints are violated, the engine becomes inoperable. This distinction can typically be ignored, but for example, soft constraints can be specified as safety constraints for training embodiments if necessary.

[0043] The physical quantities that can be measured with respect to system 210 may include some or all of the above-listed items, and may also include emission values such as carbon dioxide (CO2), carbon monoxide (CO), nitrogen oxides (NOx), particulate matter (PM), volatile organic compounds (VOC), sulfur oxides (SOx), hydrocarbons (HC), ammonia (NH3), formaldehyde (HCHO), phenol and aldehydes, lead, and heavy metals.

[0044] The physical quantity may also include quantities other than emissions, such as engine roughness. Examples of engine roughness include irregular combustion, which refers to non-uniformity in the combustion process within the cylinder that causes uneven power supply. This irregular combustion can be affected by factors such as fuel quality, air-fuel mixture, and ignition timing. Another factor is mechanical friction, in which case roughness can refer to increased resistance in the movement of the piston, or friction experienced by components moving within the engine, such as non-uniform wear and tear in the bearings. Vibration can also be an indicator, in which case "roughness" represents excessive vibration or noise generated from the engine due to imbalance in the rotating parts or misalignment of components. Finally, operational instability such as RPM (revolutions per minute) fluctuations can also be expressed by the term roughness and can be quantitatively measured using various different measurement criteria such as the root mean square of acceleration or a specialized roughness index.

[0045] The target system 210 is configured to enable measurement of at least one target safety value of the target system and a physical quantity useful for monitoring or controlling the target system. Further, the state in which the system 210 can be configured can be recorded together with the measured physical quantity and / or safety value.

[0046] The training device 220 is configured to train a target machine learning model to predict one or more physical quantities from the state in which the system 210 is configured. The state can include, for example, the settings of the system 210, input parameters such as throttle position, and / or measured physical parameters, and such parameters may be outside the user's control, such as ambient temperature, but can still affect the physical quantity. Note that the model is trained to predict such physical quantities. When selecting a new state, it can be assumed that parts of the state that cannot be changed, such as ambient temperature, are fixed while the remaining part of the state is being selected.

[0047] The training device 220 has access to auxiliary training data 224 corresponding to the auxiliary system. The auxiliary training data includes a plurality of pairs of auxiliary states and auxiliary safety values. The training device 220 uses transfer learning to train a multi-task Gaussian process as a joint model of the safety value of the target system and the safety value of the auxiliary system. The multi-task Gaussian process receives, as input, a state that can be either the state of the auxiliary system or the state of the target system, and generates a prediction regarding the target safety value or a prediction regarding the auxiliary safety value, respectively. The auxiliary training data includes, for the auxiliary system, typically a larger number of known state / safety value pairs than are available for the target system, and / or explores a wider range of regions in the domain input than the pairs available for the target, so that the learning of the joint model is faster and / or explores a wider range of regions of the domain input, e.g., a larger number of safe regions, than training a Gaussian process based only on state / safety value pairs regarding only the target system.

[0048] Using the trained joint model, new states for the system 210 to attempt can be iteratively selected, and preferably measurements regarding both a physical quantity and a safety value can be obtained. It should be noted that the system does not necessarily always measure both for each state, but it is desirable to do so.

[0049] Once the target machine learning model is trained, this target machine learning model can be uploaded to the controller 230. For example, the controller 230 can be used with the system 210 to control the system 210. For example, if the target machine learning model predicts that a physical quantity is outside a desired range, e.g., above a predetermined value or too low, an input parameter, e.g., a gas supply amount, etc., can be changed so that the predicted physical quantity approaches the desired range, e.g., the state can be changed. One embodiment may include both the controller 230 and the system 210, or may include only the controller 230.

[0050] For example, the controller 230 can be configured to better control the physical quantities of the target system 210 using the system 200. For example, the state of the system 210 can be changed such that, for example, the emission value of the engine is maintained below the target or returned below the target. For example, the state of the system 210 can be changed such that, for example, the orientation of the autonomous vehicle changes, for example, turns. The physical quantity and the safety quantity may be considered the same. For example, a model can be trained to predict the engine temperature, but during learning, the engine temperature is restricted from exceeding a specific value. In use, the model can be used to ensure that the predicted engine temperature is maintained within a safe boundary. If the predicted temperature exceeds a predetermined value, the state can be changed, for example, the vehicle speed can be reduced, for example, the gas supply can be reduced. This action can also include emergency actions, such as emergency braking, etc. This may be to preserve the integrity of the engine, but may also be a reaction to the traffic situation.

[0051] The target system 210 may include a processor system 211, a storage 212, and a communication interface 213. The training device 220 may include a processor system 221, a storage 222, and a communication interface 223. The controller device 230 may include a processor system 231, a storage 232, and a communication interface 233.

[0052] In various different embodiments of the communication interfaces 213, 223, and / or 233, the communication interface can be selected from various different options. For example, the interface can be a network interface to a local area network or a wide area network, such as the Internet, a storage interface to internal or external data storage, an application programming interface (API), etc.

[0053] Storages 212, 222, and 232 may be, for example, electronic storage, magnetic storage, etc. The storage may include local storage, such as a local hard drive or electronic memory. Storages 212, 222, and 232 may include non-local storage, such as cloud storage. In the latter case, storages 212, 222, and 232 may include a storage interface to the non-local storage. The storage may include a plurality of separate sub-storages that together constitute storages 212, 222, 232.

[0054] Storages 212, 222, and 232 may be non-transitory storage. For example, storages 212, 222, and 232 can store data in the presence of power, like a volatile memory device, such as random access memory (RAM). For example, storages 212, 222, and 232 can store data both in the presence and absence of power, like a non-volatile memory device, such as flash memory. The storage may include a volatile writable portion, such as RAM, and may include a non-volatile writable portion, such as flash. The storage may include a non-volatile non-writable portion, such as ROM. The training device 220 may be able to access a collection of auxiliary data 224 that may include an auxiliary state and a safety value, which may be in the form of a database or other ordered data storage.

[0055] Devices 210, 220, and 230 can communicate with each other internally via a computer network, with other devices, external storage, input devices, output devices, and / or one or more sensors. The computer network may be, for example, the Internet, an intranet, a LAN, a WLAN, a WAN, etc. The computer network may be the Internet. Devices 210, 220, and 230 may include a connection interface configured to communicate with the inside or outside of system 200 as needed. For example, the connection interface may include a connector, such as a wired connector, such as an Ethernet connector, an optical connector, etc., or may include a wireless connector, such as an antenna, such as a Wi-Fi antenna, a 4G antenna, or a 5G antenna.

[0056] Using communication interface 213, digital data can be transmitted or received, such as a state including a measured part of the state, an instruction for changing the state, a measured physical quantity, and / or a safety value. Using communication interface 223, this digital data can be transmitted or received. Using communication interface 223, controller 230 can be configured. Using communication interface 233, digital data can be transmitted or received, for example, a model can be received from device 220. Using communication interface 233, system 210 can be controlled.

[0057] The target system 210, the training device 220, and the controller device 230 may have a user interface that may include well-known elements such as one or more buttons, a keyboard, a display, a touch screen, etc.

[0058] The execution of apparatuses 210, 220, and 230 may be implemented in a processor system. Apparatuses 210, 220, and 230 may include functional units for implementing aspects of each embodiment. The functional units may be part of a processor system. For example, the functional units shown herein may be implemented, in whole or in part, in computer instructions stored in the storage of the apparatus and executable by the processor system.

[0059] The processor system may include one or more processor circuits, such as, for example, a microprocessor, a CPU, a GPU, etc. Apparatuses 210, 220, and 230 may include a plurality of processors. The processor circuit may be implemented, for example, in a distributed manner as a plurality of sub-processor circuits. For example, apparatuses 210, 220, and 230 may be able to use cloud computing.

[0060] Typically, the target system 210, the training apparatus 220, and the controller apparatus 230 each include a microprocessor that executes appropriate software stored in the apparatus. For example, this software may be downloaded and / or stored in a corresponding memory, such as volatile memory like RAM or non-volatile memory like flash.

[0061] Instead of implementing functions using software, apparatuses 210, 220, and / or 230 may be implemented, in programmable logic, wholly or partly, for example, as a field programmable gate array (FPGA). The apparatuses may be implemented, wholly or partly, as so-called application specific integrated circuits (ASICs), for example, as integrated circuits (ICs) customized for their respective specific applications. For example, the circuits may be implemented in CMOS using a hardware description language such as Verilog, VHDL, etc. In particular, the target system 210, the training apparatus 220, and the controller apparatus 230 may include circuits for, for example, cryptographic processing and / or arithmetic processing.

[0062] In a hybrid embodiment, the functional units are implemented partly, for example, as a coprocessor, for example, as a GPU coprocessor, in hardware, and partly in software stored and executed on the apparatus.

[0063] FIG. 2b schematically shows an example of an embodiment of the system 202. The system 202 may include one or more of the target system 210, the training apparatus 220, and the controller 230. The system and apparatuses are connected via a computer network 272, for example, via the Internet. The target system 210 and the training apparatus 220 may conform to an embodiment.

[0064] FIG. 3a schematically shows an example of an embodiment of the system 100 for training a target machine learning model for the target system 330.

[0065] Figure 3a shows a target system 330. The target system 330 can be configured to one of a plurality of possible states. As schematically shown, the target system 330 is configured to state 341. The target system 330 can be provided with sensors for measuring various different aspects of the system. In particular, the target system 330 can be provided with a physical quantity sensor 333 for measuring a physical quantity 334. The physical quantity is useful for monitoring or controlling the target system 330.

[0066] The target system 330 can also be provided with a target safety value sensor 331 for measuring a target safety value 332. When the safety value is within a safety region that has been determined in advance, for example, experimentally or by an expert user or the like, the target system 330 is defined as being safe. It is highly desirable to prevent the target system 330 from operating in a state 341 that causes it to enter an unsafe state, that is, a state 341 that causes the safety value of the target system 330 to take a value outside the safety region.

[0067] Typically, there are two or more safety values, and the values of these safety values together define a safe state. The safety region can include a plurality of non-connected sub-regions within the state space.

[0068] In one embodiment, the safety value includes one or more measurement criteria selected from the group consisting of temperature, pressure, vibration level, noise level, humidity, current level or voltage level, flow rate, force, torque, speed, RPM, concentration of a chemical substance, and state of charge.

[0069] In one embodiment, the physical quantity can include one or more of the aforementioned safety values. The physical quantity can also refer to the environment of the target system, for example, the position of an obstacle near the target machine.

[0070] Figure 3a shows an auxiliary system 310. The auxiliary system 310 can also be configured to one of a plurality of possible states. The possible states of the system 310 may be the same as the possible states regarding the target system 330, but this is not essential. As schematically shown, the auxiliary system 310 is configured to state 321. The auxiliary system 310 can also be provided with sensors for measuring various different aspects of the system. The auxiliary system 310 can also be provided with an auxiliary safety value sensor 311 for measuring an auxiliary safety value 312.

[0071] The auxiliary system 310 may have one or more auxiliary safety values. The number of safety values is typically the same for systems 310 and 330, but this is not essential.

[0072] The auxiliary safety value can define that the system 310 is safe. Usually, it is not necessary to consider the safe state of the system 310, and only the safety value itself needs to be considered. However, even when new measurements are still made regarding the auxiliary system, it may be desirable to consider the safe region of the auxiliary system in order to avoid causing an unsafe state in the auxiliary system.

[0073] Preferably, the auxiliary system 310 is also provided with a physical quantity sensor 313 for measuring a physical quantity 314. This is not essential, but as will be further described with reference to Figure 3f, it can accelerate the training of a machine learning model regarding the physical state of the target system 330 by transfer learning.

[0074] Examples of safety values are described in this specification. Typical examples include temperature, pressure, etc. indicating an unsafe operating state of the system.

[0075] The relationship between the safety value of the auxiliary system and the corresponding auxiliary state is assumed to be beneficial for the relationship between the safety value of the target system and the corresponding target state. For example, the auxiliary relationship can be similar to the relationship between the safety value of the target system and the corresponding target state. The characteristics of the auxiliary can inform the characteristics of the target.

[0076] For example, in one embodiment, the auxiliary system is a computer simulation of the target system. In this case, the safety value is obtained using simulated sensors rather than real sensors. This is an advantageous embodiment because computer simulation is a good method for obtaining safety values without the risk to the actual target system. The simulated values may not be completely accurate, but still, the simulated values can be used to provide information for training the joint model as described later. A further advantage is that simulated safety values can also be obtained from non-safe regions.

[0077] In one embodiment, the target system and the auxiliary system are both the same type of engineered process or machine, whereby measurements from the auxiliary system are predictive of the target system. For example, the target system and the auxiliary system may well be engines such as internal combustion engines, in which case both systems can represent variations of these engines in, for example, separate vehicles. The physical quantity can include parameters such as fuel efficiency or emissions, in the case of a wind turbine system, the physical quantity can include measurement criteria such as blade stress or rotational speed, there are also various parameters in the case of a chemical reactor, in the case of a CNC machine, the physical quantity can include tool wear or machine accuracy, and in the case of a 3D printer, the physical quantity can include print quality or material usage. Many other examples of systems and physical parameters are possible.

[0078] FIG. 3b schematically shows an example of an embodiment of training data for a multi-task Gaussian process that implements a joint model of the safety value of the target system and the safety value of the auxiliary system. A joint model 361 (see FIG. 3c) is trained to transfer the information contained in the data captured regarding the auxiliary system, and this joint model 361 is applied to both the auxiliary system and the target system. Accordingly, training data 350 including auxiliary training data 351 and target training data 352 is collected.

[0079] The auxiliary training data 351 includes a plurality of pairs of an auxiliary state 321 and an auxiliary safety value 312. The plurality of states are typically different from each other, typically at least partially different, but some degree of replication of the states can be tolerated, and the corresponding safety values 312 resulting from those states can also be tolerated.

[0080] The target training data 352 includes a plurality of pairs of a target state 341 and a target safety value 332. The plurality of states are typically different from each other, typically at least partially different, but some degree of replication of the states can be tolerated, and the corresponding safety values 332 resulting from those states can also be tolerated.

[0081] Typically, the auxiliary training data 351 is fixed. For example, the auxiliary training data 351 may be a collection of data obtained from other systems, for example, an older version, i.e., one that is now complete and no longer available, or may be a computer simulation where it is not even possible to obtain new data. On the other hand, when the system 310 is in operation, this is not essential, and new data including new safety values and, optionally, physical quantities can be continuously obtained, in which case these new data can be added to the auxiliary training data.

[0082] On the other hand, the target training data 352 is expected to expand. As the joint model 361 is improved, the predictions of this joint model 361 can be used to safely select new input states, and in turn, new safety values and physical quantities can be measured for these new input states. For example, as shown in FIG. 1b, at the start of training, there are a large number of data points known for the auxiliary system, but only a few starting points are known for the target system. However, as training progresses, the target training data expands.

[0083] FIG. 3c schematically shows an example of an embodiment of a system for training a target machine learning model for a target system. The joint model 361 is shown in FIG. 3c.

[0084] In one embodiment, the joint model 361 is trained using transfer learning.

[0085] There are various different types of transfer learning. For example, the model can first be trained with the auxiliary training data 351 and then fine-tuned with the target auxiliary training data 352.

[0086] However, in one embodiment, other different approaches are used. In one embodiment, the joint model receives as input a state regarding either the target system or the auxiliary system (or one of them) and an indicator (which may be a single bit in the case of a single auxiliary system) indicating which system this state belongs to, and generates a prediction of one or more safety values regarding the corresponding system. Thus, the joint model retains the ability to predict safety values regarding both the target system and the auxiliary system, rather than being retrained for a new but related task.

[0087] It was discovered in experiments that an advantageous option for the joint model 361 is a multi-task Gaussian process that implements a joint model of the safety value of the target system and the safety value of the auxiliary system. The multi-task Gaussian process receives, as input, a state regarding either the target system or the auxiliary system (or one of these) and an indicator (which may be a single bit in the case of a single auxiliary system) indicating which system this state belongs to, and generates a prediction of one or more safety values regarding the corresponding system.

[0088] To explore the relationship between the state input regarding the target system and the physical quantity and / or safety value, a new state is generated, and then this new state can be executed on the target system 330. In the process, the physical quantity and / or safety value is obtained. Typically, both the physical quantity and the safety value are obtained, but even if only one of them is obtained, it is possible to handle this situation.

[0089] Using the new physical quantity obtained for the new state, the target machine learning model 381 (see Figure 3e) can be further trained to predict the physical quantity. Using the new safety value obtained for the new state, the joint model 361 can be further trained to predict the safety value regarding the target system.

[0090] To select a new state 342 for the target system, system 10 may include a state selection unit 362. The state selection unit, on the one hand, may generate new states that would be useful, for example, to expand knowledge about the physical quantities of the target system for training a model 381, and / or new states that would be useful, for example, to expand knowledge about the safety values of the target system for training a joint model 361, provided that such new states are generated under the condition that they do not cause the target system 330 to enter an unsafe situation. That is, the new state should not cause the target system 330 to take a safety value outside the safety region. For example, the state selection unit may select candidates for new states and use the joint model to obtain a prediction of the corresponding safety value for the target system. If the prediction is outside the range of the safety region, the candidate for the new state is rejected. The joint model 361 can also generate a probability distribution of the values of the safety value, for example, typically a Gaussian distribution. The state selection unit 362 can use the probability distribution to calculate the probability that the safety value is outside the safety region. An acceptable probability outside the safety region can be defined in the selection unit 362.

[0091] For example, the selection unit 362 can be provided with an active learning algorithm. The state selection unit selects new states to obtain useful data for training a machine learning model and / or a joint model. For this purpose, various different criteria can be used, such as uncertainty sampling that focuses the model on the region with the highest uncertainty, or a query-by-committee where multiple versions of the model jointly determine the next most useful state to query. The selection unit 362 can focus only on the training of the machine learning model, for example, only on the learning of physical quantities, but in one embodiment, the state selection unit balances the useful values of the new states for both the joint model 361 and the machine learning model 381.

[0092] For example, the selection algorithm can balance exploration and exploitation within the data space. For example, in order to improve prediction accuracy, a compromise can be made between exploring new regions of the data space and exploiting known regions. For example, a sampling distribution can be assigned to a pool of candidate new states. The sampling distribution can represent a value that includes a value indicating a term for search and a value indicating a term for exploitation. One example is the Active Thompson Sampling (ATS) algorithm. For example, an acquisition function is defined for candidate new states, and this acquisition function is optimized to select new states. The acquisition function can evaluate the benefit or usefulness of a new state. For example, the acquisition function can quantify the uncertainty or potential benefit of a given new state.

[0093] The selection unit 362 does not necessarily have to focus equally on the entire domain input. For example, the selection unit 362 can be configured for Bayesian optimization, in which a physical state is sought that optimizes several criteria.

[0094] In both methods, candidate new states are evaluated based on the predicted safety value, and if the candidate new states are likely to cause an unsafe state in the target system 330, for example, if they exceed a threshold, these candidate new states are rejected.

[0095] When the selection unit 362 selects a new state 342, the target system 330 can be configured for that new state 342, new physical states and / or safety values can be obtained, and the machine learning model 381 and the joint model 361 can be updated.

[0096] Interestingly, in one embodiment, the auxiliary data portion of the training data 350 is constant while the target data portion is expanding. This can be utilized to more efficiently update the joint model 361. A portion of the joint model 361 that is only related to the auxiliary system and / or the auxiliary data 351 can be identified. Since the auxiliary data 351 does not change, there is no need to update this portion related to the auxiliary system 310 in order to update the joint model 361. This is particularly efficient when the joint model 361 includes a multi-task Gaussian process because the update of such a process can increase three-dimensionally with the number of data points. By reducing the amount of data that needs to be updated, the training process becomes more efficient.

[0097] FIG. 3d schematically shows an example of one embodiment of the training data 370 for the target machine learning model 381.

[0098] The target training data 370 includes a plurality of pairs of a target state 341 and a physical quantity 334. The plurality of states are typically different from each other and typically at least partially different, but some degree of replication of the states can be tolerated, and the corresponding physical quantities 334 generated by those states can also be tolerated.

[0099] FIG. 3e schematically shows an example of one embodiment of the target machine learning model 381.

[0100] For the target machine learning model 381, various different machine learning models can be employed. The model 381 is configured to receive the state of the target system as an input and generate a prediction of the physical quantity as an output.

[0101] For example, the target machine learning model 381 may include a neural network trained to process the state of the target system and generate predictions of physical quantities. Alternatively, the model 381 can employ a support vector machine (SVM) that can operate by finding one or more hyperplanes to separate multiple different predicted results based on the state of the system. Another option may be a random forest algorithm that uses an ensemble of decision trees to make predictions. Each tree considers a random subset of features, thereby providing a robust model. It is also possible to use gradient boosting techniques that combine multiple weak predictors to form one strong predictor by continuously correcting the errors from previous models. By updating such a model, the previous version of the model can be fine-tuned, and this fine-tuning may include retraining the model with a larger data set 370.

[0102] A particularly advantageous option may be a Gaussian process, which may be incorporated to provide measures of prediction and uncertainty.

[0103] FIG. 3f schematically shows an example of one embodiment of training data 373 for a multi-task Gaussian process that implements a joint model of the physical quantities of the target system and the physical quantities of the auxiliary system.

[0104] The training data 373 includes the training data 372 but also includes auxiliary training data 371. The auxiliary training data 371 includes a plurality of pairs of an auxiliary state 321 of the auxiliary system and a corresponding physical quantity 314.

[0105] The model 381 may include a multi-task Gaussian process that implements a joint model of the physical quantities of the target system and the physical quantities of the auxiliary system. The joint model 381 is configured to predict the physical quantities regarding both the target system and the auxiliary system, thereby causing transfer learning between the two domains.

[0106] In the following, numerous applications of embodiments for training a target machine learning model are presented.

[0107] In one embodiment, model 381 is configured to analyze data from various different types of sensors in order to obtain measurements of the environment. These sensors can generate various different forms of data including digital images such as video images, radar images, LiDAR images, ultrasonic images, motion images and thermal images, as well as audio signals. In particular, model 381 may be configured to receive sensor signals and derive additional information regarding elements encoded in those signals. This function enables indirect measurements based on direct sensor signals. For example, model 381 may be a so-called virtual sensor. Physical quantities that are theoretically measurable but not desirable to obtain in non-prototype systems can be predicted from other aspects of the state of the system, for example, from sensor readings. In one embodiment, model 381 is configured to determine continuous values from sensor data. These continuous values can include measurements such as distance, speed or acceleration. Model 381 can also track items within the data.

[0108] In one embodiment, model 381 may be incorporated into a controller configured to calculate control signals for a wide variety of technical systems. These technical systems may be diverse, ranging from computer-controlled machines such as home appliances and manufacturing machines to information transmission systems such as monitoring systems or medical imaging systems. The control signals can include, for example, start / stop signals, temperature settings, speed control, emergency stops, and the like.

[0109] The target system can be monitored using model 381. For example, monitoring can include · obtaining the state of the target system, and · applying the target machine learning model to the obtained state, thereby obtaining a prediction of a physical quantity as an output may include.

[0110] The predicted physical quantity can be reported to the user of the target system. The predicted physical quantity can be compared with a target value, and if the predicted physical quantity deviates from the target value by more than a threshold, the state can be corrected. For example, it can be tested whether the predicted physical quantity is outside a desired range, and if it is outside the range, repair can be initiated.

[0111] It should be noted that the physical quantity and the safety value may be the same, or the physical quantity may be included in the safety value. For example, in a dynamic system such as a robotic system, for example, the safety value is monitored in real time, and if the state begins to develop into an unsafe output, the target system can be shut down and then repair can be initiated. If the dynamics are sufficiently slow compared to the response time of the system, it is effective to monitor in real time.

[0112] The target system can be controlled using the model 381. For example, a plurality of states can be obtained, and the target machine learning model is applied to each of the plurality of states, whereby predictions of physical quantities for each of the plurality of states are obtained as outputs, and one state is selected from the plurality of states depending on the predicted physical quantity, and the target system is configured according to that state.

[0113] For example, the target system includes at least a partially autonomous vehicle, and the states include a control state, for example, one or a combination of one or more of throttle, brake, and steering angle, and a vehicle state, for example, one or a combination of one or more of the vehicle's position, orientation, longitudinal speed and lateral speed, gearbox position, and engine RPM, and a road state, for example, surrounding objects and traffic, and one or a combination of road information, and the output of the machine learning includes the predicted changes in the vehicle state.

[0114] In what follows, some further alternative improvements, details, and embodiments are illustrated in more mathematical language. These additional examples are used as additional embodiments and improvements, but are not intended to limit the possible scope of the embodiments.

[0115] The regression output and safety value are considered as follows. Each input

Number

Number

Number

[0116] For example,

Number

Number

Number

Number

[0117] Typically, a small number of safe observations

Mathematics

Mathematics

Mathematics

Mathematics

[0118] The goal to be achieved is often to evaluate the function

Mathematics

Mathematics

Mathematics

Mathematics

Number

[0119] This problem formulation applies to both active learning (AL) and Bayesian optimization. Embodiments focus on AL, although the embodiments may be changed to BO as needed. The goal is to make accurate predictions f(X) using the evaluation, and the points selected will prioritize a more general understanding over the space X until reaching the safety constraints.

[0120] Gaussian process (GP): GP is a stochastic process specified by a mean and a kernel function. Without loss of generality, assume that the mean of GP is zero. Furthermore, in the absence of prior knowledge of the data, it is common to assume that the dominant kernel is static. For example,

Number

Number

[0121]

Number

Number

Number

Number

Number

Number

Number

[0122] Safe learning: The core of a safe learning method is to compare the safety reliability boundary with a threshold value and define a safe set S N ⊆ X pool as

Number

Number

Number

Number

[0123] In each iteration, new points are queried by mapping candidates for safe inputs to acquisition scores,

Number

[0124] A well-known acquisition function in the AL problem is the predictive entropy, i.e.,

Number

[0125]

Number

[0126]

Table 1

[0127]

Number

Number

Number

Number

[0128] and

Number

Number

Number

Number

[0129]

Number

Number

Number

Number

Number

Number

Number

[0130]

Number

[0131] can more easily achieve global exploration. [Number] Acceleration during experiments by auxiliary precomputation:

[0132] The calculation of [Number] involves a three-dimensional time complexity [Number] This calculation is also used to fit the model. General fitting techniques include Type II maximum likelihood estimation (Type II ML), Type II maximum a posteriori estimation (Type II MAP), and Bayesian processing for the kernel and noise parameters, all of which involve the marginal likelihood [Number] includes calculating. Bayesian processing is not preferred because it takes time for MC sampling.

[0133] Here, the achievement goal is to

Number

[0134] For each

Number

Number

[0135]

Number

Number

Number

Number

[0136] The learning procedure is summarized in the following algorithm.

[0137]

Table 2

[0138] In each iteration (line 4a), the time complexity is

Number

Number

Number

Number

Number

[0139] Kernel selection: Multi-output kernel

Number

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

[0140] Hierarchical GP (HGP):

Mathematics

[0141] In the experiment, the above modular algorithm (Algorithm 2) using HGP as the main pipeline according to the present invention is implemented. As a baseline comparison, a first sequential learning algorithm is executed using a conventional single-output GP. Furthermore, a general but slow framework that utilizes generally used LMC together with a vanilla sequential learning model fitting strategy (Algorithm 1) is compared with the main pipeline. The base kernels k s , k t , k l and the kernel are all Matern-5 / 2 kernels with a length scale parameter

Number

[0142] However, the pairing of the modularized computational scheme according to the present invention with a general LMC kernel would be useful in closely related settings, for example, (i) in a data set where two or more auxiliary tasks can be utilized, or (ii) in a sequential learning scheme where GP is re-fitted only after receiving a batch of a series of query points. This combination was not used in the experiments.

[0143] Experiment Three experimental setups are compared, namely, Algorithm 2 using multi-output HGP, named efficient transfer, and Algorithm 2 using multi-output LMC, named transfer, which is a flexible but slow transfer learning framework, and the conventional Algorithm 1 using single-output GP and Matern-5 / 2 kernel, named baseline. For the safety tolerance, β Ν = 4 is always fixed, for example,

Equation

[0144] Figures 4a to 4f show the training results for approximating the function f with input dimension D = 1.

[0145] Figures 5a to 5f show the training results for approximating the function f with input dimension D = 2. GP data X = [-2, 2] 2 in the experiment of safe AL, AL in f constrained by the additional safety function q ≥ 0. X pool is the X discretized from N pool = 5000.

[0146] The drawings show the results for efficient transfer, transfer, and baseline.

[0147] All the drawings show the number of iterations on the horizontal axis.

[0148] Figures 4a and 5a show the RMSE on the vertical axis.

[0149] Figures 4b and 5b show the true positive (TP) region portions on the vertical axis.

[0150] Figures 4c and 5c show the fitting time in seconds on the vertical axis.

[0151] Figures 4d and 5d show the false positive (FP) region portions on the vertical axis.

[0152] Figures 4e and 5e show the count of the queried regions on the vertical axis.

[0153] Figures 4f and 5f show the ratio of unsafe queries as a percentage on the vertical axis.

[0154] In FIGS. 4a to 5f, a baseline 413 which is safe active learning without using auxiliary data, a transition 412 which is an embodiment of active learning using auxiliary data, in this case an LMC without using modularization is used, and an efficient transition 411 which is an embodiment of active learning using auxiliary data where the model data related to the auxiliary system remains unchanged during update, in this case, an HGP using pre-computed fixed auxiliary knowledge is used are compared. M s = 250, and N ranges from 20 (0th iteration) to 120 (100th iteration). Those results are the average of 100 experiments and 1 standard error. The safe region is predicted using an alternative GP model. The test points of RMSE are sampled from the true safe region.

[0155] Experiments are performed on the simulated data and the engine data. All simulation data have an input dimension D which is 1 or 2. Therefore, it is analytically and computationally possible to cluster the unconnected safe regions via the labeling algorithm of the connected components. This means that in each iteration of the experiment, it is possible to track which safe region each observation belongs to.

[0156] Furthermore, the safety region learned by the alternative model is tracked. Since the data set according to the present invention in the experiment is prepared as a query to be executed,

Number

[0157]

Number

Number

Number

[0158] An auxiliary data set and a target data set are generated, and each of the auxiliary data set and the target data set has two or more separate safety regions, and a part of the safety region is also safe in other data sets. Specifically, multi-output GP samples are generated. The first output is treated as an auxiliary task according to the present invention, and the second output is treated as a target task. The data set is generated such that the target task has at least two separate safety regions. In this case, each region has a common safety region shared with the auxiliary, and this shared region is larger than 10% of the entire space.

[0159] Figures 4a through 4f relate to a data set where D = 1, f is the main function, and an additional q ≥ 0 is a safety constraint.

[0160] Figures 5a through 5f relate to a data set where D = 2, f is the main function, and an additional q ≥ 0 is a safety constraint.

[0161] Additional experiments were conducted on data sets where D = 1 or D = 2 and q = f ≥ 0 is a safety constraint.

[0162] Twenty data sets are generated for each type, and the AL experiment is repeated 5 times for each data set. When D = 1, M s = 100, N = 10 (initial) is set and queried over 50 iterations (N = 10 + 50). When D = 2, M s = 250, N = 20 (initial) is set and queried over 100 iterations (N = 20 + 100). N pool is always set to 5000. Since all data sets have values centered around 0, the constraint q ≥ 0 according to the present invention indicates that approximately half of the space is safe.

[0163] In Figures 4a through 5f, transfer learning according to one embodiment achieves a significantly larger and more accurate coverage rate of the safe set (the TP region is larger and the FP region is smaller).

[0164]

Number

Number

[0165]

Table 3

[0166] It should be noted that the method according to an embodiment is faster in learning (see FIGS. 4a and 5a). This does not come at the expense of safety. Because the safety ratios of efficient transfer and transfer are not worse than those of the baseline.

[0167] Engine Modeling Safe AL experiments were performed on two data sets measured under different conditions from the same prototype engine. Both data sets measure temperature, roughness, HC emissions, and NOx emissions.

[0168] Interestingly, the safe set for this target task is not clearly separated into multiple distinct regions. Thus, traditional methods ultimately identify most of the safe region. Nevertheless, it is still confirmed that the RMSE is significantly better and the data consumption is significantly less when the coverage rate of the safe set is large.

[0169] The AL experiment for learning roughness was constrained by the normalized temperature value q ≤ 1.0. The safe set is about 0.5293 of the entire space. The data set has two free variables and two fixed context inputs. Since the context inputs are recorded with noise, the values are interpolated using a multi-output GP simulator trained on the complete data set. This experiment is carried out under semi-simulated conditions. M s = 500, N = 20 (initially), N pool = 3000 is set and queried over 100 iterations (N = 20 + 100).

[0170] FIG. 6 schematically shows an example of one embodiment of a method 600 for training a target machine learning model for a target system. The method 600 is computer-implemented. The target system is configurable to one of a plurality of possible states, · During operation, the target system enables the following measurements, namely, · The target system enables measurement of at least one target safety value of the target system, and the configuration of the target system is defined as being safe if at least one target safety value is within a safe region, · The target system enables measurement of physical quantities useful for monitoring or controlling the target system, · The target machine learning model is configured to receive the state of the target system as input and generate a prediction of a physical quantity as output. The method (600) comprises · obtaining (610) auxiliary training data corresponding to an auxiliary system, the auxiliary training data including a plurality of pairs of an auxiliary state and an auxiliary safety value, · initializing (620) a multi-task Gaussian process implementing a joint model of a safety value of the target system and a safety value of the auxiliary system, the multi-task Gaussian process receiving as input a state that can be a state of the auxiliary system or a state of the target system and generating respectively a prediction regarding a target safety value or a prediction regarding an auxiliary safety value, · iteratively training (630) the target machine learning model and iteratively training (630) the target machine learning model comprises · selecting (640) one state from a plurality of possible states regarding the target system, wherein the target safety value predicted by the multi-task Gaussian process for the selected state is within a safe region, · obtaining (650) a physical quantity and at least one target safety value for the selected state in the target system, the physical quantity and the safety value being measured via a sensor, · updating (660) the multi-task Gaussian process using the selected state and the corresponding target safety value, · updating (670) the target machine learning model using the selected state and the corresponding physical quantity and the method includes.

[0171] For example, the present method may be a computer-implemented method. For example, obtaining the auxiliary training data may use a communication interface, such as a network, or a storage interface, such as an API. For example, obtaining a physical quantity and at least one target safety value for a selected state in the target system may include issuing an instruction to the target system via, for example, a communication interface to perform a setting according to the new state and receive measured values of the physical quantity and the safety value from a sensor.

[0172] For example, a computer processor can execute initializing a multi-task Gaussian process, selecting a state, and updating the multi-task Gaussian process and the target machine learning model.

[0173] As will be apparent to those skilled in the art, numerous different ways of implementing the present method are possible. For example, the order of each step can be implemented in the presented order, but the order of each step can also be changed, or several steps can be executed in parallel. Further, other method steps can be inserted between steps. The inserted steps may represent an improved form of the present method as described herein, or may not be related to the present method. For example, several steps can be executed at least partially in parallel. Further, a given step may not be fully completed before the next step begins.

[0174] Embodiments of the method may be implemented using software, which includes instructions for causing a processor system to implement method 600. The software may include only those steps executed by a particular sub-entity of the system. The software may be stored on a suitable storage medium such as a hard disk, flexible disk, memory, optical disk, etc. The software may be transmitted as a signal, either wired or wirelessly, or using a data network, for example, the Internet. The software may be available for download and / or for remote use on a server. Embodiments of the method may be implemented using programmable logic, for example, a bitstream arranged to configure a field programmable gate array (FPGA), to implement the method.

[0175] It will be understood that the subject matter disclosed herein also extends to a computer program configured to implement the subject matter disclosed herein, particularly a computer program on or in a carrier. The program may be in the form of, for example, auxiliary code, object code, intermediate auxiliary code and object code, as in a partially compiled form, or in any other form suitable for use in implementing embodiments of the method. Embodiments related to computer program products include computer-executable instructions corresponding to each processing step of at least one of the methods described above. These instructions may be subdivided into subroutines and / or stored in one or more files that may be statically or dynamically linked. Other embodiments related to computer program products include computer-executable instructions corresponding to each of the devices, units, and / or parts of at least one of the systems and / or products described above.

[0176] FIG. 7a shows a computer-readable medium 1000 having a writable portion 1010 and a computer-readable medium 1001 having a writable portion as well. The computer-readable medium 1000 is shown in the form of an optically readable medium. The computer-readable medium 1001 is shown in the form of an electronic memory, in this case in the form of a memory card. The computer-readable media 1000 and 1001 can store data 1020, which, when executed by a processor system, can indicate instructions for causing the processor system to implement an embodiment of a method for training a target machine learning model for a target system according to an embodiment. The computer program 1020 may be embodied as a physical mark on the computer-readable medium 1000 or may be embodied by magnetization of the computer-readable medium 1000. However, any other suitable embodiments are also envisioned. Further, although the computer-readable medium 1000 is shown as an optical disk herein, the computer-readable medium 1000 may be any suitable computer-readable medium such as a hard disk, solid state memory, flash memory, etc., and may be non-recordable or recordable, as will be understood. The computer program 1020 includes instructions for causing the processor system to implement the aforementioned method for training a target machine learning model for a target system.

[0177] FIG. 7b shows a schematic diagram of a processor system 1140 according to an embodiment of a system for training a target machine learning model for a target system. The processor system includes one or more integrated circuits 1110. The architecture of the one or more integrated circuits 1110 is schematically shown in FIG. 7b. Circuit 1110 includes a processing unit 1120, such as a CPU, for executing computer program components, implementing the method according to an embodiment, and / or implementing modules or units of computer program components. Circuit 1110 includes a memory 1122 for storing programming code, data, etc. A part of the memory 1122 may be considered read-only. Circuit 1110 may include a communication element 1126, such as an antenna, a connector, or both. Circuit 1110 may include an application-specific integrated circuit 1124 for performing some or all of the processing defined in the method. The processor 1120, the memory 1122, the application-specific IC 1124, and the communication element 1126 may be interconnected via an interconnect 1130, such as a bus. The processor system 1110 may be configured to communicate in a contact and / or non-contact manner using an antenna and / or a connector, respectively.

[0178] For example, in one embodiment, the processor system 1140, such as a system for training a target machine learning model for a target system, may include a processor circuit and a memory circuit, and the processor is configured to execute software stored in the memory circuit. For example, the processor circuit may be an Intel Core i7 processor, an ARM Cortex-R8, etc. The memory circuit may be a ROM circuit or a non-volatile memory, such as a flash memory. The memory circuit may be a volatile memory, such as an SRAM memory. In the latter case, the device may include a non-volatile software interface configured to provide software, such as a hard drive, a network interface, etc.

[0179] Memory 1122 may be regarded as a “non-transitory machine-readable medium.” As used herein, the term “non-transitory” is understood to include all forms of storage that do not include transient signals, but include both volatile and non-volatile memory.

[0180] Note that the above-described embodiments are illustrative rather than limiting of the subject matter disclosed herein, and those skilled in the art will be able to design numerous alternative embodiments.

[0181] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The use of the verb “comprise” and its conjugations does not exclude the presence of elements or steps other than those recited in the claim. The indefinite article “a” or “an” preceding an element does not exclude the presence of a plurality of such elements. Expressions such as “at least one” when preceding a list of elements represent alternatives selected from all or any subset of the elements in the list. For example, the expression “at least one of A, B, and C” should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The subject matter disclosed herein may be implemented as a combination of hardware that includes several distinct elements and by a suitably programmed computer. In apparatus claims listing several parts, some of these parts may be embodied by the same hardware item. The mere fact that a particular means is recited in several mutually different dependent claims does not indicate that a combination of these means cannot be advantageously used.

[0182] In the claims, the reference signs in parentheses refer to the reference signs in the drawings illustrating the embodiments or the mathematical formulas of the embodiments, thereby enhancing the understandability of the claims. These reference signs should not be construed as limiting the claims.

Explanation of Signs

[0183] Explanation of Signs The following list of reference signs and abbreviations corresponds to FIGS. 1a to 3f, 7a, 7b and is provided to facilitate the interpretation of the drawings and should not be construed as limiting the claims.

[0184] 101 Safety Region 111 Initial Target State 112 Selected Target State 113 Auxiliary State 200, 202 Training System 210 Target System 220 Training Device 230 Controller Device 211 Processor System 212 Storage 213 Communication Interface 221 Processor System 222 Storage 223 Communication Interface 231 Processor System 232 Storage 233 Communication Interface 224 Database 272 Computer Network 100 System for Training a Target Machine Learning Model for a Target System 310 Auxiliary System 311 Auxiliary Safety Value Sensor 312 Auxiliary Safety Value 313 Physical Quantity Sensor 314 Physical Quantity 321 State 330 Target System 331 Target safety value sensor 332 Target safety value 333 Physical quantity sensor 334 Physical quantity 341 State of the target model 342 Selected state of the target model 350 Training data for safety values 351 Training data for auxiliary safety values 352 Training data for target safety values 361 Joint model 362 State selection unit 370 Training data for the target machine learning model 371 Training data for auxiliary physical quantities 372 Training data for target physical quantities 373 Training data for the joint model of the physical quantities of the target system and the physical quantities of the auxiliary system 381 Machine learning model 411 Efficient transfer 412 Transfer 413 Baseline 1000, 1001 Computer-readable medium 1010 Writable part 1020 Computer program 1110 Integrated circuit 1120 Processing unit 1122 Memory 1124 Application-specific integrated circuit 1126 Communication element 1130 Interconnect 1140 Processor system

Claims

1. 1. A computer-implemented method (600) for training a target machine learning model for a target system in an engineered process or machine, the target system being configurable into one of a plurality of possible states; During operation, the target system allows the following measurements: the target system allows for the measurement of at least one target safety value of the target system, a configuration of the target system being defined as safe by the at least one target safety value being within a safety region; the target system allows the measurement of physical quantities that are useful for monitoring or controlling the target system; the target machine learning model is configured to receive as input a state of the target system and to generate as output a prediction of the physical quantity; The method for training a target machine learning model comprises: - acquiring (610) auxiliary training data corresponding to an auxiliary system, the auxiliary training data including a plurality of pairs of auxiliary states and auxiliary safety values; - initializing (620) a multitasking Gaussian process implementing a joint model of the safety value of the target system and the safety value of the auxiliary system, the multitasking Gaussian process receiving as input a state of the auxiliary system or a state that may be a state of the target system, and generating a prediction for the target safety value or a prediction for the auxiliary safety value, respectively; - iteratively training (630) the target machine learning model; Including, Iteratively training (630) the target machine learning model comprises: selecting (640) a state from the plurality of possible states for the target system, where a target safety value predicted by the multitasking Gaussian process for the selected state is within the safety region; and - obtaining (650) the physical quantity and at least one target safety value for the selected state of the target system, comprising configuring the target system according to the selected state, the physical quantity and the safety value being measured via sensors; Updating (660) the multitasking Gaussian process with the selected states and the corresponding target safety values; - updating (670) the target machine learning model using the selected states and the corresponding physical quantities; A method comprising:

2. the target system is a vehicle or robotic device configured for at least partially autonomous movement; The method of claim 1.

3. the auxiliary system being a computer simulation of the target system; The method according to claim 1 or 2.

4. the target system and the auxiliary system each comprise the same type of engineered process or machine, such that measurements from the auxiliary system are predictive of the target system; The method according to claim 1 or 2.

5. the target machine learning model comprises a multitask Gaussian process; the auxiliary training data includes an auxiliary physical quantity for the auxiliary state within the auxiliary training data; The method comprises: initializing a multitasking Gaussian process implementing a joint model of the physical quantities of the target system and the physical quantities of the auxiliary system, the multitasking Gaussian process receiving as inputs states of the auxiliary system or states which may be states of the target system, and generating predictions for the target physical quantities or predictions for the auxiliary physical quantities, respectively; Including, 5. The method according to any one of claims 1 to 4.

6. the multi-tasking Gaussian process having a portion associated with the auxiliary system; updating the multi-tasking Gaussian process with the selected state and the corresponding target safety value leaves the portion of the multi-tasking Gaussian process associated with the auxiliary system unchanged.

6. The method according to any one of claims 1 to 5.

7. selecting one state of the plurality of possible states includes optimizing an acquisition function.

7. The method according to any one of claims 1 to 6.

8. The target system allows the measurement of at least one target safety value from temperature, pressure, vibration level, noise level, humidity, current or voltage level, flow rate, force, torque, speed, RPM, chemical concentration, state of charge.

8. The method according to any one of claims 1 to 7.

9. The target system comprises: Any of the amounts as set forth in claim 8, or - Location of obstacles in the vicinity of the target machine enabling measurement of at least one physical quantity from 9. The method according to any one of claims 1 to 8.

10. A computer-implemented method for targeting systems in engineered processes and machines using a target machine learning model trained according to any one of claims 1 to 7, comprising: - obtaining a state of the target system; - applying the target machine learning model to the obtained states, thereby obtaining as output a prediction of a physical quantity; A method comprising:

11. The method comprises: monitoring the target system, the acquired state being a state for which the target system is configured; testing whether the predicted physical quantity is outside a desired range; and initiating repairs if so; and / or Controlling the target system, comprising: a plurality of states being obtained, the target machine learning model being applied to each of the plurality of states, thereby obtaining as output a prediction of the physical quantity for each of the plurality of states; selecting one of the plurality of states depending on the predicted physical quantity; and configuring the target system according to the state. The method of claim 10.

12. the target system includes an at least partially autonomous vehicle; The state is A control state, e.g., one or more combinations of throttle, brake, steering angle; Vehicle conditions, such as one or more combinations of vehicle position, orientation, longitudinal and lateral speeds, gearbox position, engine RPM; Road conditions, e.g., one or more combinations of surrounding objects and traffic and road information; [0033] The machine learning output includes a predicted change in vehicle state.

12. The method according to any one of claims 1 to 11.

13. one or more processors; one or more storage devices; A system comprising: The one or more storage devices store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for a method according to any one of claims 1 to 12. system.

14. A transitory or non-transitory computer-readable medium (1000) comprising data (1020) representing instructions that, when executed by a processor system, cause the processor system to perform a method according to any one of claims 1 to 12.