System, device, and method for controlling physical or chemical process
Patent Information
- Application Number
- JP2022152669
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-27
- Filing Date
- 2022-09-26
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-09-26
AI Technical Summary
Existing model-based optimization methods for production and machining processes require significant measurement data, leading to high costs and potential inaccuracies due to variations in process characteristics and measurement errors, which are not adequately addressed by current transfer learning techniques.
A method that reduces the number of required measurements by utilizing transfer learning with Gaussian processes, considering the uncertainty of both the source and target models, allowing for efficient adaptation of models with fewer data points and improved accuracy through the use of sequential and enhanced hierarchical Gaussian processes.
This approach significantly reduces training costs and time while enhancing the accuracy of models by accounting for uncertainties, enabling efficient optimization and control of physical or chemical processes with limited data.
Smart Images

Figure 00000029_0000 
Figure 00000029_0001 
Figure 00000029_0002
Abstract
Description
[Technical Field]
[0001] Various embodiments relate generally to systems, devices, and methods for controlling physical or chemical processes. [Background technology]
[0002] In production and machining processes (e.g., drilling, milling, heat treatment, etc.), process parameters, such as process temperature, process time, vacuum or gas atmosphere, etc., are adjusted to obtain desired properties, such as hardness, strength, thermal conductivity, electrical conductivity, density, microstructure, macrostructure, chemical composition, etc., of the workpiece. These process parameters can be determined by model-based optimization methods, such as Bayesian optimization. A model for the production or machining process can then be determined based on measurement data. However, this may require a large amount of measurement data and thus significant costs (e.g., time and / or financial expenditure). These costs can be reduced by determining a model in association with the measurement data based on an already trained model that represents a process related to the production or machining process (also known as transfer learning). For example, two models can represent drilling or milling on different machines (and with equivalent process parameters). In this case, the already trained model can be used as the basis for the model to be trained, thereby reducing the amount of measurement data required.
[0003] The publication "Google Vizier: A Service for Black-Box Optimization" by D. Golovin et al. (KDD 2017 Applied Data Science Paper, 2017) (hereafter referred to as Reference [1]) describes hierarchical transfer learning, where a model is derived as a Gaussian process based on a model that has already been trained as a Gaussian process and based on measurement data.
[0004] To train a model according to Bayesian optimization, new measurement points can be found at each iteration using an acquisition function. Examples of acquisition functions are given in the publications "Efficient global optimization of expensive black-box functions" by D. Jones et al., Journal of Global Optimization, 1998 (hereafter referred to as Reference [2]) and "Gaussian process optimization in the bandit setting: No regret and experimental design" by N. Srinivas et al., Proc. International Conference on Machine Learning (ICML), 2010 (hereafter referred to as Reference [3]). [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Publication “Google Vizier: A Service for Black-Box Optimization” by D. Golovin et al. (KDD 2017 Applied Data Science Paper, 2017) [Non-patent document 2] Publication: "Efficient global optimization of expensive black-box functions," by D. Jones et al. (Journal of Global Optimization, 1998) [Non-patent document 3] Publication “Gaussian process optimization in the bandit setting: No regret and experimental design” by N. Srinivas et al. (Proc. International Conference on Machine Learning (ICML), 2010) Summary of the Invention [Problem to be solved by the invention]
[0006] Measurements characterizing at least one property of a production or machining process and / or at least one effect of this process on a workpiece may vary even when process parameters are identical. Such variation may arise from the process itself, the workpiece, or measurement errors. Such variation and / or relatively little measurement data may cause a trained model to have high inaccuracies and / or uncertainties in regions with little or no measurement data. According to various embodiments, it has been recognized that the training cost (e.g., optimization efficiency) in transfer learning can be reduced when the uncertainty of both the already trained model and the uncertainty of the model to be trained are taken into account. [Means for solving the problem]
[0007] The method with the features of independent claim 1, the device with the features of independent claim 5, and the system with the features of independent claim 6 enable efficient optimization of models. In this case, the number of required measurement data is reduced, which can reduce the costs (e.g., time and / or financial expenditure) for training the model. In particular, within the framework of transfer learning, fewer measurements are needed to adapt a model already trained for a similar process to a new process. Furthermore, the consideration of the uncertainty of the already trained model as described herein increases the accuracy of the trained model for the new process (e.g., increased accuracy of the expected value of the trained model and increased accuracy of the considered uncertainty of this expected value). In various embodiments, it is possible to train a model even when there is little data and / or when the model already trained for a similar process has high uncertainty.
[0008] A method having the features of independent claim 1 constitutes a first example, in which in particular at least one (e.g. exactly one or more than one) hyperparameter of the first posterior model cannot be re-optimized.
[0009] The models described herein may be statistical models or any kind of mathematical representation that describes the relationship between input variables and output variables of a physical or chemical process, such as a Gaussian process. The input and / or output variables of a physical or chemical process may be input or output variables of another system. For example, the input variables may be parameters of a simulation (e.g., of a physical or chemical process). For example, the output variables may be an approximation error. For example, the output variables may be hyperparameters and one or more loss values of a machine learning-based model.
[0010] When determining the second posterior model, all hyperparameters of the covariance function of the first posterior model can be kept unchanged. The features described in this paragraph are combined with the first example to form a second example. Specifically, the second posterior model can be determined using hyperparameter optimization of the second model and tuning to the known second measurement point, where the hyperparameters of the first posterior model can be kept unchanged. Specifically, all hyperparameters of the first posterior model cannot be re-optimized.
[0011] These two methods for obtaining the second posterior model according to the first example allow for a reduction in computational complexity and therefore computational cost by defining the hyperparameters of the first posterior model (see, for example, Table 1 and the accompanying description).
[0012] The method further comprises: selecting new input parameter values for at least one input variable of the physical or chemical process using an acquisition function based on the second posterior model; measuring an output value of at least one output variable that is assigned to the input parameter value, the selected new input parameter value and the measured new output value forming a new measurement point; and adapting a second posterior model using the new measurement points, the adapted second posterior model representing a physical or chemical process for the known and new measurement points; The adapted second posterior model is used to control a physical or chemical process. The features described in this paragraph may be combined with the first or second examples to form a third example.
[0013] These two approaches for determining the second posterior model according to the first or second example may have reduced computational complexity compared to other methods that also allow propagation of uncertainty (see, for example, Table 1 and the accompanying description).
[0014] The third posterior model may represent a relationship between at least one input variable and at least one output variable of a further process related to a physical or chemical process, and each further measurement point of the known further measurement points may have an input parameter value of at least one input variable of the further process and an output value of at least one output variable of the further process that is assigned to the input parameter value. This method is and determining the first posterior model incorporating a third posterior model, wherein determining the first posterior model incorporating the third posterior model comprises: determining each separate Gaussian process by forming an expectation of the Gaussian process, where the function is derived from the third posterior model, and determining a plurality of other Gaussian processes using a separate common covariance function; determining a second prior model as the average of a number of other Gaussian processes and adjusting the second prior model to other known measurement points to determine the first posterior model; or adjusting each of the plurality of other Gaussian processes to the known other measurement points, and determining a first posterior model as an average value of the adjusted plurality of other Gaussian processes; The features described in this paragraph may be combined with one or more of the first to third examples to form a fourth example. Optionally, the first posterior model may be determined by another method. For example, the first posterior model may be determined using a hierarchical Gaussian process according to Reference [1], where β=1.
[0015] Specifically, the second posterior model can be determined as a sequence of multiple other models (e.g., the third posterior model and the first posterior model). These two approaches for determining the second posterior model from one or more examples of the first to third examples can significantly improve the efficiency of training a model based on multiple other models compared to other methods (e.g., see illustration 352 in FIG. 3B).
[0016] Controlling a physical or chemical process using a second posterior model determining input parameter values for at least one input variable that are assigned to a desired output value of at least one output variable of the physical or chemical process according to the second posterior model; and controlling a physical or chemical process in accordance with the determined input parameter values. The features described in this paragraph may be combined with one or more of the first through fourth examples to form a fifth example.
[0017] Specifically, the second posterior model can take into account the uncertainty of the first posterior model (and the third posterior model according to the fourth example), which results in improved accuracy of the second posterior model (e.g., of information regarding expected values and / or uncertainty).
[0018] The apparatus may be configured to perform the method according to one or more of the first through fifth examples. An apparatus having the features described in this paragraph constitutes a sixth example.
[0019] The apparatuses for performing (e.g., physical or chemical) processes described herein may be any type of computer-controlled device, such as, for example, a robot (e.g., a manufacturing robot, a maintenance robot, a household robot, a medical robot, etc.), a vehicle (e.g., an autonomous vehicle), a household appliance, a production machine, a personal assistant, an access control system, etc. However, the apparatuses for performing (e.g., physical or chemical) processes described herein may also be non-computer-controlled devices. For example, the apparatuses may be manually controlled by a user to perform (e.g., physical or chemical) processes. In general, the apparatuses may be, for example, production devices for manufacturing products or processing devices for processing workpieces (e.g., drilling or milling devices), robotic devices for performing movements, devices for, for example, pharmaceutical agent design, devices for, for example, designing systems for aeronautics and spaceflight, etc.
[0020] The system is an apparatus configured to perform a physical or chemical process; a control device configured to determine, for a desired output value of at least one output variable of a physical or chemical process, input parameter values assigned to the desired output value using the second posterior model determined according to one or more of the first to fifth examples, and to control the device to perform the physical or chemical process according to the determined input parameter values; A system having the features described in this paragraph constitutes a seventh example.
[0021] In this system, the physical or chemical processes are Machining of workpieces, Adjustment (e.g., calibration) of equipment (e.g., instruments and / or measuring devices); Manufacturing of products, or Robot arm movement The features described in this paragraph may be combined with the seventh example to form an eighth example.
[0022] The computer program may include instructions that, when executed by a processor, cause the processor to perform a method according to one or more of the first through fifth examples. A computer program having the features described in this paragraph constitutes a ninth example. It should be understood that the processor may generate instructions for controlling a physical or chemical process (e.g., instructions for controlling an apparatus).
[0023] A computer-readable medium (e.g., a computer program product, a non-transitory storage medium, a non-transitory storage medium, a non-volatile storage medium) may store instructions that, when executed by a processor, cause the processor to perform a method according to one or more of the first through fifth examples. A computer-readable medium having the features described in this paragraph constitutes a tenth example.
[0024] An embodiment of the invention is shown in the drawing and is explained in more detail below. [Brief explanation of the drawings]
[0025] [Figure 1A] FIG. 1 illustrates the training of a first posterior model representing a process performed by a first device, according to various embodiments. [Figure 1B] FIG. 1 illustrates the training of a first posterior model representing a process performed by a first device, according to various embodiments. [Figure 1C] FIG. 1 illustrates control of a first device by a trained first posterior model, according to various embodiments. [Figure 2A] FIG. 1 illustrates the implementation of a physical or chemical process using a second device, according to various embodiments. [Figure 2B] FIG. 10 illustrates control of a second device by a trained second posterior model, according to various embodiments. [Figure 2C] FIG. 10 illustrates the adaptation of a second posterior model, according to various embodiments. [Figure 2D] FIG. 10 illustrates control of a second device by an adapted second posterior model, according to various embodiments. [Figure 3A] 1A-1C illustrate exemplary visualizations of various models or the progress of optimization based on various models, according to various embodiments. [Figure 3B] 1A-1C illustrate exemplary visualizations of various models or the progress of optimization based on various models, according to various embodiments. [Figure 4] FIG. 1 illustrates a method for controlling a physical or chemical process, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0026] In one embodiment, a "computer" may be understood as any kind of logic-implementing entity, which may be hardware, software, firmware, or a combination thereof. Thus, in one embodiment, a "computer" may be a hardwired logic circuit or a programmable logic circuit, such as a programmable processor, such as a microprocessor (e.g., CISC (a processor with a large instruction stock) or RISC (a processor with a reduced instruction stock)). A "computer" may include one or more processors. A "computer" may also be software implemented or executed by a processor, such as any kind of computer program, for example, a computer program using virtual machine code such as Java. Other modes of implementation of the functions described in detail below may be understood as a "computer" in accordance with optional embodiments.
[0027] Various embodiments relate to systems, devices, and methods for controlling physical or chemical processes, where a model representing a physical or chemical process may be trained using transfer learning based on another previously trained model representing a process related to the physical or chemical process, taking into account the uncertainty of both the previously trained model and the uncertainty of the model to be trained. Specifically, this can significantly improve training efficiency, thereby allowing the model to be trained with significantly less data and / or reducing the computational costs of developing the model. This not only reduces the time and / or financial costs of training, but also improves the accuracy of the trained model for the physical or chemical process.
[0028] According to various embodiments, the second device 208 may perform (e.g., complete) a physical or chemical process. Figure 1A illustrates a first system 100 that includes a first device 108. The first device 108 may be configured to perform another process related to the physical or chemical process.
[0029] A physical or chemical process as used herein may be any kind of technological process, such as, for example, a manufacturing process (e.g., manufacturing a finished product or intermediate product), a machining process (e.g., processing a workpiece), a control process (e.g., moving a robot arm), an adjustment process (e.g., calibrating a measuring instrument), etc. For example, it may be necessary to adjust different operating variables of an apparatus (e.g., within a calibration range), and a physical or chemical process may be such an adjustment. For example, a physical or chemical process during heat treatment using a furnace may be the calibration of the temperature and / or vacuum of the furnace.
[0030] It should be understood that other processes related to physical or chemical processes may also be physical or chemical processes as described herein.
[0031] Two processes may be related to each other in different ways. For example, this may be a generally similar process, such as drilling or milling a component, performed by different machines. It should be understood that generally similar processes on different machines may also produce distinct results. Two processes performed on the same machine may also be related to each other. For example, one process may be drilling a metal component and another process may be drilling a ceramic component. In general, two processes may be related to each other if the input variables of each process at least partially overlap and the output variables of each process at least partially overlap. Specifically, two related processes may have one or more identical input variables (e.g., process temperature, process time, and / or vacuum pressure in the case of heat treatment) and one or more identical output variables (e.g., hardness, strength, density, microstructure, macrostructure, and / or chemical composition in the case of heat treatment). Two processes may be related to each other if their respective models are suitable for transfer learning.
[0032] The first system 100 may include a first control device 106. The first control device 106 may be configured to control a first device 108 according to each (provided) first input parameter value 102 of at least one (i.e., exactly one or more) input variable. Specifically, the first control device 106 may control the interaction of the first device 108 with the environment according to one (or more) first input parameter values 102. As each first input parameter value 102 may be provided for each input variable (of the at least one input variable), this is also referred to hereinafter as at least one first input parameter value 102.
[0033] The term "controller" (also referred to as "control mechanism") may be understood as any type of logically implemented unit that may include, for example, circuits and / or processors capable of executing software, firmware, or a combination thereof stored on a storage medium, and that may instruct the device to perform a process, in this example. A controller may be configured to control, for example, by program code (e.g., software), the operation and / or adjustment (e.g., calibration) of a system, for example, a production system, a processing system, a robot.
[0034] An input parameter value as used herein may be a parameter value that represents an input variable, e.g., a physical or chemical variable, an applied voltage, a valve opening, etc. For example, an input variable may be a process-related property of one or more materials, e.g., hardness, thermal conductivity, electrical conductivity, density, microstructure, macrostructure, chemical composition, etc.
[0035] The first device 108 may be configured to perform another process related to a physical process or a chemical process according to at least one first input parameter value 102 .
[0036] The first system 100 may include one or more first sensors 110. The one or more first sensors 110 may be configured to detect a result of a process. The result of the process may be, for example, a property of a manufactured product or a property of a processed workpiece (e.g., hardness, strength, density, microstructure, macrostructure, chemical composition, etc.), the success or failure of a robot's skill (e.g., picking up an object), the resolution of an image captured by a camera, etc. The result of the process may be represented using at least one (i.e., exactly one or more) output variable. The one or more first sensors 110 may be configured to detect a first output value 112 of the at least one output variable. For example, a respective first output value 112 for each output variable of the plurality of output variables is detectable. Because a respective first output value 112 for each output variable (of the at least one output variable) is detectable, this is also referred to hereinafter as at least one first output value 112.
[0037] As described herein, detecting the results of a process using one or more sensors may occur during the process (e.g., in situ) and / or after the process (e.g., ex situ). For example, the model may represent a relationship between one or more input variables and at least two output variables, where an output value of one of the at least two output variables may be detected during the process and an output value of another of the at least two output variables may be detected after the process. As a specific example of detecting output values after the process, the process may be curing a workpiece in a furnace having temperature as an input variable. In this case, the output variable may be the hardness of the workpiece at room temperature after the curing process.
[0038] An output value, as used herein, may be a value representing an output variable of a process. The process output variable may be a characteristic of a product, a workpiece, a captured image, or other deliverable. However, the process output variable may also be success or failure (e.g., of a robot's skill). Specifically, at least one first output value 112 may result from at least one first input parameter value 102. The output variable may include an application-specific quality criterion. The output variable may be a parameter related to a component, such as a dimension or layer thickness, or a parameter related to a material, such as hardness, thermal conductivity, electrical conductivity, density, chemical composition, etc.
[0039] At least one first input parameter value 102 and at least one first output value 112 may form a first measurement point 130. The first control device 106 may be configured to control the first device 108 in tandem with respect to each (e.g., mutually different) input parameter value of the at least one input variable. One or more first sensors 110 may each detect at least one first output value 112. In particular, a plurality of first measurement points may be determined. The determined first measurement points may be determined from a known first measurement point D s(2) (i.e., D s or D s2 ) is also called.
[0040] First known measurement point D s(2) can be expressed by equation (1).
number
[0041] The first system 100 may include a storage device with at least one memory. s(2) may be stored or storable in memory, which may be used during processing performed by a processor.
[0042] The memory used in this embodiment may be a volatile memory, such as a DRAM (Dynamic Random Access Memory), or a non-volatile memory, such as a PROM (Programmable ROM), an EPROM (Erasable PROM), an EEPROM (Electrically Erasable PROM), or a flash memory, such as a floating gate memory device, a charge trap memory device, an MRAM (Magnetoresistive Memory) or a PCRAM (Phase Change Memory).
[0043] In various embodiments, model training or learning is described (see, for example, Algorithms 1, 2, 4, and 5). During model training or learning, a prior model (also referred to as a model a priori) is defined. Specifically, this prior model may exist without data (e.g., without known measurement points). The model may include hyperparameters that specify general characteristics of the distribution represented by the model. During model training or learning, optimization of the model's hyperparameters (also referred to as hyperparameter optimization) can be performed. The model's hyperparameter optimization can be performed based on known measurement points (e.g., measured measurement points and / or simulated measurement points). A model with optimized hyperparameters can also be referred to as a hyperparameter-optimized model. Specifically, the hyperparameters can be adapted using known measurement points. For example, if the known measurement points reveal that output values change relatively significantly when input parameter values are changed, the model can be adapted, for example, to increase the probability of a more rapidly changing function. During model training or learning, a posterior model (also referred to as a posteriori for a model) can be determined by tuning the model (e.g., a hyperparameter optimization model) to known measurement points. Here, posterior distributions of random variables associated with the hyperparameters can be determined (e.g., calculated). In this case, the hyperparameters can be kept unchanged. According to various embodiments, the values of the hyperparameters can be optimized based on the known measurement points. According to various embodiments, at least one first model and a second model can be trained or learned in this manner (see, e.g., Algorithms 1, 2, 4, and 5). According to various embodiments, the second model can be learned based on the first model, and the previously optimized hyperparameters of the first model can be kept unchanged during the training or learning of the second model.
[0044] When adapting a learned or trained posterior model to new measurement points, hyperparameter optimization or tuning can also be performed as described above.
[0045] According to various embodiments, a known first measurement point D s(2) Using the first posterior model 114, a first posterior model 114 (also referred to as a first model a posteriori) can be determined. The first posterior model 114 can represent a relationship between at least one input variable and at least one output variable of another process. Specifically, even if the input parameter value or the output value is not part of the known first measurement point, the first posterior model 114 can be used to determine an assigned expected output value for the input parameter value, and for the output value, an associated probability density for the input parameter value. Then, for example, the input value with the highest probability density can be selected. Specifically, the first posterior model 114 represents another process.
[0046] For each first measurement point, the user selects at least one first input parameter value 102 to determine the known first measurement point D. s(2) However, the first measurement point D s(2) It is also possible to measure a subset of the known first measurement points D s(2) A temporary first posterior model 113 can be determined for a subset of the known first measurement points D s(2)The first control device 106 may be configured to determine at least one input parameter value 116 (also referred to as new input parameter value) for the new measurement point 132 using an acquisition function (e.g., according to Reference [2] and / or Reference [3]). The first control device 106 may control the first device 108 according to the determined at least one input parameter value 116, and the one or more first sensors 110 may detect an output value 118 belonging to the new measurement point 132. For example, in this way, the first control device 106 may determine the output value 118 for the new measurement point D s(2) In this case, the temporary first posterior model 113 can be adapted based on the determined new measurement points to determine the first posterior model 114. Thus, the first posterior model 114 is adapted based on all known first measurement points D s(2) For a given parameter, a relationship between at least one input variable and at least one output variable can be expressed.
[0047] However, the first measurement point D s(2)Alternatively, the input parameter values 116 may be determined using a simulation. In this case, the system 100 may include a computer. In a first example, the simulation may include determining the input parameter values 116 for the new measurement point 132 and simulating a physical or chemical process. In this case of the first example, the first controller 106, the first device 108, and the one or more first sensors 110 are not required. In a second example, the simulation may include only simulating a physical or chemical process. In this case of the second example, for example, the first controller 106 may also determine the input parameter values 116 for the new measurement point 132, and the first device 108 and the one or more first sensors 110 are not required. The computer may use a memory during processing. As mentioned above, the computer may be any type of circuit, i.e., any type of logic-implementing entity. The computer may be configured to simulate another process and thus determine multiple times at least one associated output value 112 for at least one input parameter value 102. In particular, at a known first measurement point D s(2) The output values of may be simulated rather than measured.
[0048] 1C illustrates control of a first device 108 by a trained first posterior model 114, according to various embodiments. For example, the first posterior model 114 may be used to determine an input parameter value 126 of at least one input variable based on a desired output value 122 of at least one output variable of another process. The first controller 106 may be configured to control the first device 108 according to the determined input parameter value 126. Optionally, one or more first sensors 110 may detect an output value that may substantially correspond to the desired output value 122.
[0049] FIG. 2A shows a second system 200 that includes a second device 208 that can perform a physical or chemical process.
[0050] The second system 200 may be similar to the first system 100. Accordingly, the second system 200 may include a second controller 206 and one or more second sensors 210. The second controller 206 may be configured to control the second device 208 according to each (provided) second input parameter value 202 of at least one (i.e., exactly one or more) input variable. Specifically, the second controller 206 may control the interaction of the second device 208 with the environment according to the one(or more) second input parameter values 202. Because each second input parameter value 202 may be provided for each input variable (of the at least one input variable), this is also referred to hereinafter as at least one second input parameter value 202. The second device 208 may be configured to perform a physical or chemical process according to the at least one second input parameter value 202.
[0051] The one or more second sensors 210 may be configured to detect the results of a physical or chemical process. The results of a physical or chemical process may be, for example, a characteristic of a manufactured product or a characteristic of a processed workpiece (e.g., hardness, strength, density, microstructure, macrostructure, chemical composition, etc.), the success or failure of a robot's skill (e.g., picking up an object), the resolution of an image captured by a camera, etc., as described with respect to FIG. 1A . The results of a process may be represented using at least one (i.e., exactly one or more) output variable. For example, the results of another process may be represented using multiple first output variables, and the results of a physical or chemical process may be represented using multiple second output variables. In this case, the multiple first output variables may have one or more (e.g., all) output variables of the second output variables, or vice versa. For example, one or more output variables of the multiple first output variables may match one or more output variables of the multiple second output variables. The result may be determined during or following a physical or chemical process, as described in connection with one or more sensors 110.
[0052] The one or more second sensors 210 may be configured to detect a second output value 212 of at least one output variable. For example, a respective second output value 212 for each output variable of the plurality of output variables may be detected. The at least one second input parameter value 202 and the at least one second output value 212 may form a second measurement point 230. According to various embodiments, the plurality of second measurement points may be referred to as known second measurement points D. t It can be calculated as:
[0053] Known second measurement point D t can be expressed by equation (2).
number
[0054] Similar to the first system 100, the second system 200 is configured to measure a known second measurement point D t The device may include a storage device having at least one memory for storing:
[0055] Both the first posterior model 114 and the second posterior model 214 may have uncertainties regarding the accuracy of at least one output variable for a particular at least one input variable, and thus the accuracy of each model. These uncertainties are relative to known measurement points. Each model's prediction may include both an expected value and a measure of uncertainty regarding this expected value. That is, each model's prediction may include a distribution of values (e.g., a normal distribution) for each output value. Specifically, the posterior model may represent the relationship between the input and output variables for known measurement points, where the area between the known measurement points is also predicted. It should be understood that these predicted areas h have some uncertainty, and that the greater the number of measurement points, the less uncertainty there may be. Such uncertainty may be significant, for example, when only a few measurement points are present. It should also be understood that each measurement point may have uncertainty due to noise and / or disturbances.
[0056] According to various embodiments, the first posterior model 114 and the known second measurement point D t to determine the second posterior model 214. According to various embodiments, when determining the second posterior model 214, the known first measurement point D s(2) The uncertainty of the first posterior model 114 about the known second measurement point D tThe uncertainty of the second posterior model 214 with respect to (i.e., input parameter values of one or more input variables) at measurement points where no measured first output value is present may be crucial. In the following, two concepts (also referred to as approaches) are described that can consider the uncertainty of the two posterior models 114, 214. According to various embodiments, the system 200 may include a computer configured to determine the second posterior model 214.
[0057] I) First concept for finding the second posterior model 214 (also called sequential hierarchical Gaussian process): The first posterior model 114 is a first posterior probability distribution p(f s(2) |D s(2) ) may be included.
[0058] The computer calculates a first posterior probability distribution p(f s(2) |D s(2) ) (e.g., p(f s |D s )), which may be configured to determine a Gaussian process based on a first posterior probability distribution p(f s |D s ) (and thus the first posterior model 114) to the function
number
number
number
number
[0059] The computer can calculate the average of multiple Gaussian processes (for example, using Bayesian Model Averaging).
number
number
number
number
number
[0060] The computer calculates the second known measurement point D t Second advance model f t The second posterior model 214 may be configured to determine a second posterior probability distribution p(f t |D t ) may be included.
[0061] In general, the prior model of a model may be considered as a prior belief before adjusting the model to known measurement points. The prior belief of the prior model can be expressed as a Gaussian process gp(m,k) with expectation function m(·) and covariance function (also called kernel function) k(·,·). By adjusting the prior model to known measurement points, the posterior model can be obtained. The posterior model is calculated by the following equation: * The posterior probability distribution may be a Gaussian process with a posterior expectation function and a posterior covariance function at x, y ...* The posterior expectation function of the Gaussian process at t A second prior model f t For the adjustment of , the posterior covariance function is shown in equation (3) and the corresponding posterior covariance function is shown in equation (4).
number
number
number
number
[0062] The covariance function of the second posterior model 214 is calculated by multiplying the covariance function k s (also called the first kernel function) and a common covariance function k t (also referred to as a second kernel function or common covariance function of multiple Gaussian processes). The common covariance function of the first model and the second model can be expressed according to Equation (5).
number
[0063] During Bayesian optimization, hyperparameters of kernel functions (e.g., the first kernel function and / or the second kernel function) can be adapted (also referred to as hyperparameter optimization). For example, the hyperparameters of the first prior model can be adapted (e.g., optimized) at a known first measurement point D s The first prior model may be adapted (e.g., optimized) to maximize the probability that the second prior model is represented by the first prior model. For example, the hyperparameters of the second prior model may be adapted (e.g., optimized) to maximize the probability that the second prior model is represented by the first prior model. t The first kernel function k can be calculated so that the probability that k is represented by this second prior model is high (e.g., maximized). In this case, the hyperparameters of the first prior model can be kept constant. s By fixing the hyperparameters of , the uncertainty of the first posterior model 114 can be maintained as the second prior model is adjusted to the known second measurement points.
[0064] According to various embodiments, a known first measurement point D s, the known second measurement point D t , the first kernel function k s and the second kernel function k t Based on this, the second posterior model 214 can be determined according to Algorithm 1, incorporating the first posterior model 114.
[0065] Algorithm 1: Learning the second posterior model 214 according to the first concept Input: D s , D t , k s , k t Output: p(f t |D t ) 1. (If the first measurement point is known and pre-normalized, e.g.
number
number
[0066] Here, as mentioned above, the second prior model f t is the average value of multiple Gaussian processes.
number
[0067] As mentioned above, the first kernel function k s The hyperparameters of (i) and (ii) can be kept unchanged. Specifically, for new measurement points, the second posterior model 214 can be learned by Bayesian optimization by executing only step 4 of Algorithm 1.
[0068] The transfer of uncertainties of the first posterior model 114 to the second prior model and the second posterior model (also referred to as uncertainty communication / propagation) is both data-efficient and optimization-efficient because it reduces the number of measurement data required, thereby reducing the cost (e.g., time cost and / or financial expenditure) of training the model.
[0069] II) A second concept for finding the second posterior model 214 (also called an enhanced hierarchical Gaussian process): Similar to the first concept, the second concept involves multiple Gaussian processes. common covariance function k t By
number
number
number
number
number
number
[0070] The common kernel function of the two models is the first kernel function k of the first posterior model. s and the second kernel function k t Based on this, it can be expressed according to equation (6).
number
number
number
number
number
number
[0071] The second posterior model 214 can be trained according to Algorithm 2.
[0072] Algorithm 2: Learning a second posterior model according to the second concept Input: D s , D t , k s , k t Output: p(f t |D t ) 1. (If the first measurement point is known and pre-normalized, e.g.
number
number
[0073] 2B illustrates control of a second device 208 by a trained second posterior model 214, according to various embodiments. The second controller 206 may be provided with a desired output value 222 of at least one output variable. Optionally, the controller 206 may determine the desired output value based on other parameter values and / or adjustments. The controller 206 may be configured to determine an input parameter value 226 of at least one input variable, which is assigned to the desired output value 222 according to the second posterior model 214. The controller 206 may control the second device 208 according to the determined input parameter value of the at least one input variable to perform a physical or chemical process. Optionally, one or more second sensors 210 may detect an output value that may substantially correspond to the desired output value 222.
[0074] FIG. 2C illustrates adapting the second posterior model 214 according to various embodiments. According to various embodiments, adapting the second posterior model 214 can be performed using Bayesian optimization, where additional second measurement points can be determined using Bayesian optimization, and the second posterior model 214 can be adjusted to the additional second measurement points. The new second measurement points can be adjusted based on the second posterior model 214 using an acquisition function α(f t |x,D t ) (for example, the acquisition function according to [2] and / or [3]), for example, according to equation (8).
number
number
number
[0075] For example, the control device 206 can use an acquisition function to select input parameter values 216 for the new second measurement point 232. The control device 206 can control the second device 208 according to the determined input parameter values 216, and one or more second sensors 210 can detect output values 218 associated with the new second measurement point 232. According to various embodiments, the second posterior model 214 can additionally be adjusted for the new second measurement point 232. The adapted second posterior model 234 can then adapt the relationship between at least one input variable and at least one output variable to the previously known second measurement point D. t and the new second measurement point 232. Specifically, the known second measurement point D t may similarly include new second measurement points 232. In this way, new second measurement points may be determined as frequently as desired, and at each iteration, each second posterior model may be adjusted to the now newly known second measurement points. Again, in the first and second concepts, the first kernel function k sThe hyperparameters of the second posterior model can be kept unchanged. This can increase the efficiency of the Bayesian optimization. It should be appreciated that it is not necessary to explicitly determine multiple Gaussian processes, but rather an effective formula that represents the limit of an infinitely large number of such Gaussian processes can be evaluated. This effective formula can be evaluated, for example, when the hyperparameters of the second posterior model are optimized again based on new measurement points.
[0076] The determination of the new second measurement points and the adaptation of the second posterior model can be described according to Algorithm 3.
[0077] Algorithm 3: Bayesian optimization of the second posterior model 214 for i←1,2,…do: 1. Optimization of the acquisition function via the second posterior model x i =argmax x α(f t |x,D t ) to calculate the input parameter value x i Find out. 2. Input parameter value x i The output value y belongs to i =y i (x i ) is found. 3. Known second measurement point D t Then, a new measurement point (x i ,y i )
number
[0078] 2D illustrates the control of a second device 208 by an adapted second posterior model 234, according to various embodiments. This can be done similarly to the control of a second device 208 described in connection with FIG. 2B. A second controller 206 can be provided with a desired output value 222 of at least one output variable, and the controller 206 can determine input parameter values 246 of at least one input variable, which are assigned according to the adapted second posterior model 234. The controller 206 can control the second device 208 according to the determined input parameter values 246 of the at least one input variable to perform a physical or chemical process.
[0079] According to various embodiments, the second posterior model can be trained sequentially based on multiple other models. Specifically, before using the first posterior model to train the second posterior model, the first posterior model can be trained based on yet another posterior model (hereinafter referred to as a third posterior model), which may also have been trained based on yet another posterior model.
[0080] For example, the third posterior model can represent a relationship between at least one input variable and at least one output variable of yet another process related to a physical or chemical process. The first posterior model can be determined incorporating the third posterior model according to Algorithm 1 or Algorithm 2, as described above for incorporating the first posterior model to determine the second posterior model. The third posterior model can be determined by incorporating the third known measurement point D s(1) Alternatively, the third posterior model can be determined based on a known third measurement point D s(1) As described herein, known measurement points, such as the known third measurement point, may also be determined using simulation.
[0081] As specific examples of such transfer learning, a known third measurement point can be determined from a simulation, and a third posterior model can represent this relationship; a known first measurement point can be measured in a laboratory in an experimental setup similar to the production system, and a first posterior model can represent this relationship; a known second measurement point can be measured in the production system, and a second posterior model can represent the relationship between the input and output variables of the production system.
[0082] According to various embodiments, a plurality of other n s The second posterior model 214 may be trained sequentially for each of the n models, where each separate model may be derived incorporating each preceding separate model as described herein for training the second posterior model incorporating the first posterior model. Specifically, the second posterior model 214 may be trained for n s It can be viewed as a stack of Gaussian processes layered on top of each other.
[0083] Regarding the first concept, the common kernel function for all models can be expressed according to equation (9).
number
[0084] The second posterior model 214 can be determined (eg, trained) using Algorithm 4 according to the first concept.
[0085] Algorithm 4: Learning the second posterior model 214 according to the first concept Input: Each known measurement point (D s1 ,D s2 ,…,D t ) and each kernel function (k s1 ,k s2 ,…,k t ) Output: p(f t |D t ) 1. (Known measurement point D s1 If is pre-normalized, e.g.,
number
number
number
[0086] Regarding the second concept, all other n s The covariance function of the second posterior model 214 can be expanded to have the covariance of the models. In this case, the common kernel function for all models can be expressed according to equation (10):
number
[0087] Specifically, the expanded covariance function is s Optionally, additional covariance terms may be added to predict each other model except the first other model. This allows for a model of size (2n s -1)*(2n s -1) kernel function matrix is obtained. In this case, the expression in Eq. (6)
number
number
[0088] A second posterior model 214 can be determined (e.g., trained) using Algorithm 5 according to a second concept.
[0089] Algorithm 5: Learning a second posterior model according to the second concept Input: Each known measurement point (D s1 ,D s2,…,D t ) and each kernel function (k s1 ,k s2 ,…,k t ) Output: p(f t |D t ) 1. (Known measurement point D s1 If is pre-normalized, e.g.,
number
number
number
[0090] Other n s The training of the second posterior model 214 based on a model (e.g., by Algorithm 1, Algorithm 2, Algorithm 3, or Algorithm 4) described herein reduces the overall computational complexity for determining the second posterior model 214 compared to traditional Bayesian kernel methods that consider model uncertainty. The overall computational complexity can be divided into (i) the computational complexity for training another model, (ii) the computational complexity for training a posterior model (e.g., the second posterior model 214) assuming all other models have been trained, and (iii) the computational complexity for predicting using the trained model. Table 1 compares the computational complexities (i), (ii), and (iii) of the traditional Bayesian kernel method, the hierarchical Gaussian process according to Reference [1], the sequential hierarchical Gaussian process according to the first concept, and the enhanced hierarchical Gaussian process according to the second concept. In Table 1, N s is v=1,2,…,n s Assuming that
number
[0091] [Table 1]
[0092] The other models can be trained at the beginning of the Bayesian optimization and do not need to be trained again during the Bayesian optimization. Therefore, the computational complexity (i) of training the other models has a relatively small impact on the overall computational complexity. Training the posterior model may be performed once in each iteration of the Bayesian optimization, thereby adapting the posterior model with respect to newly detected measurement points in each iteration (e.g., see 4. in Algorithm 3). As shown in Table 1, training the posterior model by the traditional Bayesian kernel method requires a large number of iterations for the total number of known measurement points N t +N s is scaled cubed (i.e., by a power of 3) by , and therefore this can be a limiting factor in Bayesian optimization. Typically, the number of known measurement points N of the target model to be trained is t is the number of all other n s The number of known measurement points of the model, N s (i.e., N t < <N s In such a general case, training a posterior model with the first concept and training a posterior model with the second concept increases computational complexity from a cubic scaling to N scan be reduced to a squared relationship in . Furthermore, since only the hyperparameters of the models to be trained need to be optimized when training the posterior model using the first concept and the posterior model using the second concept, fewer hyperparameters are optimized compared to the traditional Bayesian kernel method in which all models are optimized jointly, thereby further reducing the computational complexity. The prediction of the trained model may be performed multiple times in each iteration of Bayesian optimization during the optimization of the acquisition function (e.g., see 1. in Algorithm 3). Therefore, the computational complexity (iii) for the prediction of the trained model is also relevant to the overall computational complexity. As shown in Table 1, the computational complexity (iii) of the first concept and the second concept is comparable to the computational complexity (iii) of the traditional Bayesian kernel method. Reference [1] reduces this computational complexity (iii), but Reference [1] does not provide a solution for other n s It does not provide a well-founded method for considering the uncertainty of a particular model. Essentially, it is not possible to map the covariance between individual measurements of other models. Within the framework of a statistical model, it is not possible to consider the uncertainty of other n s By properly considering the uncertainty of individual models, the efficiency of Bayesian optimization can be improved.
[0093] For illustrative purposes, Fig. 3A exemplarily visualizes a hierarchical Gaussian process 316 according to Reference [1], a sequential hierarchical Gaussian process 320 according to a first concept, and an enhanced hierarchical Gaussian process 318 according to a second concept, with β=1. Fig. 3A shows a source function 308 modeled by an exemplary first posterior model and an objective function 314 modeled by an exemplary second posterior model, which respectively represent the relationship between input parameter values 302 and output values 304. The first posterior model representing the source function 308 is modeled by a known first measurement point D s , 306, or the first posterior model representing the source function 308 may be determined based on the known first measurement point D s, 306. For the objective function 314 represented by the second posterior model, the known second measurement point D t , 312. Illustration a) shows the first known measurement point D s , 306 does not exist, the first posterior model has a relatively high uncertainty 310. Illustration (b) shows that the posterior model trained using the hierarchical Gaussian process 316 according to Reference [1] with β=1 underestimates the uncertainty in this right-hand region. In contrast, illustrations (b) and (c) show that both the posterior model trained using the sequential hierarchical Gaussian process 320 according to the first concept and the posterior model trained using the enhanced hierarchical Gaussian process 318 according to the second concept take the uncertainty in this right-hand region into account.
[0094] Figure 3B shows a comparison between a traditional Bayesian kernel method 360 (without transfer learning), a traditional Gaussian process 316 according to reference [1] with β=1, a sequential hierarchical Gaussian process 320 according to the first concept, and an enhanced hierarchical Gaussian process 318 according to the second concept. s =1), illustratively with simple regret 334 as a function of optimization iteration 336 for a 1-dimensional Forrester function 338, a 2-dimensional Brain function 340, a 3-dimensional Hartmann3 function 342, and a 6-dimensional Hartmann6 function 344. Illustration 352 illustrates Bayesian optimization for multiple other models, illustratively with three other models (n s = 3), including the Hartmann 6 function 354, five other models (n s= 5), with simple regret 334 as a function of optimization iteration 336. The smaller the simple regret value, the better the optimization found, and the fewer optimization iterations 336 where small simple regret values are found, the more efficient the Bayesian optimization. Specifically, FIG. 3B shows that the improvement in the efficiency of Bayesian optimization according to the first and second concepts, compared to other methods, increases with the number of dimensions (see illustration 332, e.g., 6-dimensional Hartmann6 function 344). Multiple other models (n s >1), significant efficiency gains are also shown (see illustration 352, e.g., n s = 5 (see Alpine function 356).
[0095] Specifically, the transfer learning concept described herein, in which all the above-mentioned model uncertainties are propagated (i.e., uncertainty propagation), facilitates Bayesian optimization. For example, when only a few measurement points exist or can be detected to obtain a model (e.g., at a relatively high cost (e.g., time cost, energy consumption, financial expenditure, etc.)), the transfer learning described herein can significantly improve learning efficiency based on uncertainty propagation. Conventional Bayesian optimization methods either increase computational costs by fully considering uncertainty, or cause defects in the trained model by not considering uncertainty at all, and may result in significant inaccuracy of the model when only a few measurement points exist. The concept described herein can achieve both consideration of uncertainty and relatively high learning efficiency. The transfer learning described herein provides relatively high data efficiency and optimization efficiency by propagating uncertainty.
[0096] FIG. 4 illustrates a flowchart of a method 400 for controlling a physical or chemical process, according to various embodiments.
[0097] The method 400 may include determining (at 402) a second posterior model by incorporating the first posterior model. The second posterior model may represent a relationship between at least one input variable and at least one output variable of a physical or chemical process. The first posterior model may represent a relationship between at least one input variable and at least one output variable of a process related to the physical or chemical process.
[0098] Incorporating the first posterior model to determine a second posterior model (at 402) may include determining a plurality of Gaussian processes with a common covariance function (at 402A), where a function may be derived from the first posterior model to determine each Gaussian process of the plurality of Gaussian processes by forming an expectation of each Gaussian process.
[0099] According to various embodiments, the second posterior model can be determined using at least one of two concepts (also referred to as approaches): According to the first concept, multiple Gaussian processes can be first averaged with a common covariance function and then adjusted to known measurement points; According to the second concept, multiple Gaussian processes can each be adjusted to known measurement points and then the adjusted Gaussian processes can be averaged.
[0100] According to the first concept, incorporating (at 402) the first posterior model to determine a second posterior model may further include determining a second prior model as an average of multiple Gaussian processes and determining the second posterior model using an adjustment of the prior model to the known measurement point (at 402B). For example, incorporating (at 402) the first posterior model to determine the second posterior model may include adjusting the first posterior model to the known second measurement point D toptimizing the hyperparameters of the second posterior model so that the probability that is represented by the second prior model is increased (e.g., maximized); and using the optimized hyperparameters to calculate the second measurement point D t Second advance model f t and adjusting the
[0101] According to various embodiments, during training or learning (e.g., hyperparameter optimization) of the second model, the hyperparameters of the covariance function of the first model may be kept unchanged.
[0102] According to the second concept, determining a second posterior model incorporating the first posterior model (at 402) may further include adjusting each Gaussian process of the plurality of Gaussian processes to the known measurement points and determining the second posterior model as an average of the adjusted Gaussian processes (at 402B). The second concept also has the advantage that the hyperparameters of the first model do not need to be newly optimized when training the second model.
[0103] The method 400 may further include controlling a physical or chemical process using the second a posteriori model (at 404).
Claims
1. A method (400) for controlling a physical process or a chemical process, comprising: Each measurement point of the known measurement points has an input parameter value of at least one input variable of the physical process or the chemical process, and an output value of at least one output variable of the physical process or the chemical process assigned to the input parameter value; The first posterior model represents the relationship between at least one input variable and at least one output variable of another process related to the physical process or the chemical process; The method comprises: - Incorporating the first posterior model to obtain a second posterior model (402), the second posterior model representing the relationship between the at least one input variable and the at least one output variable of the physical process or the chemical process, and obtaining the second posterior model (402) by incorporating the first posterior model comprises: - Deriving a function from the first posterior model to obtain each Gaussian process by forming the expected value of the Gaussian process, and obtaining a plurality of Gaussian processes using a common covariance function (402A); - Obtaining a prior model as the average value of the plurality of Gaussian processes, and optimizing the hyperparameters of the model by using the known measurement points so that at least one hyperparameter of the covariance function of the first posterior model remains unchanged, thereby obtaining the second posterior model and adjusting the model according to the known measurement points (402B), or - Adjusting each Gaussian process of the plurality of Gaussian processes according to the known measurement points, and obtaining the second posterior model as the average value of the adjusted plurality of Gaussian processes (402B); Including; The method further comprises: - Controlling the physical process or the chemical process using the second posterior model (404). Method (400).
2. The method further comprises: - Based on the second posterior model, using an acquisition function to select a new input parameter value of the at least one input variable of the physical process or the chemical process. - Measuring the output value assigned to the new input parameter value of the at least one output variable, wherein the selected new input parameter value and the measured new output value form a new measurement point, - Adaptively modifying the second posterior model using the new measurement point, wherein the adaptively modified second posterior model represents the physical or chemical process for the known measurement points and the new measurement points, - Controlling the physical or chemical process using the adaptively modified second posterior model, The method (400) according to claim 1.
3. The third posterior model represents the relationship between at least one input variable and at least one output variable of yet another process related to the physical or chemical process, and each distinct measurement point of the known other measurement points includes an input parameter value of at least one input variable of the other process and an output value assigned to the input parameter value of at least one output variable of the other process. The method includes - Obtaining the first posterior model by incorporating the third posterior model, and obtaining the first posterior model by incorporating the third posterior model includes - Deriving a function from the third posterior model to obtain each distinct Gaussian process by forming an expected value of the Gaussian process, and obtaining a plurality of other Gaussian processes using another common covariance function; and - Obtaining another prior model as an average value of the plurality of other Gaussian processes, and obtaining the first posterior model by adjusting the other prior model according to the known other measurement points, or - Adjusting each distinct Gaussian process of the plurality of other Gaussian processes according to the known other measurement points, and obtaining the first posterior model as an average value of the adjusted plurality of other Gaussian processes, including The method (400) according to claim 1.
4. Controlling the physical or chemical process using the second posterior model (404) includes - Determining an input parameter value of the at least one input variable assigned to a desired output value of the at least one output variable of the physical or chemical process according to the second posterior model, - Controlling the physical process or the chemical process according to the obtained input parameter values; comprising; The method (400) according to claim 1.
5. An apparatus (208) configured to implement the method (400) according to any one of claims 1 to 4.
6. - An apparatus (208) configured to execute a physical process or a chemical process; - Using a second post-model (214, 234) obtained according to the method (400) according to any one of claims 1 to 4, for a desired output value (222) of the at least one output variable of the physical process or the chemical process, determining input parameter values (226, 246) assigned to the desired output value (222), and a control device (206) configured to control the apparatus (208) to execute the physical process or the chemical process according to the determined input parameter values (226, 246); A system (200) comprising.
7. The physical process or the chemical process is - Machining of a workpiece, - Adjustment of a device, - Manufacture of a product, or - Movement of a robotic arm is, The system (200) according to claim 6.
8. A computer program comprising instructions for causing a processor to implement the method (400) according to any one of claims 1 to 4 when executed by the processor.
9. A computer-readable medium storing instructions for causing a processor to implement the method (400) according to any one of claims 1 to 4 when executed by the processor.