Parameter adjustment device and parameter adjustment method

The parameter adjustment device addresses the issue of simulator-actual device behavior discrepancies by using a difference learning unit and prediction unit to determine optimal parameters, ensuring high-evaluation values are found efficiently in the actual device.

WO2026083605A1PCT designated stage Publication Date: 2026-04-23MITSUBISHI ELECTRIC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MITSUBISHI ELECTRIC CORP
Filing Date
2025-01-27
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for adjusting parameters in devices using simulators fail to account for significant differences between simulator and actual device behaviors, leading to the potential exclusion of high-evaluation parameter values in the actual device.

Method used

A parameter adjustment device and method that includes a difference learning unit to predict evaluation value differences, a simulator learning unit to predict simulated evaluation values, and a prediction unit to determine the next adjustment parameter based on these predictions, thereby improving efficiency by accounting for simulator-actual device discrepancies.

Benefits of technology

Enables the identification of high-evaluation parameter values in the actual device even when simulator and actual device behaviors differ significantly, enhancing parameter adjustment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025002375_23042026_PF_FP_ABST
    Figure JP2025002375_23042026_PF_FP_ABST
Patent Text Reader

Abstract

This parameter adjustment device (1A) is characterized by comprising: a difference learning unit (17) that predicts, from the value of an adjustment parameter to be tried by a device to be adjusted, an evaluation value difference that is the difference between an actual evaluation value obtained by evaluating the operation of the device to be adjusted in an actual device (5) and a simulated evaluation value obtained by evaluating the operation of the device to be adjusted in a simulator (6); a simulator learning unit (18) that predicts the simulated evaluation value from the value of the adjustment parameter to be tried; a prediction unit (19) that predicts the actual evaluation value for the value of the adjustment parameter to be tried on the basis of the prediction result of the evaluation value difference and the prediction result of the simulated evaluation value; and an actual device trial point determination unit (20) that determines the value of the adjustment parameter to be tried next on the basis of the prediction result of the actual evaluation value.
Need to check novelty before this filing date? Find Prior Art

Description

Parameter Adjustment Device and Parameter Adjustment Method

[0001] The present disclosure relates to a parameter adjustment device and a parameter adjustment method for adjusting adjustment parameters of a device to be adjusted using a simulator.

[0002] The parameter adjustment device is a device that automatically searches for adjustment parameters that increase the evaluation for the device to be adjusted. As a parameter target device, there is a known device that improves the parameter adjustment efficiency for the actual device by utilizing a simulator of the device to be adjusted. Here, the adjustment efficiency refers to the height of the evaluation with respect to the number of trials performed on the actual device.

[0003] For example, in Patent Document 1, adjustment of adjustment parameters for a simulator is performed in advance, and the adjustment efficiency is improved by narrowing the parameter adjustment range in the actual device based on the values of the adjustment parameters with high evaluation obtained at that time.

[0004] Japanese Patent Application Laid-Open No. 2009-122779

[0005] However, since there is a difference in behavior between the simulator and the actual device, the value of the adjustment parameter with high evaluation in the simulator may not necessarily result in high evaluation in the actual device. When narrowing the adjustment range in the actual device based on the value of the adjustment parameter with high evaluation in the simulator as in the method disclosed in Patent Document 1, there is a possibility that the parameter value that can obtain high evaluation in the actual device may not be included in the adjustment range. Therefore, when the difference in behavior between the simulator and the actual device is large, there is a problem that it may not be possible to find the value of the adjustment parameter with high evaluation.

[0006] The present disclosure has been made in view of the above, and an object thereof is to obtain a parameter adjustment device that can find the value of an adjustment parameter with high evaluation even when the difference in behavior between the simulator and the actual device is large while improving the adjustment efficiency by utilizing the simulator.

[0007] To solve the above-mentioned problems and achieve the objective, the parameter adjustment device according to this disclosure is characterized by comprising: a difference learning unit that predicts the difference between an actual evaluation value obtained by evaluating the operation of the device to be adjusted on an actual device and a simulated evaluation value obtained by evaluating the operation of the device to be adjusted on a simulator, based on the value of the adjustment parameter to be adjusted for testing of the device to be adjusted; a simulator learning unit that predicts a simulated evaluation value from the value of the adjustment parameter to be adjusted for testing; a prediction unit that predicts the actual evaluation value for the value of the adjustment parameter to be adjusted for testing based on the prediction result of the difference in evaluation value and the prediction result of the simulated evaluation value; and an actual device trial point determination unit that determines the value of the adjustment parameter to be tested next based on the prediction result of the actual evaluation value.

[0008] According to this disclosure, by utilizing a simulator, it is possible to obtain a parameter adjustment device that can improve adjustment efficiency while also finding highly-rated adjustment parameter values ​​even when there is a large difference in behavior between the simulator and the actual device.

[0009] Figure 1 shows an example of the configuration of the parameter adjustment system according to Embodiment 1. Figure 3 shows an example of the functional configuration of the difference learning unit. Figure 4 shows an example of a layered neural network. Figure 5 shows an example of the functional configuration of the simulator learning unit. Figure 6 is a flowchart for explaining the parameter adjustment method according to Embodiment 1. Figure 7 shows an example of the configuration of the parameter adjustment system according to Embodiment 2. Figure 8 is a flowchart for explaining the parameter adjustment method according to Embodiment 2. Figure 9 shows an example of the configuration of the parameter adjustment system according to Embodiment 3. Figure 1 shows dedicated hardware for realizing the functions of the parameter adjustment device according to Embodiments 1 to 3. Figure 1 shows the configuration of the control circuit for realizing the functions of the parameter adjustment device according to Embodiments 1 to 3.

[0010] The parameter adjustment device and parameter adjustment method according to the embodiments of this disclosure will be described in detail below with reference to the drawings.

[0011] Embodiment 1. Figure 1 shows an example of the configuration of a parameter adjustment system 100A according to Embodiment 1. The parameter adjustment system 100A includes a parameter adjustment device 1A, a real device 5, and a simulator 6.

[0012] The parameter adjustment device 1A has the function of evaluating the value of the adjustment parameter based on the actual device operation results obtained from the actual device 5 of the device to be adjusted and the simulator operation results obtained from the simulator 6 of the device to be adjusted, and adjusting the value of the adjustment parameter based on the evaluation results. The device to be adjusted can be any device that has adjustment parameters, and the parameter adjustment system 100A can be used to adjust the parameters of a device that has adjustment parameters. The parameter adjustment system 100A can be used, for example, to adjust the control parameters of a refrigeration cycle or to adjust parameters in motor positioning control. When the parameter adjustment system 100A is used to adjust the control parameters of a refrigeration cycle, the device to be adjusted is an air conditioner, a refrigerator, etc., and the adjustment parameters are, for example, actuator operating amounts for a compressor, expansion valve, fan, etc. When the parameter adjustment system 100A is used to adjust parameters in motor positioning control, the device to be adjusted is a servo motor, and the adjustment parameters are, for example, parameters that control the command waveform shape.

[0013] The parameter adjustment device 1A includes a real device acquisition unit 11, a simulator acquisition unit 12, a simulation data acquisition unit 13, an evaluation value calculation unit 14, a real device storage unit 15, a simulator storage unit 16, a difference learning unit 17, a simulator learning unit 18, a prediction unit 19, and a real device trial point determination unit 20.

[0014] The actual device acquisition unit 11 has the function of acquiring the actual device operation results, which are the operating results when the actual device 5 of the device to be adjusted is operated. The actual device operation results include at least the adjustment parameters input to the actual device 5. In addition to the values ​​of the adjustment parameters, the actual device operation results may also include information indicating the operating conditions of the actual device 5 and the operating status of the actual device 5. Examples of operating conditions for the actual device 5 include information indicating the configuration of the actual device 5. Examples of information indicating the operating status of the actual device 5 include measured values ​​obtained using sensors installed inside or outside the actual device 5, and values ​​calculated inside the actual device 5. The actual device acquisition unit 11 outputs the acquired actual device operation results to the evaluation value calculation unit 14.

[0015] The simulator acquisition unit 12 has the function of acquiring simulator operation results, which are the results of operating the simulator 6 of the device to be adjusted. The simulator operation results include at least the values ​​of the adjustment parameters input to the simulator 6. In addition to the values ​​of the adjustment parameters, the simulator operation results may also include information indicating the operating conditions of the simulator 6 and the operating status of the simulator 6. Examples of operating conditions for the simulator 6 include information indicating the configuration of the simulator 6. Examples of information indicating the operating status of the simulator 6 include values ​​calculated internally by the simulator 6. The simulator acquisition unit 12 outputs the acquired simulator operation results to the evaluation value calculation unit 14.

[0016] The simulation data acquisition unit 13 has the function of acquiring simulation data from outside the parameter adjustment device 1A. The simulation data includes the results of multiple simulator operations. The simulation data can be of any type, but for example, it may be data acquired when adjustment parameters were adjusted for the simulator 6 in advance. The simulation data acquisition unit 13 outputs the acquired simulation data to the evaluation value calculation unit 14.

[0017] The evaluation value calculation unit 14 has the function of calculating an evaluation value based on the actual device operation results output by the actual device acquisition unit 11, the simulator operation results output by the simulator acquisition unit 12, or the simulator operation results included in the simulation data output by the simulation data acquisition unit 13. The evaluation value can be such that, for example, a larger value indicates more desirable performance. In the following, the evaluation value will be such that a larger value indicates more desirable performance, but if an evaluation value is used where a smaller value indicates more desirable performance, the method described below can be applied by multiplying the evaluation value by "-1".

[0018] The evaluation value is a value that evaluates the operation of the device being adjusted. For example, when applied to the adjustment of control parameters of a refrigeration cycle, the evaluation value may be a value indicating capacity such as heating capacity, cooling capacity, or refrigeration capacity, or it may be the energy consumption efficiency (COP: Coefficient of Performance), or a combination of these. Also, when applied to the adjustment of parameters in motor positioning control, the evaluation value may be the time it takes for the controlled object, which is moved by the drive of the servo motor that is the device being adjusted, to reach the target travel distance, or it may be the magnitude of residual vibration, or a combination of these.

[0019] The adjustment parameters and evaluation values ​​are vectors containing one or more variables. In the following description, the letter 'm' represents the adjustment parameter and the letter 'n' represents the evaluation value.

[0020] The actual device storage unit 15 has the function of storing actual evaluation values, which are evaluation values ​​that evaluate the operation of the actual device 5 and are calculated based on the actual device operation results, in association with the values ​​of the corresponding adjustment parameters. The data stored in the actual device storage unit 15 is Data D r This can be expressed by the following formula (1).

[0021]

[0022] Here, m ri is the value of the i-th stored adjustment parameter, and n ri is the i-th stored actual evaluation value, where i is a non-negative integer.

[0023] The simulator storage unit 16 has a function of storing, in association with the value of an adjustment parameter, a simulation evaluation value which is an evaluation value obtained by evaluating the operation in the simulator 6 calculated based on the simulator operation result. The simulator storage unit 16 can store, in association with the value of an adjustment parameter, a simulation evaluation value calculated based on the simulator operation result output from the simulator acquisition unit 12 and a simulation evaluation value calculated based on the simulator operation result included in the simulation data output from the simulation data acquisition unit 13. The data stored in the simulator storage unit 16 is data D s and is represented by the following mathematical formula (2).

[0024]

[0025] Here, m si is the value of the adjustment parameter stored at the i-th position, and n si is the simulation evaluation value stored at the i-th position.

[0026] The differential learning unit 17 has a function of learning the relationship between the adjustment parameter and the evaluation value difference using the data stored in the actual device storage unit 15 and the data stored in the simulator storage unit 16, and a function of predicting the evaluation value difference from the input value of the adjustment parameter. Further, the differential learning unit 17 may further have a function of calculating the uncertainty with respect to the prediction of the evaluation value difference. The evaluation value difference is the result of subtracting the simulation evaluation value from the actual evaluation value for a certain value of the adjustment parameter. When the actual evaluation value is n r and the simulation evaluation value is n s , the evaluation value difference δ is represented by the following mathematical formula (3).

[0027]

[0028] To calculate the evaluation value difference δ, both the actual evaluation value and the simulation evaluation value at the same value of the adjustment parameter are required. Therefore, the data actually used for learning is only the data in which both data D r and data D s exist for the same value of the adjustment parameter among the data stored in the actual device storage unit 15 and the simulator storage unit 16.

[0029] There are no particular restrictions on the learning method used by the differential learning unit 17. For example, the differential learning unit 17 can use commonly known machine learning methods such as Gaussian process regression, linear regression, k-nearest neighbors, neural networks, random forests, and gradient boosting trees. The differential learning unit 17 may also use a combination of multiple methods from the above. Depending on the method used by the differential learning unit 17, prediction uncertainty may not be calculated. In this case, for example, the prediction uncertainty can be calculated by raising the distance from the value of the tuning parameter to be predicted to the k-th closest data point in the training data to the power of d. Here, k is an arbitrary positive integer and d is a real number greater than 0.

[0030] Here, as an example, we will describe the case where the differential learning unit 17 uses a neural network. Figure 2 is a diagram showing an example of the functional configuration of the differential learning unit 17. The differential learning unit 17 includes a training data acquisition unit 171, a differential model generation unit 172, a differential model storage unit 173, a prediction data acquisition unit 174, and a prediction unit 175.

[0031] The learning data acquisition unit 171 acquires data D stored in the actual device storage unit 15. r and data D stored in the simulator memory unit 16 s From there, training data is generated and acquired. Specifically, the training data acquisition unit 171 acquires data D r and Data D s From the combinations of adjustment parameters and evaluation values ​​included, those corresponding to the same adjustment parameter value are extracted, and the adjustment parameter and actual evaluation value n are selected. r , simulated evaluation value n s The associated values ​​are acquired as training data. Alternatively, the training data acquisition unit 171 acquires the extracted actual evaluation values ​​n r and simulated evaluation value n s The evaluation value difference δ can be calculated using the relationship in formula (3), and the extracted adjustment parameters and the calculated evaluation value difference δ can be associated and obtained as training data. The training data acquisition unit 171 outputs the acquired training data to the difference model generation unit 172.

[0032] The difference model generation unit 172 learns the evaluation value difference δ based on the training data output from the training data acquisition unit 171. In other words, the difference model generation unit 172 generates a difference model, which is a trained model for predicting the evaluation value difference δ from the values ​​of the adjustment parameters of the device to be adjusted.

[0033] The difference model generation unit 172 learns the evaluation value difference δ by so-called supervised learning, for example, according to a neural network model. Here, supervised learning is a method in which a learning device is given a data set of input and result labels, learns the features in that training data, and infers the result from the input.

[0034] A neural network consists of an input layer made up of multiple neurons, a hidden layer made up of multiple neurons, and an output layer made up of multiple neurons. The hidden layer is also called a hidden layer. There may be one hidden layer or two or more hidden layers.

[0035] Figure 3 shows an example of a three-layer neural network. For example, in a three-layer neural network like the one shown in Figure 3, when multiple inputs are input to the input layer (X1-X3), these values ​​are multiplied by weights W1 (w11-w16) and input to the hidden layer (Y1-Y2), and the result is further multiplied by weights W2 (w21-w26) and output from the output layer (Z1-Z3). This output result changes depending on the values ​​of weights W1 and W2.

[0036] In this disclosure, the neural network learns the evaluation value difference δ by so-called supervised learning, according to the training data acquired by the training data acquisition unit 171.

[0037] In other words, a neural network learns by inputting the values ​​of tuning parameters into the input layer and adjusting the weights W1 and W2 so that the result output from the output layer approaches the evaluation value difference δ.

[0038] The difference model generation unit 172 generates a difference model, which is a trained model, by performing the learning described above, and outputs it to the difference model storage unit 173.

[0039] The difference model storage unit 173 stores the trained model output from the difference model generation unit 172 as a difference model.

[0040] The prediction data acquisition unit 174 acquires the values ​​of the adjustment parameters.

[0041] The prediction unit 175 predicts the evaluation value difference δ obtained using the trained difference model. That is, the prediction unit 175 can output the evaluation value difference δ inferred from the values ​​of the adjustment parameters by inputting the values ​​of the adjustment parameters obtained by the prediction data acquisition unit 174 into the trained difference model. The prediction unit 175 can output the prediction uncertainty of the evaluation value difference δ together with the evaluation value difference δ. As described above, depending on the method used by the difference learning unit 17, the prediction unit 175 may also have a function to calculate the prediction uncertainty.

[0042] In this embodiment, the prediction unit 175 is described as outputting the evaluation value difference δ using the difference model learned by the difference model generation unit 172 of the difference learning unit 17 of the parameter adjustment device 1A, but the system is not limited to this example. The prediction unit 175 may also acquire a learned difference model from an external source, such as another parameter adjustment device, and output the evaluation value difference δ based on this difference model.

[0043] Furthermore, the difference model generation unit 172 may learn the evaluation value difference δ according to the training data created for multiple target devices. The difference model generation unit 172 may acquire training data from multiple target devices used in the same area, or it may learn the evaluation value difference δ using training data collected from multiple target devices operating independently in different areas. It is also possible to add or remove target devices from which training data is collected. Moreover, a training device that has learned the evaluation value difference δ for one target device may be applied to another target device, and the evaluation value difference δ for that other target device may be retrained and updated.

[0044] In this example, the parameter adjustment device 1A is assumed to include all functions, including the learning data acquisition unit 171 and the difference model generation unit 172 corresponding to the functions of the learning device, the difference model storage unit 173, and the prediction data acquisition unit 174 and the prediction unit 175 corresponding to the functions of the prediction device. However, the example is not limited to this. Some of the functions shown in Figure 2 may be implemented by a device separate from the parameter adjustment device 1A. For example, some of the functions of the learning device and the prediction device of the parameter adjustment device 1A may reside on a cloud server.

[0045] Returning to the explanation of Figure 1, the simulator learning unit 18 processes the data D stored in the simulator storage unit 16. s The simulator learning unit 18 has the function of generating a simulated evaluation value model, which is a trained model for predicting simulated evaluation values ​​that evaluate the operation in the simulator 6 from the values ​​of the adjustment parameters, and the function of predicting simulated evaluation values ​​for the input adjustment parameters. The simulator learning unit 18 may further have a function for calculating the prediction uncertainty of the predicted simulated evaluation values. The learning method and the method for calculating the prediction uncertainty are the same as those of the differential learning unit 17.

[0046] Figure 4 shows an example of the functional configuration of the simulator learning unit 18. Here, we will explain an example using a neural network, similar to the differential learning unit 17. The simulator learning unit 18 includes a learning data acquisition unit 181, a simulated evaluation value model generation unit 182, a simulated evaluation value model storage unit 183, a prediction data acquisition unit 184, and a prediction unit 185.

[0047] The learning data acquisition unit 181 acquires data D stored in the simulator storage unit 16. s From this, training data is generated and acquired. Specifically, the training data acquisition unit 181 acquires data D s The combinations of adjustment parameters and simulated evaluation values ​​included are extracted in a corresponding manner, and the adjustment parameters and simulated evaluation values ​​n s The data obtained by associating these elements is acquired as training data. The training data acquisition unit 181 outputs the acquired training data to the simulated evaluation value model generation unit 182.

[0048] The simulated evaluation value model generation unit 182 generates a simulated evaluation value n based on the training data output from the training data acquisition unit 181. s The system learns the following: In other words, the simulated evaluation value model generation unit 182 generates a simulated evaluation value n from the values ​​of the adjustment parameters of the device to be adjusted. s Generate a simulated evaluation model, which is a pre-trained model for predicting the outcome.

[0049] The simulated evaluation value model generation unit 182 generates simulated evaluation values ​​n by so-called supervised learning, for example, according to a neural network model. s This involves learning. Here, supervised learning is a method in which a learning device is given pairs of data consisting of inputs and resulting labels, learns the features in that training data, and infers the results from the inputs.

[0050] The neural network used by the simulated evaluation value model generation unit 182 may be the same as the one described with reference to Figure 3.

[0051] In this disclosure, the neural network uses so-called supervised learning to obtain a simulated evaluation value n according to the training data acquired by the training data acquisition unit 181. s Learn about it.

[0052] In other words, the neural network takes the value of the tuning parameter into the input layer, and the result output from the output layer is the simulated evaluation value n. s The learning process involves adjusting weights W1 and W2 to approximate the target value.

[0053] The simulated evaluation value model generation unit 182 generates a simulated evaluation value model, which is a trained model, by performing the learning described above, and outputs it to the simulated evaluation value model storage unit 183.

[0054] The simulated evaluation value model storage unit 183 stores the trained model output from the simulated evaluation value model generation unit 182 as a simulated evaluation value model.

[0055] The prediction data acquisition unit 184 acquires the values ​​of the adjustment parameters to be tested.

[0056] The prediction unit 185 obtains a simulated evaluation value n using the trained model, which is a simulated evaluation value model.s The prediction unit 185 inputs the values ​​of the adjustment parameters acquired by the prediction data acquisition unit 184 into the trained simulated evaluation value model, thereby predicting the simulated evaluation value n inferred from the values ​​of the adjustment parameters. s The prediction unit 185 can output a simulated evaluation value n. s The prediction uncertainty is simulated using the evaluation value n. s It can be output together with the following. Similar to the case of the differential learning unit 17, depending on the method used by the simulator learning unit 18, the prediction unit 185 may have a function to calculate prediction uncertainty.

[0057] Returning to the explanation of Figure 1, the prediction unit 19 receives the prediction result of the evaluation value difference δ output by the difference learning unit 17 and the simulated evaluation value n output by the simulator learning unit 18. s Based on the prediction results, the actual evaluation value n for the adjustment parameter value r The prediction unit 19 predicts the actual evaluation value n for the value of the adjustment parameter under trial. r It has the function of predicting the actual evaluation value n. r It may also have a function to calculate the prediction uncertainty.

[0058] μ is the predicted value of the evaluation value difference δ. δ , simulated evaluation value n s The predicted value of μ s In this case, the actual evaluation value n r Predicted value μ r This is expressed by the following formula (4). In other words, the prediction unit 19 predicts the value μ of the evaluation value difference δ. δ and simulated evaluation value n s Predicted value μ s The sum of these two values ​​gives the actual evaluation value n. r Predicted value μ r It can be done this way.

[0059]

[0060] The prediction unit 19 uses the actual evaluation value n r Predicted value μ r As a prediction uncertainty for this, the predicted value μ is the difference in evaluation values ​​δ. δ The prediction uncertainty may be used as is. Alternatively, the prediction unit 19 uses the actual evaluation value n rPredicted value μ r As a prediction uncertainty for this, the predicted value μ is the difference in evaluation values ​​δ. δ Prediction uncertainty and simulated evaluation value n s Predicted value μ s Alternatively, the prediction uncertainty and the prediction uncertainty can be raised to the power of p, the sum of these values, and the result of raising the sum to the power of p (1) may be used. The predicted value μ of the evaluation difference δ. δ The prediction uncertainty σ δ Let the simulated evaluation value n s Predicted value μ s The prediction uncertainty σ s In this case, the actual evaluation value n r Predicted value μ r Predictive uncertainty σ r This can be expressed by the following formula (5).

[0061]

[0062] The prediction unit 19 uses the actual evaluation value n r Predicted value μ r And the actual evaluation value n r Predicted value μ r Predictive uncertainty σ r The results are output to the actual device trial point determination unit 20.

[0063] The actual device trial point determination unit 20 uses the actual evaluation value n output by the prediction unit 19. r Predicted value μ r And the actual evaluation value n r Predicted value μ r Predictive uncertainty σ rBased on this, the system has the function of determining the value of the next adjustment parameter to be tested and transmitting it to the actual device 5 and the simulator 6. The method by which the actual device trial point determination unit 20 determines the value of the adjustment parameter should be such that it increases the adjustment efficiency, that is, it should be such that the value of the adjustment parameter that yields a high evaluation value in the actual device 5 can be found at an early stage in the future. For example, the actual device trial point determination unit 20 can determine the value of the next adjustment parameter to be tested by maximizing an acquisition function commonly used in Bayesian optimization. Examples of acquisition functions include PI (Probability of Improvement), EI (Expected Improvement), UCB (Upper Confidence Bound), PES (Predictive Entropy Search), and MES (Max-value Entropy Search). If the vector length of the evaluation value is 2 or more, the actual device trial point determination unit 20 may use, for example, EHVI (Expected Hyper Volume Improvement), which is an acquisition function for multi-objective optimization. The actual device trial point determination unit 20 may use any method to maximize the acquisition function. For example, the actual device trial point determination unit 20 can use methods such as random search, CMA-ES (Covariance Matrix Adaptation Evolution Strategy), L-BFGS (Limited-memory Broyden-Fletcher-Goldfarb-Shanno), and sequential quadratic programming.

[0064] Figure 5 is a flowchart illustrating the parameter adjustment method according to Embodiment 1. First, the parameter adjustment device 1A acquires simulation data in the simulation data acquisition unit 13, and then calculates a pseudo-evaluation value for each of the multiple simulator operation results included in the simulation data in the evaluation value calculation unit 14 (step S101). The calculated pseudo-evaluation values ​​are stored in the simulator storage unit 16 in association with the values ​​of the corresponding adjustment parameters.

[0065] The parameter adjustment device 1A transmits the initial value of the adjustment parameter to be tested to the actual device 5 and the simulator 6 (step S102). The initial value may be determined using random numbers, or it may be determined by receiving a command from outside the parameter adjustment device 1A. Alternatively, the initial value may be determined based on data stored in the simulator storage unit 16. For example, the initial value may be the value of the adjustment parameter with the highest evaluation value among the data stored in the simulator storage unit 16.

[0066] Next, using the values ​​of the adjustment parameters set for the trial, the operation of the actual device 5 and the simulator 6 is started (step S103).

[0067] After the simulator 6 has finished running, the parameter adjustment device 1A acquires the simulator operation results in the simulator acquisition unit 12 and calculates simulated evaluation values ​​for the adjustment parameters to be tested from the simulator operation results in the evaluation value calculation unit 14 (step S104). The calculated simulated evaluation values ​​are stored in the simulator storage unit 16 in association with the adjustment parameters.

[0068] The simulator learning unit 18 learns the relationship between the adjustment parameters and the simulated evaluation values ​​using the data stored in the simulator storage unit 16 (step S105). The simulated evaluation value model generated as a result of the learning is stored in the simulated evaluation value model storage unit 183.

[0069] After the operation of the actual device 5 is completed, the parameter adjustment device 1A acquires the operation results of the actual device in the actual device acquisition unit 11 and calculates the actual evaluation value from the operation results in the evaluation value calculation unit 14 (step S106). The calculated actual evaluation value is stored in the actual device storage unit 15 in association with the adjustment parameters.

[0070] The parameter adjustment device 1A, in the difference learning unit 17, uses the data stored in the actual device storage unit 15 and the data stored in the simulator storage unit 16 to learn the relationship between the adjustment parameters and the difference in evaluation values ​​(step S107). The difference model generated as a result of the learning is stored in the difference model storage unit 173.

[0071] The parameter adjustment device 1A determines the value of the adjustment parameter to be tested next in the actual device trial point determination unit 20, and transmits the determined value of the adjustment parameter to the actual device 5 and the simulator 6 (step S108).

[0072] After step S108, the parameter adjustment device 1A repeats its operation from step S103 using the value of the adjustment parameter to be tried next.

[0073] Note that Figure 5 does not show the conditions for terminating the operations from step S103 to step S108. For example, the processes from step S103 to step S108 may be repeated until the user performs a termination operation. Alternatively, termination conditions may be set for the evaluation value, and after the processing of step S108 is completed, before returning to the processing of step S103, it may be determined whether or not the termination conditions are met. If the termination conditions are met, the process in Figure 5 may be terminated.

[0074] As described above, Embodiment 1 provides a parameter adjustment device 1A characterized by comprising: a difference learning unit 17 that predicts the difference between the actual evaluation value obtained by evaluating the operation of the device to be adjusted on the actual device 5 and the simulated evaluation value obtained by evaluating the operation of the device to be adjusted on the simulator 6, based on the value of the adjustment parameter to be adjusted for testing of the device to be adjusted; a simulator learning unit 18 that predicts the simulated evaluation value from the value of the adjustment parameter to be adjusted for testing; a prediction unit 19 that predicts the actual evaluation value for the value of the adjustment parameter to be adjusted for testing based on the prediction result of the difference in evaluation value and the prediction result of the simulated evaluation value; and an actual device trial point determination unit 20 that determines the value of the adjustment parameter to be tested next based on the prediction result of the actual evaluation value.

[0075] The parameter adjustment device 1A can improve its adjustment efficiency by utilizing the simulator 6. The parameter adjustment device 1A also predicts simulated evaluation values ​​that evaluate the operation in the simulator 6, and evaluation value differences that show the difference between the behavior of the actual device 5 and the behavior of the simulator 6. Then, the parameter adjustment device 1A predicts the actual evaluation values ​​from the predicted results of the simulated evaluation values ​​and the predicted results of the evaluation value differences, and determines the value of the adjustment parameter to be tried next based on the predicted result of the actual evaluation values. In this way, it becomes possible to find a highly evaluated adjustment parameter value even when there is a large difference in the behavior between the simulator 6 and the actual device 5.

[0076] Furthermore, the parameter adjustment device 1A includes a simulation data acquisition unit 13 that acquires simulation data containing multiple simulator operation results that correlate the values ​​of the adjustment parameters and simulated evaluation values ​​from past simulations performed on the simulator 6 of the device to be adjusted. The simulator learning unit 18 has the function of generating a simulated evaluation value model, which is a trained model for predicting simulated evaluation values ​​from the adjustment parameters, using the simulation data acquired by the simulation data acquisition unit 13. Since the parameter adjustment device 1A performs learning separately in the differential learning unit 17 and the simulator learning unit 18, the simulation data acquired by the simulation data acquisition unit 13 can be utilized during learning, and the prediction accuracy of the actual evaluation values ​​that evaluate the operation of the actual device 5 can be improved. This results in the effect of improving adjustment efficiency.

[0077] Furthermore, the prediction unit 19 can use the sum of the predicted value of the evaluation value difference by the difference learning unit 17 and the predicted value of the simulated evaluation value by the simulator learning unit 18 as the predicted value of the actual evaluation value. The prediction unit 19 can further use the predicted uncertainty of the evaluation value difference by the difference learning unit 17 and the predicted uncertainty of the simulated evaluation value by the simulator learning unit 18 as the predicted uncertainty of the actual evaluation value, after adding the sum of the two values ​​raised to the power of p and raising the sum to the power of p, or using the predicted uncertainty of the evaluation value difference by the difference learning unit 17 as the predicted uncertainty of the actual evaluation value. Here, p is a real number greater than 0.

[0078] Furthermore, Embodiment 1 provides a parameter adjustment method. The parameter adjustment method is characterized by including the steps of: predicting the evaluation value difference, which is the difference between the actual evaluation value obtained by evaluating the operation of the device to be adjusted on the actual device 5 and the simulated evaluation value obtained by evaluating the operation of the device to be adjusted on the simulator 6, based on the value of the adjustment parameter to be adjusted for testing; predicting the simulated evaluation value from the value of the adjustment parameter to be tested; predicting the actual evaluation value for the value of the adjustment parameter to be tested based on the prediction result of the evaluation value difference and the prediction result of the simulated evaluation value; and determining the value of the adjustment parameter to be tested next based on the prediction result of the actual evaluation value.

[0079] Embodiment 2. In Embodiment 1, the simulator 6 was operated only at the same time as the actual device 5. Generally, the simulator 6 completes its operation in a shorter time than the actual device 5, so in many cases, the simulator 6 can be operated multiple times while the actual device 5 is operated once. Therefore, Embodiment 2 shows an example in which the simulator 6 is operated multiple times while the actual device 5 is operated once, thereby improving the prediction accuracy of the simulator learning unit 18 and the prediction unit 19 and achieving high adjustment efficiency.

[0080] Figure 6 shows an example of the configuration of the parameter adjustment system 100B according to Embodiment 2. The parameter adjustment system 100B includes a parameter adjustment device 1B, a real device 5, and a simulator 6.

[0081] The parameter adjustment device 1B has the same configuration as the parameter adjustment device 1A, plus a simulator trial point determination unit 21. The following description will mainly focus on the differences from Embodiment 1, and will omit the description of parts that are the same as Embodiment 1.

[0082] The simulator trial point determination unit 21 determines the value of the adjustment parameter to be tested next in the simulator 6, based on the predicted value and prediction uncertainty of the actual evaluation value obtained from the prediction unit 19 and the prediction uncertainty regarding the simulated evaluation value obtained from the simulator learning unit 18. The simulator trial point determination unit 21 transmits the determined value of the adjustment parameter only to the simulator 6 and not to the actual device 5.

[0083] The simulator trial point determination unit 21 should determine the value of the adjustment parameter to be the next trial target in a manner that increases the adjustment efficiency in the actual device 5 and selects an adjustment parameter value that has high predictive uncertainty regarding the simulated evaluation value.

[0084] For example, the simulator trial point determination unit 21 can maximize the acquisition function UCB by using the prediction uncertainty of the simulated evaluation value instead of the prediction uncertainty of the actual evaluation value. The acquisition function UCB for the predicted value μ of the evaluation value and the prediction uncertainty σ of the evaluation value is a UCB When (μ, σ) is used, the acquisition function a with respect to the adjustment parameter x is s (x) is expressed by the following formula (6).

[0085]

[0086] As another example, the predicted value μ of the actual evaluation value is shown in formula (7) below. r , prediction uncertainty of actual evaluation value σ r A general acquisition function a(μ) based on this r , σ r ) Prediction uncertainty σ of simulated evaluation value s The result of adding this may be used as a new acquisition function. The simulator trial point determination unit 21 uses the acquisition function a shown in formula (7) s By maximizing (x), we can determine the value of the next adjustment parameter to try.

[0087]

[0088] Note that β is a real number greater than 0. As another example, the simulator trial point determination unit 21 may extract values ​​for multiple adjustment parameters using random samples, and among those whose predicted values ​​for the actual evaluation value fall in the top q%, the one with the highest prediction uncertainty for the simulated evaluation value may be selected as the value of the adjustment parameter to be tested next.

[0089] Figure 7 is a flowchart illustrating the parameter adjustment method according to Embodiment 2. The operations from step S101 to step S105 are the same as those described using Figure 5 for the parameter adjustment method according to Embodiment 1, so their explanation is omitted here.

[0090] In step S105, after learning the relationship between the adjustment parameters and the simulated evaluation values, the parameter adjustment device 1B determines in the prediction unit 19 whether or not to adjust the actual device 5 (step S201). The decision of whether or not to adjust the actual device 5 may be based, for example, on whether or not the operation of the actual device 5 has finished, or on whether or not the number of trials in the simulator 6 alone has exceeded a predetermined number.

[0091] When adjusting the actual device 5 (step S201: Yes), the parameter adjustment device 1B proceeds to step S106. The processes from step S106 to step S108 are the same as the parameter adjustment method according to Embodiment 1 described with reference to Figure 5, so their explanation is omitted here.

[0092] If the actual device 5 is not adjusted (step S201: No), the prediction unit 19 instructs the simulator trial point determination unit 21 to determine the value of the adjustment parameter to be tested next. The simulator trial point determination unit 21 determines the value of the adjustment parameter to be tested next and transmits it to the simulator 6 (step S202).

[0093] Next, the simulator 6 is started using the values ​​of the adjustment parameters set for the trial (step S203). After that, the parameter adjustment device 1B proceeds to the process in step S104.

[0094] As described above, Embodiment 2 provides a parameter adjustment device 1B that, in addition to the functions of the parameter adjustment device 1A according to Embodiment 1, includes a simulator trial point determination unit 21 that determines the value of the adjustment parameter for the next trial target by the simulator 6 based on the predicted result of the actual evaluation value and the predicted result of the simulated evaluation value. The parameter adjustment device 1B makes it possible to run the simulator 6 multiple times while the actual device 5 is running once, using the value of the adjustment parameter for the trial target determined by the simulator trial point determination unit 21. This makes it possible to improve the prediction accuracy of the simulator learning unit 18 and improve the prediction accuracy of the actual evaluation value.

[0095] Furthermore, the simulator trial point determination unit 21 determines the value of the adjustment parameter for the next trial target based on the level of predictive uncertainty of the simulated evaluation value. The simulator trial point determination unit 21 selects trial points that reduce the predictive uncertainty of the simulated evaluation value in the simulator 6 at adjustment parameter values ​​that are likely to increase the adjustment efficiency in the actual device 5, thereby achieving a significant improvement in adjustment efficiency.

[0096] Embodiment 3. Figure 8 shows an example of the configuration of the parameter adjustment system 100C according to Embodiment 3. The parameter adjustment system 100C includes a parameter adjustment device 1C, a real device 5, and a simulator 6.

[0097] The parameter adjustment device 1C has a display unit 22 in addition to the configuration of the parameter adjustment device 1B according to Embodiment 2. Below, we will mainly describe the parts that differ from Embodiments 1 and 2, and omit the description of parts that are the same as Embodiments 1 and 2.

[0098] The display unit 22 has the function of displaying information indicating the internal state of the parameter adjustment device 1C in a way that can be seen by the user, based on information obtained from the differential learning unit 17, the simulator learning unit 18, the prediction unit 19, the actual device trial point determination unit 20, and the simulator trial point determination unit 21. The information displayed by the display unit 22 may include, for example, at least one of the learning status of the differential learning unit 17, the learning status of the simulator learning unit 18, the prediction status of the prediction unit 19, the adjustment status of the actual device trial point determination unit 20, and the adjustment status of the simulator trial point determination unit 21.

[0099] Information indicating the adjustment status may include the value of the adjustment parameter determined as the next trial target, and at least one of the predicted value and prediction uncertainty for the entire or partial adjustment range, relating to the actual evaluation value, simulated evaluation value, and evaluation value difference. If an acquisition function is used, the information indicating the adjustment status may also include the value of the acquisition function for the entire or partial adjustment range.

[0100] Information indicating the prediction status may include the prediction accuracy for all or part of the data used for training. In addition to information indicating the prediction status, information indicating the internal state of the trained model may also include information indicating the internal state of the trained model. Internal state generally includes what are called hyperparameters of the machine learning model, Shapley values, permutation importance and other indicators of the importance of each feature, and PDP (Partial Dependence Plot) and ICE (Individual Conditional Expectation) which indicate the change in predicted values ​​in response to changes in features. When using Gaussian process regression, kernel hyperparameters such as length scale may also be included as internal state.

[0101] As described above, Embodiment 3 provides a parameter adjustment device 1C having a display unit 22 that displays at least one of the following: a differential learning status including the prediction accuracy of the differential learning unit 17 predicting the difference in evaluation values ​​and information indicating the internal state of the differential model; a simulated learning status including the prediction accuracy of the simulator learning unit 18 predicting simulated evaluation values ​​and information indicating the internal state of the simulated evaluation value model; the prediction accuracy of the prediction unit 19; the adjustment status of the actual device trial point determination unit 20; and the adjustment status of the simulator trial point determination unit 21. The parameter adjustment device 1C has a function to provide the user with the status of machine learning and parameter adjustment, so that the user can grasp the status of parameter adjustment and changes in the status. This makes it possible for the user to review the parameter adjustment conditions, and further improvement in parameter adjustment efficiency can be expected. The parameter adjustment conditions referred to here include the method for calculating evaluation values, the parameter adjustment range, the configuration of the actual device 5, the configuration of the simulator 6, and environmental conditions.

[0102] Here, the hardware configuration of the parameter adjustment devices 1A, 1B, and 1C will be described. The functions of each part of the parameter adjustment devices 1A, 1B, and 1C are realized by processing circuits. These processing circuits may be realized by dedicated hardware, or they may be control circuits using a CPU (Central Processing Unit).

[0103] When the above processing circuits are implemented using dedicated hardware, they are implemented by the processing circuit 90 shown in Figure 9. Figure 9 is a diagram showing dedicated hardware for realizing the functions of the parameter adjustment devices 1A, 1B, and 1C according to Embodiments 1 to 3. The processing circuit 90 may be a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof.

[0104] When the above processing circuit is implemented using a control circuit with a CPU, this control circuit is, for example, a control circuit 91 with the configuration shown in Figure 10. Figure 10 is a diagram showing the configuration of a control circuit 91 for realizing the functions of parameter adjustment devices 1A, 1B, and 1C according to embodiments 1 to 3. As shown in Figure 10, the control circuit 91 includes a processor 92 and a memory 93. The processor 92 is a CPU, also called a central processing unit, processing unit, arithmetic unit, microprocessor, microcomputer, DSP (Digital Signal Processor), etc. The memory 93 is, for example, a non-volatile or volatile semiconductor memory such as RAM (Random Access Memory), ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable ROM), EEPROM (Registered Trademark) (Electrically EPROM), magnetic disk, flexible disk, optical disk, compact disk, minidisc, DVD (Digital Versatile Disk), etc.

[0105] When the above processing circuit is implemented by the control circuit 91, it is implemented by the processor 92 reading and executing a program corresponding to the processing of each component stored in the memory 93. The memory 93 is also used as temporary memory for each process executed by the processor 92. The program executed by the processor 92 may be provided in a state stored on a storage medium, or it may be provided via a communication channel such as the internet.

[0106] The configurations shown in the embodiments described above are merely examples of the content of this disclosure and can be combined with other known technologies, and parts of the configuration can be omitted or modified without departing from the gist of this disclosure.

[0107] 1A, 1B, 1C Parameter adjustment device, 5 Actual device, 6 Simulator, 11 Actual device acquisition unit, 12 Simulator acquisition unit, 13 Simulation data acquisition unit, 14 Evaluation value calculation unit, 15 Actual device storage unit, 16 Simulator storage unit, 17 Difference learning unit, 18 Simulator learning unit, 19, 175, 185 Prediction unit, 20 Actual device trial point determination unit, 21 Simulator trial point determination unit, 22 Display unit, 90 Processing circuit, 91 Control circuit, 92 Processor, 93 Memory, 100A, 100B, 100C Parameter adjustment system, 171, 181 Learning data acquisition unit, 172 Difference model generation unit, 173 Difference model storage unit, 174, 184 Prediction data acquisition unit, 182 Simulated evaluation value model generation unit, 183 Simulated evaluation value model storage unit, a, a s (x) Acquisition function, D r , D s data, n r Actual evaluation value, n s Simulated evaluation value, x adjustment parameter, δ evaluation value difference, μ, μ r , μ s , μ δ Predicted value, σ, σ r , σ s Predictive uncertainty.

Claims

1. A parameter adjustment device comprising: a difference learning unit that predicts an evaluation value difference, which is the difference between an actual evaluation value obtained by evaluating the operation of the device to be adjusted on an actual device and a simulated evaluation value obtained by evaluating the operation of the device to be adjusted on a simulator, based on the value of the adjustment parameter to be adjusted on a trial; a simulator learning unit that predicts the simulated evaluation value from the value of the adjustment parameter to be adjusted on a trial; a prediction unit that predicts the actual evaluation value for the value of the adjustment parameter to be adjusted on a trial based on the prediction result of the evaluation value difference and the prediction result of the simulated evaluation value; and an actual device trial point determination unit that determines the value of the adjustment parameter to be tried next based on the prediction result of the actual evaluation value.

2. The parameter adjustment device according to claim 1, further comprising: a simulation data acquisition unit that acquires a plurality of simulation data including simulator operation results that associate the values ​​of the adjustment parameters with the simulated evaluation values ​​when a simulation was previously performed in the simulator of the device to be adjusted, wherein the simulator learning unit has a function of generating a simulated evaluation value model, which is a trained model for predicting the simulated evaluation values ​​from the adjustment parameters, using the simulation data acquired by the simulation data acquisition unit.

3. The parameter adjustment device according to claim 1 or 2, comprising: a simulator trial point determination unit that determines the value of the adjustment parameter for the next trial target by the simulator based on the predicted result of the actual evaluation value and the predicted result of the simulated evaluation value, wherein while the actual device is running once, the simulator is run multiple times using the value of the adjustment parameter for the trial target determined by the simulator trial point determination unit.

4. The parameter adjustment device according to claim 3, comprising a display unit for displaying the adjustment status of the simulator trial point determination unit, wherein the adjustment status includes the value of the adjustment parameter for the next trial target, the actual evaluation value, the simulated evaluation value, the predicted value of the difference between the evaluation values, and the prediction uncertainty.

5. The parameter adjustment device according to claim 3 or 4, characterized in that the simulator trial point determination unit determines the value of the adjustment parameter for the next trial target based on the level of predictive uncertainty of the simulated evaluation value.

6. The parameter adjustment device according to any one of claims 1 to 5, wherein the difference learning unit has a function to generate a difference model which is a trained model for predicting the difference in evaluation values ​​from the adjustment parameters, the simulator learning unit has a function to generate a simulated evaluation value model which is a trained model for predicting the simulated evaluation values ​​from the adjustment parameters, and the display unit displays at least one of the following: a difference learning status which includes information indicating the prediction accuracy of the difference learning unit for predicting the difference in evaluation values ​​and the internal state of the difference model, a simulated learning status which includes information indicating the prediction accuracy of the simulator learning unit for predicting the simulated evaluation values ​​and the internal state of the simulated evaluation value model, the prediction accuracy of the prediction unit, and the adjustment status of the actual device trial point determination unit, and the adjustment status which includes at least one of the following: the value of the adjustment parameter for the next trial target, and at least one of the actual evaluation value, the simulated evaluation value, the predicted value of the difference in evaluation values, and the prediction uncertainty.

7. The parameter adjustment device according to any one of claims 1 to 6, characterized in that the prediction unit takes the predicted value of the evaluation value difference by the difference learning unit and the predicted value of the simulated evaluation value by the simulator learning unit as the predicted value of the actual evaluation value, and adds the predicted uncertainty of the evaluation value difference by the difference learning unit and the predicted uncertainty of the simulated evaluation value by the simulator learning unit to the power of p, which is a real number greater than 0, and then takes the result of raising the result to the power of 1 / p, or takes the predicted uncertainty of the evaluation value difference by the difference learning unit as the predicted uncertainty of the actual evaluation value.

8. A parameter adjustment method characterized by comprising: predicting an evaluation value difference, which is the difference between an actual evaluation value obtained by evaluating the operation of the device to be adjusted on an actual device and a simulated evaluation value obtained by evaluating the operation of the device to be adjusted on a simulator, based on the value of the adjustment parameter to be adjusted for testing of the device to be adjusted; predicting the simulated evaluation value from the value of the adjustment parameter to be adjusted for testing; predicting the actual evaluation value for the value of the adjustment parameter to be adjusted for testing based on the prediction result of the evaluation value difference and the prediction result of the simulated evaluation value; and determining the value of the adjustment parameter to be tested next based on the prediction result of the actual evaluation value.

Citation Information

Patent Citations

  • Device and program for searching optimal solution by evolutionary method and controller of control object by evolutionary method

    JP2002245434A

  • Model prediction controller and model prediction control method

    JP2006146764A

  • Automatic adaptation system for control parameter

    JP2008217155A

  • Operation condition determination device for plant, control system for plant, operation condition determination method, and program

    JP2020181296A