Information processing apparatus, control system, control variable determination method, and control variable determination program
By using predictive distribution calculations from an information processing device and robust Gaussian process regression algorithms, the problem of automatic control for non-constant waste properties was solved, achieving efficient and reliable waste handling control for cranes.
Patent Information
- Application Number
- CN202080097673.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-28
- Filing Date
- 2020-11-30
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2040-11-30
AI Technical Summary
Existing technologies struggle to determine suitable control variables when automatically controlling waste with inconsistent properties, resulting in poor mixing effects and failure to achieve the desired waste treatment results.
An information processing device is used to determine the optimal control variables through predictive distribution calculation, control variable retrieval and update mechanisms. Bayesian optimization and robust Gaussian process regression algorithms are used to optimize the determination process of control variables, reduce the impact of unreliable data and improve the reliability of control results.
It enables effective automatic control of waste with inconsistent properties, improves the uniformity and efficiency of waste treatment, reduces the number of experiments, and improves the reliability of control results.
Smart Images

Figure CN115175868B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to an information processing apparatus and the like, which can be used for automatic control of a crane for carrying garbage. BACKGROUND
[0002] For garbage carried into a garbage disposal facility, the garbage is temporarily stored in a storage device called a pit, and then is sent to an incinerator to be incinerated. In a general garbage disposal facility, a crane is used to move garbage stored in a pit. The crane is basically operated by an operator by hand, but currently, automatic control is also being tried.
[0003] For example, Patent Literature 1 below discloses a technology of quantifying a degree of stirring of garbage, and automatically controlling a crane based on the quantified degree of stirring of garbage. For the technology described in this document, automatic control is achieved by quantifying a degree of stirring based on a number of stirring, and generating a crane control instruction that specifies a position at which garbage is to be grabbed, and a position at which the grabbed garbage is to be dropped.
[0004] PRIOR ART DOCUMENTS
[0005] PATENT LITERATURE
[0006] Patent Literature 1: Japanese Patent Application Publication No. 2010-275064 SUMMARY
[0007] (1) PROBLEMS TO BE SOLVED BY THE INVENTION
[0008] However, garbage stored in a garbage pit is mixed with various garbage that differs in material and state, and the properties thereof are not constant. Therefore, if the technology of Patent Literature 1 is applied to stirring of garbage in an actual garbage pit, it can be impossible to perform stirring as intended.
[0009] For example, only specifying a position at which garbage is to be grabbed, depending on the properties of garbage at the position, there can be a case where a large amount of garbage is grabbed, and a case where only a small amount of garbage is grabbed. Also, if the amount of garbage grabbed is unstable, during automatic control of a crane, a difference between an actual amount of movement of garbage and an intended amount of movement can cumulatively become large. Therefore, it can eventually be impossible to obtain an intended stirring effect. In addition, even in a case where garbage as intended is grabbed, it is considered that an intended stirring effect cannot be obtained due to a property deviation of grabbed garbage. These problems are not limited to stirring, and also exist in control when any work such as lifting, scattering, and dropping of garbage by a crane.
[0010] In the case of automatically controlling a crane that handles garbage whose properties are not constant, it is necessary to determine a control variable of the crane in order to obtain a desired control result, but the related art has a problem that such a control variable cannot be determined. An object of one embodiment of the present application is to provide an information processing apparatus and the like which can determine a control variable of a crane that handles garbage, the control variable being able to obtain a desired control result.
[0011] (II) Technical Solution
[0012] To solve the above problem, an information processing apparatus according to one embodiment of the present application includes a predictive distribution calculation portion, a control variable search portion, and a control variable determination portion. The predictive distribution calculation portion calculates a predictive distribution of a function representing a relationship between a control variable of a crane that handles garbage and a control result using control result data obtained by associating the control variable with the control result of the crane controlled using the control variable. The control variable search portion searches for a candidate control variable which is a candidate of an optimal value of the control variable, based on the predictive distribution. The predictive distribution calculation portion updates the predictive distribution using the candidate control variable searched for by the control variable search portion and the control result of the crane controlled using the candidate control variable. The control variable determination portion determines the optimal value of the control variable using a function formed based on the updated predictive distribution.
[0013] In addition, to solve the above problem, a control variable determination method according to one embodiment of the present application is performed by one or a plurality of information processing apparatuses. The method includes a predictive distribution calculation step of calculating a predictive distribution of a function representing a relationship between a control variable of a crane that handles garbage and a control result using control result data obtained by associating the control variable with the control result of the crane controlled using the control variable; a control variable search step of searching for a candidate control variable which is a candidate of an optimal value of the control variable, based on the predictive distribution; an update step of updating the predictive distribution using the candidate control variable searched for in the control variable search step and the control result of the crane controlled using the candidate control variable; and a control variable determination step of determining the optimal value of the control variable using a function formed based on the updated predictive distribution.
[0014] (III) Advantageous Effects
[0015] According to one embodiment of the present application, it is possible to determine a control variable which can obtain a desired control result. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a block diagram illustrating an example of the main part structure of the information processing apparatus of Embodiment 1 of the present application.
[0017] Figure 2 is a diagram showing an outline of a control system including the above information processing device.
[0018] Figure 3 is a flowchart showing an example of a process of determining a control variable of a crane.
[0019] Figure 4 is a graph showing a function configured based on a mean and a variance of a prediction distribution calculated using a Gaussian process regression and a function configured based on a mean and a variance of a prediction distribution calculated using a robust Gaussian process regression.
[0020] Figure 5 is a block diagram showing an example of a main part structure of the information processing device of Embodiment 2 of the present application.
[0021] Figure 6 is a flowchart showing an example of a process of optimizing a parameter of a kernel function.
[0022] Figure 7 is a graph showing a result of an experiment for verifying an effect of the above information processing device.
[0023] Figure 8 is a graph showing a task kernel at the time of optimization end in Experiments 10 to 12. DETAILED DESCRIPTION
[0024] (Embodiment 1)
[0025] (System outline)
[0026] Based on Figure 2 An outline of a control system 9 of one embodiment of the present application will be described. Figure 2 is a diagram showing an outline of the control system 9. As shown in the figure, the control system 9 includes an information processing device 1A, a control device 3, and a crane 5. The information processing device 1A is a device that calculates a control variable for controlling the crane 5.
[0027] The control system 9 is a system that controls an action of the crane 5 using the control device 3. The information processing device 1A calculates a control variable that specifies a content of control performed by the control device 3. By the information processing device 1A calculating an appropriate control variable, appropriate automatic control of the crane 5 using the control device 3 is realized.
[0028] The crane 5 is a crane for garbage handling, for example, used in garbage disposal facilities and the like. The crane 5 can have, for example, a grab having a plurality of claws that grab garbage, an opening and closing mechanism that opens and closes the claws of the grab, a lifting mechanism that lifts the grab, and a moving mechanism that moves the grab in a horizontal direction, and the like. At this time, the control device 3 is able to cause the crane 5 to perform garbage stirring and the like by controlling the opening and closing mechanism, the lifting mechanism, and the moving mechanism.
[0029] In the control system 9, when the information processing device 1A is caused to calculate the optimal control variable, first, a task to be performed by the crane 5 is set, and a control variable at the time when the crane 5 performs the task is set.
[0030] For example, after grabbing garbage with the grab, the grab is moved in the horizontal direction and opened and closed, and the garbage is dispersed on the moving path of the grab, and thus the crane 5 performs a task of stirring the garbage. At this time, as long as the control variable is set so that the garbage is uniformly dispersed, that is, as long as the control variable that determines the timing of the opening and closing control of the grab is set.
[0031] In the above case, as long as a series of controls in which, after starting the opening operation of the grab that grabs garbage, the closing operation of the grab is started when a predetermined amount of garbage falls from the grab, and the opening operation is started again after a predetermined time elapses, is repeated, the garbage is dispersed. Therefore, the above predetermined amount and the above predetermined time can be control variables.
[0032] In addition, for example, the weight of the garbage that falls during the period from the start of the opening operation of the grab to the start of the closing operation, the remaining amount or the rate of change of the weight of the garbage in the grab during the period, the length of the period, and the moving distance of the grab during the period, and the like can be control variables. In addition, for example, it can be configured so that the closing operation is automatically started after the opening operation ends, and the above period is set to be the period from the start of the opening operation to the end of the closing operation. Furthermore, the time of the opening operation, the time of the closing operation, and the like can be control variables.
[0033] In addition, data used by the control device 3 when controlling the crane 5 using the above control variables is not particularly limited. For example, in controlling the crane 5, in addition to the weight of the grabbed garbage, information indicating the amount of moisture, the kind, the degree of stirring, the surface state (for example, an image of the surface of the garbage), and the like can be used. The data form of such data is not particularly limited, and can be numerical data, image data, or the like.
[0034] After the control variable is set, the crane 5 is controlled by the control device 3 to perform the set task. Also, the appropriateness of the control result is evaluated, and the evaluation result is input to the information processing device 1A together with the control variable in the control. For example, if the task is to uniformly disperse garbage and perform stirring, the evaluation can be performed in such a manner that the more uniform the amount of garbage dispersed on the movement path of the grab is, the higher the evaluation value is.
[0035] The information processing device 1A performs optimization of the control variable based on the input control variable and evaluation value, and the control device 3 causes the crane 5 to perform the task again using the optimized control variable. By repeating such processing, the information processing device 1A can determine the control variable that can obtain the desired control result. Also, as a result, appropriate automatic control of the crane 5 by the control device 3 can be achieved.
[0036] (Main Part Structure)
[0037] Based on Figure 1 The structure of the information processing device 1A will be described. Figure 1 is a block diagram showing an example of the main part structure of the information processing device 1A. Also, the following describes an example in which the information processing device 1A determines the control variable that can obtain the desired control result using Bayesian optimization (hereinafter referred to as BO), that is, an example in which the control variable is optimized using BO.
[0038] As shown in the drawing, the information processing device 1A includes a control section 10A that comprehensively controls each section of the information processing device 1A, a storage section 20 that stores various data used by the information processing device 1A, an input section 30 that accepts input to the information processing device 1A, and an output section 40 that is used for data output from the information processing device 1A.
[0039] The data acquisition section 101, the predicted distribution calculation section 102, the control variable search section 103, and the control variable determination section 104 are included in the control section 10A. Also, the control result data 201 is stored in the storage section 20.
[0040] The data acquisition section 101 acquires learning data used when optimization is performed using BO. Specifically, since the control variable used in the control of the crane 5 and the evaluation value of the control result thereof are stored as the control result data 201, the data acquisition section 101 acquires the control result data 201 as the learning data.
[0041] In the case where N points of control variables are included in the control result data 201, these control variables are expressed as
[0042] [Numeral 1]
[0043]
[0044] Its evaluation value is expressed as
[0045] [Number 2]
[0046]
[0047] The prediction distribution calculation unit 102 uses the control result data 201 acquired by the data acquisition unit 101 to calculate the prediction distribution of a function representing the relationship between the control variables and the control results. This function will be referred to as the evaluation function f(θ) below. Furthermore, when new data is added to the control result data 201, the prediction distribution calculation unit 102 updates the prediction distribution in a manner that reflects that data.
[0048] If Gaussian noise ε is used n ~N(0, β) assumes the relationship between the control variables and the control results is as follows:
[0049] [Number 3]
[0050] y n =f(θ) n )+ε n
[0051] The following distribution is obtained as the prediction distribution of the evaluation function based on the Gaussian process.
[0052] [Number 4]
[0053]
[0054] [Number 5]
[0055] μ(θ) = k * (K Θ +βI) -1 f(2)
[0056] [Number 6]
[0057]
[0058] Here, k * =k(θ, θ), K θ It utilizes [K] θ ] i,j =k(θ) i θ j The fundamental matrix obtained. Additionally,
[0059] [Number 7]
[0060]
[0061] k Θ,* It is [k]Θ,* ] i =k(θ) i Let α be the vertical vector of θ, and k(·,·) be the kernel function. Here, the parameter of the kernel function is set to α. k .
[0062] The mean function μ(θ) represents the average value of the evaluation function predicted based on the control result data 201. Additionally, the variance function σ(θ) is the variance of the evaluation function predicted based on the control result data 201. σ(θ) represents the unreliability of the prediction, and this value tends to increase in regions where the control result data 201 is insufficient. Furthermore, as described later... Figure 4 The thinner gray portion represents the variance σ. If σ is large, the width of this gray portion widens, indicating that the prediction is unreliable. That is, it can be seen that the control result data required to improve the reliability of the prediction is insufficient. According to equation (3), the kernel function contained in the variance function σ(θ) and the parameter α of the kernel function are... k This affects the calculation of the predicted distribution. When calculating the predicted distribution, the parameter α is... k The optimizations will be discussed in detail later.
[0063] The control variable retrieval unit 103 retrieves candidates for the optimal control variable (candidate control variables) in order to determine the optimal control variable. Specifically, the control variable retrieval unit 103 uses the mean function μ(θ) and the variance function σ(θ) to retrieve the control variable that maximizes the following acquisition function a(θ). The control variable retrieved through this retrieval becomes a candidate for the optimal control variable. This retrieval is based on the UCB (Upper Confidence Bound) strategy. Furthermore, κ in equation (5) is a parameter used to adjust the retrieval and its use. Of course, other methods can also be used to retrieve new control variables. For example, the PI (Probability of Improvement) strategy and the EI (Expected Improvement) strategy can also be used to retrieve candidates for the optimal control variable.
[0064] Furthermore, when determining the control variable that minimizes the value of the evaluation function as the optimal control variable (for example, when the time required to complete the task is taken as the evaluation value for a task that is preferably completed in a short time), it is sufficient to search for the control variable that minimizes the obtained function a(θ).
[0065] [Number 8]
[0066]
[0067] [Number 9]
[0068] a(θ)=μ(θ)+κσ(θ) (5)
[0069] The candidate optimal control variable detected by the control variable retrieval unit 103 is used for the control of the crane 5. Furthermore, the control result (more specifically, the evaluation value of the control result) is obtained and input into the information processing device 1A. The input data (the candidate optimal control variable and the evaluation value) is appended to the control result data 201. Then, the predicted distribution is updated using the control result data 201 with the appended data. Furthermore, the calculation of the evaluation value can be performed by the information processing device 1A, or by other devices or a user.
[0070] The control variable determination unit 104 determines the optimal value of the control variable based on an evaluation function constructed from the updated predicted distribution based on the predicted distribution calculation unit 102. The optimal value of the control variable is the value inferred from the updated predicted distribution, and can also be described as the solution of the optimization calculation of the control variable performed by the information processing device 1A. By setting the value of the control variable when controlling the crane 5 to this optimal value, the best control result can be expected.
[0071] There is no particular limitation on the method for determining the optimal value; various methods can be applied. For example, the control variable determination unit 104 can determine the candidate control variable detected by the control variable retrieval unit 103 as the optimal control variable when the candidate control variable has been evaluated. This is because when the candidate control variable detected by the control variable retrieval unit 103 is evaluated, there is a higher probability that the control variable is not an extreme value of the evaluation function but corresponds to the maximum (or minimum) value.
[0072] As described above, the information processing device 1A includes: a prediction distribution calculation unit 102, which calculates the prediction distribution of the evaluation function using control result data 201; and a control variable retrieval unit 103, which retrieves candidate values for the optimal values of the control variables based on the prediction distribution. Furthermore, the prediction distribution calculation unit 102 updates the prediction distribution using the new candidate control variables detected by the control variable retrieval unit 103 and the control results of controlling the crane 5 using these candidate control variables. The information processing device 1A also includes a control variable determination unit 104, which determines the optimal value of the control variables using a function constructed based on the updated prediction distribution. More specifically, the function is constructed based on the mean and variance of the updated prediction distribution (Equation (5)).
[0073] As explained in the section on "Technical Problems to be Solved," the properties of the waste in the landfill are not constant. Therefore, the relationship between the control variables and control results of crane 5 is difficult to formulate.
[0074] Therefore, based on the above structure, the optimal value of the control variable is determined by the predicted distribution of the function that associates the control variable and the control result. Thus, for a crane handling waste whose handling properties are not constant, it is possible to determine the control variable that yields the desired control result.
[0075] Furthermore, based on the above structure, candidate control variables are retrieved based on the predicted distribution. Therefore, even if the detected candidate control variable is not the optimal control variable, it is still useful data that can be used to appropriately update the predicted distribution. Thus, for example, the number of trials can be reduced to a smaller number compared to the case where trials are repeatedly performed to randomly select control variables for crane 5 and observe the control results of crane 5 in order to determine the optimal control variable.
[0076] (Processing flow)
[0077] based on Figure 3 This describes the process of information processing device 1A determining the control variables of crane 5 (control variable determination method). Figure 3 This is a flowchart illustrating an example of the process for determining the control variables of crane 5.
[0078] In S1, the data acquisition unit 101 reads the control result data 201 stored in the storage unit 20 and sets it as the initial data. At this stage, it is sufficient that the control result data 201 contains control results based on at least one experiment (a control variable and an evaluation value for evaluating the control results using the control variable).
[0079] In S2, the prediction distribution calculation unit 102 optimizes the parameters of the kernel function. As described above, the parameter of the kernel function is α. k There are no particular limitations on the optimization methods; for example, optimization methods that can also be applied to ordinary business logic (BO) can be used.
[0080] In step S3 (prediction distribution calculation step), the prediction distribution calculation unit 102 uses the initial data set in step S1 and the parameters of the kernel function optimized in step S2 to calculate the prediction distribution of the evaluation function that evaluates the control results of the crane 5. As described above, this prediction distribution is represented by equations (1) to (3).
[0081] In step S4 (control variable retrieval step), the control variable retrieval unit 103 retrieves the control variable θ* of the crane 5 that maximizes the function obtained. The control variable θ* is a candidate for the optimal value of the control variable θ. As described above, this process is represented by the above equations (4) and (5).
[0082] In S5, the control variable determination unit 104 determines whether the control variable θ* determined in S4 is the optimal value. The method for determining whether it is the optimal value is not particularly limited. For example, the control variable determination unit 104 may determine that the control variable θ* is the optimal value if it matches the previously detected control variable in S4, and determine that it is not the optimal value if they do not match. Furthermore, the previously detected control variable refers to the control variable included in the control result data 201, that is, the control variable used to control the crane 5, and the control variable used to calculate the evaluation value for that control.
[0083] If the optimal value is determined in S5 (yes in S5), the process proceeds to S10. In S10 (control variable determination step), the control variable determination unit 104 determines the optimal value of the control variable of the crane 5 as θ*, and the process ends. Figure 3 The processing. Furthermore, in S10, the control variable determination unit 104 can cause the output unit 40 to output a determined θ*.
[0084] On the other hand, when it is determined in S5 that the control variable θ* is not the optimal value (not in S5), the control variable determination unit 104 notifies the user of the information processing device 1A by outputting the control variable θ* through the output unit 40. Based on this notification, the user causes the control device 3 to execute control of the crane 5 according to the control variable θ*, and observes and evaluates the control result. The evaluation method is not particularly limited; for example, the error between the ideal control result and the actual control result can be used as the evaluation value for calculation. The evaluation result is input to the information processing device 1A via the input unit 30.
[0085] In S6, the data acquisition unit 101 acquires the evaluation results input as described above. Furthermore, in S7, the data acquisition unit 101 associates the evaluation results acquired in S6 with the control variable θ* determined in the previous S4 and appends it to the control result data 201.
[0086] In S8, the prediction distribution calculation unit 102 uses the control result data 201, which includes the evaluation result and control variable θ*, added in S7, to optimize the parameters of the kernel function. Then, in S9 (update step), the prediction distribution calculation unit 102 uses the control result data 201, which includes the evaluation result and control variable θ*, added in S7, and the optimized kernel function parameters from S8, to calculate the prediction distribution of the evaluation function. The process then returns to S4, where the control variable retrieval unit 103 retrieves the control variables. By repeatedly adding control variables and updating the prediction distribution, the control variables from which the desired control result can be obtained can be determined.
[0087] (Implementation Method 2)
[0088] Another embodiment of the present invention will be described below. Furthermore, for ease of explanation, components having the same function as those described in the above embodiments will be marked with the same reference numerals and their descriptions will be omitted.
[0089] (Device Structure)
[0090] based on Figure 5 The structure of the information processing apparatus 1B in this embodiment will be described. Figure 5 This is a block diagram showing an example of the main component structure of the information processing device 1B. As shown, the information processing device 1B includes a control unit 10B that comprehensively controls all parts of the information processing device 1B. The control unit 10B and... Figure 1 The difference between the control unit 10A of the information processing device 1A shown is that the prediction distribution calculation unit 301 is included in the control unit 10B instead of the prediction distribution calculation unit 102.
[0091] The prediction distribution calculation unit 301, like the prediction distribution calculation unit 102, uses the control result data 201 to calculate and update the prediction distribution. However, as will be explained below, the method of calculation and updating is different from that of the prediction distribution calculation unit 102.
[0092] The prediction distribution calculation unit 301 calculates or updates the prediction distribution by using the contribution of each of the multiple control result data in the prediction distribution calculation as a contribution corresponding to the reliability of the control result data. Therefore, even when the control result data used in the calculation or update of the prediction distribution includes data with low reliability, the impact of such control result data on the prediction distribution can be relatively reduced. Furthermore, this allows for the rapid determination of appropriate control variables.
[0093] Furthermore, the reliability of control outcome data is an indicator of whether the control outcome data represents an appropriate value from the perspective of the overall predicted distribution. For example, when a control outcome data point is removed from multiple control outcome data points, if the predicted distribution of the remaining control outcome data is close to a Gaussian distribution, it can be said that the removed control outcome data is more likely to be a deviation from the true function (evaluation function), and therefore has lower reliability. Conversely, if a control outcome data point is not removed from multiple control outcome data points, and the predicted distribution is closer to a Gaussian distribution compared to the case where removal was performed, then the reliability of the control outcome data can be said to be higher.
[0094] In Implementation 1, the prediction distribution calculation unit 102 uses Gaussian process regression to calculate and update the prediction distribution. In contrast, the prediction distribution calculation unit 301 of this embodiment uses robust Gaussian process regression, which makes Gaussian process regression robust, to calculate and update the prediction distribution. By making Gaussian process regression robust, the prediction distribution can be calculated and updated stably even if the control result data contains deviation values.
[0095] exist Figure 4 The example shown is a comparison between robust Gaussian process regression and Gaussian process regression. Figure 4 The diagram shows: functions based on the mean and variance of the predicted distribution calculated using Gaussian process (GP) regression, and functions based on the mean and variance of the predicted distribution calculated using robust Gaussian process (RGP) regression.
[0096] These functions are all constructed based on the same control result data. It should be noted that, compared to using all control result data to construct functions in GP, RGP reduces or eliminates the influence of deviation values as shown in the figure when constructing functions.
[0097] When using control outcome data containing deviations to perform Gaussian process regression to calculate the predicted distribution, the deviations may lead to a predicted distribution inconsistent with the true function (evaluation function). Even when applying Gaussian process regression (GP), using a large amount of control outcome data can make the predicted distribution approximate the true function. However, in... Figure 4 When functions are constructed based on the same amount of control result data, the results are as follows: the function constructed using RGP is extremely consistent with the true function, while the function constructed using GP deviates significantly from the true function.
[0098] And, as Figure 4 As shown, the function constructed using GP reaches its maximum value when θ = 0.4, according to the true function; however, it is not actually the maximum value when θ = 0.4. On the other hand, since a function that is roughly consistent with the true function is constructed using RGP, by using the function constructed using RGP, it is possible to find that the θ value that maximizes the evaluation value is 2.0.
[0099] When constructing functions based on the same control result data, it may be impossible to find the θ that maximizes the evaluation value using GP, but it is possible to find the θ that maximizes the evaluation value using RGP. This is because, as explained later, RGP uses the Student's t-distribution as the likelihood function, thereby allowing control result data with low reliability to be treated as deviation values, reducing their contribution.
[0100] (Regarding the formulas used in the calculation and updating of the predicted distribution)
[0101] Similar to Implementation 1, when the control result data 201 includes N control variables, these control variables are represented as
[0102] [Number 10]
[0103]
[0104] The evaluation value for it is expressed as
[0105] [Number 11]
[0106]
[0107] In addition, the function between input and output data is represented as follows.
[0108] [Number 12]
[0109]
[0110] Here, the prior distribution of the above function is set as follows.
[0111] [Number 13]
[0112] f = N(f|0, K) Θ (6)
[0113] In this embodiment, the regression of the evaluation function can be performed stably even with deviations. Therefore, a distribution robust to deviations is used instead of the Gaussian distribution as the likelihood function in Gaussian process regression. For example, the student's t-distribution can be used as the likelihood function. In this case, the likelihood function is represented by the following equation (7). Furthermore, a and b in equation (7) are the parameters of the likelihood function, and Γ represents the gamma function.
[0114] [Number 14]
[0115]
[0116] Here, the Gaussian distribution is not the conjugate of the student's t-distribution. Therefore, the post-hoc distribution cannot be calculated analytically. Thus, an analytical solution for the post-hoc distribution is approximated. For example, as explained below, the variational Bayes method can be used to approximate the analytical solution for the post-hoc distribution.
[0117] First, a scale-mixture representation is used, which uses Gaussian and gamma distributions to represent the likelihood function, i.e., the student's t-distribution.
[0118] [Number 15]
[0119] p(y n|f n )=∫p(y n |f n , τ n )p(τ n )dτ n (8)
[0120] [Number 16]
[0121]
[0122] [Number 17]
[0123] p(τ n )=Gam(τ n |a, b) (10)
[0124] Therefore, the likelihood function can be viewed as a Gaussian distribution with a gamma distribution in the reciprocal of the variance. Furthermore, τ in equations (8) to (10) n τ is the reciprocal of the variance of the Gaussian distribution for the nth control result data 201. n This indicates the reliability of the nth control result data 201.
[0125] The prediction distribution calculation unit 301 approximates the analytical solution of the posterior distribution of the model through variational inference. Specifically, the prediction distribution calculation unit 301 calculates the variational distribution that maximizes the lower bound of the logarithmic periphery likelihood. Since this variational distribution is an approximation of the posterior distribution, the prediction distribution calculation unit 301 can approximate the posterior distribution.
[0126] [Number 18]
[0127] log p(Y)=log∫p(Y|f,T)p(T)p(f)dfdT (11)
[0128] Here,
[0129] [Number 19]
[0130]
[0131] Assuming that the distributions of f and T are independent, if we introduce the variational distribution q(f) and
[0132] [Number 20]
[0133]
[0134] Then the prediction distribution calculation unit 301 can calculate the lower bound Fv according to the following formula (12).
[0135] [Number 21]
[0136]
[0137] Furthermore, the prediction distribution calculation unit 301 uses the above formula (12) to calculate the variational distribution that maximizes the lower bound of the peripheral likelihood. This variational distribution is, as described above, an approximation of the post-hoc distribution.
[0138] Variational distributions q(f), q(τ) n The update rule for τ can be parsed as follows. As mentioned above, τ n This represents the reliability of the nth control result data 201. Therefore, the prediction distribution calculation unit 301 derives q(τ) that maximizes the lower bound Fv. n That is, τ n The post-hoc distribution is then used to calculate the predicted distribution (mean and variance) of the evaluation function, thereby enabling the calculation of the reliability-based predicted distribution. In other words, τ represents the reliability of the control outcome data 201. n The ex-post distribution is used to weight the control result data 201 in the calculation of the prediction distribution, thus reducing the contribution of the relatively unreliable control result data 201 to the calculation of the prediction distribution. Therefore, the impact of the relatively unreliable control result data 201 on the prediction distribution can be reduced to zero or minimized.
[0139] [Number 22]
[0140] q(f)=N(f|μ f , ∑ f (13)
[0141] [Number 23]
[0142]
[0143] [Number 24]
[0144]
[0145] [Number 25]
[0146]
[0147] [Number 26]
[0148] q(τ n )=Gam(τ n |a n b n (17)
[0149] [Number 27]
[0150]
[0151] [Number 28]
[0152]
[0153] The prediction distribution calculation unit 301 uses the approximation of the post-hoc distribution obtained through the above formula to calculate the prediction distribution for any input θ. * The evaluation function is the predicted mean function and variance function. Specifically, the prediction distribution calculation unit 301 calculates the mean function and variance function using the following formulas (20) and (21).
[0154] [Number 29]
[0155]
[0156] [Number 30]
[0157]
[0158] Furthermore, the control variable retrieval unit 103 uses the aforementioned mean function and variance function to retrieve the point that maximizes the obtained function, i.e., the candidate for the optimal control variable. For example, when applying the UCB strategy, the control variable retrieval unit 103, in the same manner as in Embodiment 1, calculates the obtained function using formula (5) and retrieves the point that maximizes the obtained function.
[0159] (Use of control result data for similar tasks)
[0160] As described above, the objective of this embodiment is to enable the crane 5 to perform actions or tasks. If the objective changes, the optimal control variable will also change. However, even for other objectives, as long as they are similar (hereinafter referred to as similar tasks), the predicted distribution of the evaluation function may be similar. In this case, control result data from similar tasks can be used. The following describes a method for calculating or updating the predicted distribution using control result data from other tasks.
[0161] When using control result data from other tasks, the prediction distribution calculation unit 301, when calculating and updating the prediction distribution for a task, uses the contribution of the prediction distribution calculation of the control result data from other tasks as a contribution corresponding to the similarity between the other tasks and the aforementioned task, to calculate or update the prediction distribution. Furthermore, the aforementioned other tasks include similar tasks. Additionally, the aforementioned other tasks may also include dissimilar tasks.
[0162] Based on this structure, the prediction distribution is calculated and updated using control result data from other tasks. Therefore, compared to using control result data from only one task, fewer updates are needed to determine the appropriate control variables. Furthermore, since the contribution of control result data from other tasks to the prediction distribution calculation is reflected in the similarity to a single task, it eliminates the need to sort through similar tasks from multiple tasks.
[0163] The following provides a detailed explanation of the method for using control result data from other tasks. When using control result data from other tasks, the prediction distribution calculation unit 301 will use the previously retrieved control variables...
[0164] [Number 31]
[0165]
[0166] Evaluation value
[0167] [Number 32]
[0168]
[0169] Task tags for data points
[0170] [Number 33]
[0171]
[0172] As learning data, it enables the regression of evaluation functions for similar tasks.
[0173] When processing M tasks, set the task label to t. n ∈{1, ..., M}. Furthermore, the same real value is assigned to the task label for the control result data of the same task. That is, the task label indicates whether the control result data of Θ and Y are data from when each task was executed. In other words, the task label is a label used to distinguish the same task.
[0174] To regress the evaluation function according to the task, the task label is processed as input to a robust Gaussian process. Therefore, as shown in equation (22) below, the input kernel k(θ, θ') is compared with the task kernel t. n The product of (t, t') is used as the kernel function.
[0175] k((θ,t), (θ',t'))=k t (t,t')k θ (θ,θ') (22)
[0176] The task kernel is a function representing task similarity, while the task label t is the input. n These are tags used to distinguish tasks within the same task. Therefore, you cannot determine a task's identity based solely on its tag. n The similarity of the tasks is calculated using the values of the input task labels. Furthermore, since there are M tasks, the output of the task kernel is an M×M pattern. Therefore, the task kernel is represented by an M-th power square matrix Kt, and the values of the elements represented by the task labels input to the task kernel are used as the output of the task kernel.
[0177] kt (t, t') = [K t ] t,t’ (twenty three)
[0178] Additionally, since the task kernel is used as a kernel function, K is required. t It is a positive definite matrix. Therefore, K is decomposed using Korysky decomposition and by utilizing the lower triangular matrix L. t Decompose into Kt=LL T Therefore, the M(M+1) / 2 elements of the lower triangular matrix L can be used as the parameter α of the task kernel. t , this parameter α t Optimization is performed within the framework of variational inference, learning task similarity from control outcome data. Furthermore, learning task similarity means updating the post-hoc distribution in a manner that reflects task similarity (making the contribution of control outcome data from similar tasks greater than the contribution of control outcome data from dissimilar tasks).
[0179] The parameter α is optimized in this way. t This represents the contribution (or weight) of the control result data from other tasks. Therefore, the prediction distribution calculation unit 301 uses the optimized parameter α. t By determining the prediction distribution of the evaluation function, we can use the contribution of control outcome data from other tasks as the contribution corresponding to the similarity between the other tasks and the target task to calculate the prediction distribution. The same applies to updating the prediction distribution.
[0180] (Processing flow)
[0181] The processing flow for determining the control variables of crane 5 using information processing device 1B is described. This processing flow is similar to... Figure 3 The processing flow of the information processing device 1A shown is generally the same, but the processing steps S2, S8, and S3 are different. The following explanation focuses on this difference.
[0182] Figure 6 This is a flowchart illustrating an example of the process of optimizing the parameters of a kernel function. Figure 6 The handling is in conjunction with Figure 3 After the same processing as S1, that is, after the initial data of the data acquisition unit 101 is set, the processing is the same as... Figure 3 The processing corresponds to S2. Additionally, information processing device 1B replaces... Figure 3 S8 processing and execution Figure 6 The processing.
[0183] In S21, the prediction distribution calculation unit 301 initializes the parameters of the kernel function. The initialized kernel function parameter is α. k and α tThese two. Next, in S22, the prediction distribution calculation unit 301 updates the variational distributions q(f) and q(τ). n Variational distributions q(f), q(τ) n The update rules for ) are as described in equations (13) to (19) above.
[0184] In S23, the prediction distribution calculation unit 301 determines whether the variational lower bound has converged. The variational distributions q(f) and q(τ) at which the variational lower bound converges are... n ) is the optimized variational distribution. Furthermore, convergence conditions can be set appropriately. For example, it is also possible to find the convergence values at q(f) and q(τ). n The algorithm calculates Fν before and after the update and determines convergence when the difference is less than a specified value (e.g., 0.1).
[0185] If convergence is determined in S23 (yes in S23), the process proceeds to S24. On the other hand, if non-convergence is determined (no in S23), the process returns to S22, and the variational distribution is updated again.
[0186] In S24, the prediction distribution calculation unit 301 determines the parameter α of the kernel function that maximizes the variational lower bound. k * α t * The above formula (12) is used in this operation. Furthermore, q(f), q(T), and p(f) in formula (12) contain a matrix K obtained using the kernel function. Therefore, Fν can be used as a function with parameter α. k α t The function is processed. Therefore, for example, any nonlinear optimization method can be used for optimization. As an example of a nonlinear optimization method, the gradient method can be cited.
[0187] In S25, the prediction distribution calculation unit 301 determines whether to end the optimization. This can be done by setting appropriate termination conditions. For example, Fν can be calculated before and after the processing in S22 to S25, and the optimization can be terminated when the difference is less than a specified value (e.g., 0.1).
[0188] When it is determined to be the end in S25 (yes in S25), then it ends. Figure 6 The processing. Afterwards, proceed with... Figure 3 The same process applies after S3. On the other hand, if it is determined that the process does not end (no in S25), the process returns to S22 and the variational distribution is updated again.
[0189] As mentioned above, in Figure 6In the processing, the calculation of the variational distribution that maximizes the lower bound of the variation is performed alternately with the optimization of the kernel function parameters. Thus, an approximation of the post-hoc distribution, i.e., the variational distribution, can be obtained. Furthermore, in... Figure 6 In the processing, the parameter α of the kernel function can be made k Optimize and optimize α t Therefore, it is able to learn the similarity between tasks.
[0190] Furthermore, it can be said that the variational lower bound Fν is obtained by approximately calculating whether the robust Gaussian process with multitasking can well represent the control result data 201. Therefore, by finding the parameter α that maximizes the variational lower bound Fν... t This allows us to determine the similarity of the control result data 201.
[0191] By repeating the processes S22 to S25, the parameter α is adjusted in the following way: t Optimization involves using control result data 201 from similar tasks and reducing the contribution of control result data 201 from dissimilar tasks. In other words, by repeatedly performing the processes S22 to S25, the control result data of other tasks are weighted according to the similarity between other tasks and the target task.
[0192] Based on the above processing, when calculating the prediction distribution for a task, appropriate consideration can be made based on the contribution of other tasks to the similarity between them and the aforementioned task, thereby reusing the control result data of those other tasks. Therefore, the amount of control result data for a task can be reduced, and appropriate control variables can be determined.
[0193] [Example]
[0194] Experiments were conducted to verify the effectiveness of information processing devices 1A and 1B. Based on Figure 7 and Figure 8 Explain the results. Figure 7 This is a graph representing the experimental results. Figure 8 This is a diagram showing the task kernel at the end of the optimization in experiments 10-12.
[0195] Furthermore, instead of using an actual crane 5, the experiment used a small simulated crane to the extent that it could be used in a laboratory, and the simulated garbage consisted of a mixture of paper cut by a shredder and toy rubber balls.
[0196] The task of the crane is to grab and lift garbage, then move it a predetermined distance while evenly distributing the garbage within that distance. Specifically, the crane performs the following actions: after initially opening the grab bucket that has grabbed the garbage, when garbage of weight θ1 falls from the grab bucket, the grab bucket begins to close, and after a time θ2, it reopens. θ1 and θ2 are the controlled variables.
[0197] In evaluating the control results, the ideal process is one where the weight of the garbage grabbed by the grab decreases at a constant rate as the crane's travel distance increases. The evaluation value is calculated based on the difference between this ideal and actual process. Specifically, the root mean square (RMS) is used to calculate a series of standardized data *w* representing the actual grab weight of the garbage grabbed by the crane, and a series of data *w* representing the ideal grab weight. I The difference was evaluated using the following formula (24).
[0198] E(θ)=5-10×RMS(w(θ-w) I ) (twenty four)
[0199] As mentioned above, since the simulated waste is just as non-uniform as the waste stored in the actual landfill, even with the same action parameters, w(θ) may be significantly different, which will affect the evaluation value E(θ).
[0200] In addition, the initial weight of the garbage to be grabbed was set to 120-300g, and the moving distance was set to 40cm. In one experiment, the optimized control variables θ1 and θ2 were used to perform the task ten times, and their control results were evaluated using the above equation (24).
[0201] A total of 12 experiments, numbered 12, were conducted. In experiments 1-3, the control variables θ1 and θ2 were optimized using information processing device 1A. Furthermore, in experiments 4-9, the control variables θ1 and θ2 were optimized using information processing device 1B. It should be noted that data from a similar task was not used in experiments 4-9. In experiments 10-12, the control variables θ1 and θ2 were optimized using information processing device 1B and data from a similar task. This similar task involved setting the crane's movement distance to 30 cm.
[0202] like Figure 7 The results confirm that although the optimized control variables θ1 and θ2 in experiments 1-12 had some deviations, the evaluation values were all at a high level, indicating that appropriate optimization had been performed.
[0203] A comparison of the results of experiments 1-3 and 4-6 reveals a difference in the number of trials required for optimization. This indicates that, for optimization using information processing device 1B, compared to optimization using information processing device 1A, fewer trials are needed to calculate the appropriate control variable. Furthermore, the number of trials required for optimization refers to the number of trials required until the optimal control variable is determined (until...). Figure 3 The number of times the crane is activated and a new control result is obtained by using the control variables determined by the obtained function (until it is determined in S5).
[0204] Furthermore, comparing the results of experiments 7–9 with those of 10–12 also reveals a difference in the number of trials required for optimization. This indicates that by using control outcome data from similar tasks, it is possible to calculate the appropriate control variables with fewer trials.
[0205] In addition, Figure 8 The graph shows the task kernels at the end of optimization for experiments 10-12. The vertical and horizontal axes represent task labels, and the values represent the similarity between tasks. Figure 8 As shown, the off-diagonal component 9-2 (1.35), representing the similarity (contribution in the calculation of the prediction distribution) between a task with a movement distance of 40cm and a task with a movement distance of 30cm (similar tasks), is significantly higher than the off-diagonal component 9-1 (0.96). This indicates that control data from similar tasks were used in the calculation of the prediction distribution.
[0206] Furthermore, although not shown in the figure, the same experiment was conducted as follows: the crane's movement distance in the above task was changed to 20cm, and similar tasks with crane movement distances of 30cm and 40cm were set up. The result was that control variables θ1 and θ2 could be calculated with the same precision as the results mentioned above with a relatively small number of trials (around 10). It can be seen that in this case, the task kernel at the end of optimization is also similar to... Figure 8 The example also shows that the off-diagonal components have larger values and uses control outcome data from a similar task.
[0207] In addition, an experiment was conducted using a real crane 5 to disperse garbage in a landfill. The results showed that, similar to the examples above, the appropriate control variables could be calculated using information processing device 1A, and information processing device 1B could calculate the appropriate control variables with fewer trials.
[0208] Furthermore, in the experiment using the actual crane 5, the operator was also asked to perform the task, and the results were evaluated using the above formula (24). Moreover, when t-tests were performed, there was no substantial difference between the evaluation value of the control result using the control variables optimized by the information processing device 1B and the evaluation value of the operator's control result. That is to say, it can be said that the control using the control variables optimized by the information processing device 1B is a high-level control of the same degree as the operator's control.
[0209] (Software-based implementation example)
[0210] The control modules of the information processing devices 1A and 1B (especially the components included in the control unit 10A and the control unit 10B) can be implemented by logic circuits (hardware) formed on integrated circuits (IC chips) or by software.
[0211] In the latter case, i.e., software, information processing devices 1A and 1B include a computer that executes commands of software, i.e., programs (control variable determination programs), to implement various functions. This computer, for example, includes one or more processors and a computer-readable storage medium storing the aforementioned program. Furthermore, in the computer, the processor reads the program from the storage medium and executes it, thereby achieving the object of the present invention. As the processor, for example, a CPU (Central Processing Unit) can be used. As the storage medium, a "non-transitory tangible medium" can be used, such as storage tapes, storage disks, memory cards, semiconductor memories, programmable logic circuits, etc., in addition to ROM (Read Only Memory). Additionally, RAM (Random Access Memory) for deploying the program can also be included. Furthermore, the program can be provided to the computer via any transmission medium capable of transmitting the program (communication network, radio waves, etc.). Moreover, one aspect of the present invention can also be implemented by using a data signal with an embedded carrier wave, which is embodied by electronically transmitting the program.
[0212] (Modified Example)
[0213] This invention is not limited to the embodiments described above, and various modifications can be made within the scope of the claims. Embodiments obtained by appropriately combining technical solutions disclosed in different embodiments are also included in the technical scope of this invention.
[0214] For example, for the information processing apparatus 1A of Embodiment 1, optimization of control result data using a similar task can also be performed. In this case, simply using the kernel function shown in equation (22) and through...Figure 3 S2 and 8 can be used to optimize the parameter αt of the kernel function.
[0215] Furthermore, the entities that execute the processes described in the above embodiments can be appropriately changed. Figure 3 The method for calculating control variables shown can also be executed by multiple information processing devices. Similarly, Figure 6 The method for determining the control variables shown can also be executed by multiple information processing devices.
[0216] Furthermore, in the above embodiments, examples of optimizing control variables in the task of agitating dispersed waste have been described, but the content is not particularly limited as long as the task is performed by a crane that transports waste. For example, it is also possible to optimize control variables for tasks such as: tasks that cause the crane to grab waste, tasks that cause the crane to lift the grabbed waste, and tasks that cause the lifted waste to drop.
[0217] Explanation of reference numerals in the attached figures
[0218] 1A - Information processing device; 102 - Predictive distribution calculation unit; 103 - Control variable retrieval unit; 104 - Control variable determination unit; 201 - Control result data; 1B - Information processing device; 301 - Predictive distribution calculation unit; 3 - Control device; 5 - Crane; 9 - Control system.
Claims
1. An information processing device comprising: Predictive distribution calculation department Control variable retrieval unit, and Control variable determination unit, The prediction distribution calculation unit uses control variable data obtained by correlating the control result data of the crane handling waste with the control result data obtained by controlling the crane using the control variable, to calculate a prediction distribution of a function representing the relationship between the control variable and the control result. This prediction distribution is represented by the mean and variance of the function predicted based on the control result data. The control variable retrieval unit retrieves candidate control variables, i.e., alternative control variables, based on the predicted distribution. The prediction distribution calculation unit updates the prediction distribution using the candidate control variables retrieved by the control variable retrieval unit and the control result of controlling the crane using the candidate control variables. The control variable determination unit uses a function based on the mean and variance of the updated predicted distribution to determine the optimal value of the control variable.
2. An information processing device, comprising: Predictive distribution calculation department Control variable retrieval unit, and Control variable determination unit, The prediction distribution calculation unit uses control variable data obtained by correlating it with control result data obtained by controlling the crane using the control variable to control the crane, to calculate the prediction distribution of a function representing the relationship between the control variable and the control result. The control variable retrieval unit retrieves candidate control variables, i.e., alternative control variables, based on the predicted distribution. The prediction distribution calculation unit updates the prediction distribution using the candidate control variables retrieved by the control variable retrieval unit and the control result of controlling the crane using the candidate control variables. The control variable determination unit uses a function based on the updated predicted distribution to determine the optimal value of the control variable. The prediction distribution calculation unit uses the contribution of each of the multiple control result data in the prediction distribution calculation as a contribution corresponding to the reliability of the control result data, and calculates or updates the prediction distribution.
3. An information processing device, comprising: Predictive distribution calculation department Control variable retrieval unit, and Control variable determination unit, The prediction distribution calculation unit uses control variable data obtained by correlating it with control result data obtained by controlling the crane using the control variable to control the crane, to calculate the prediction distribution of a function representing the relationship between the control variable and the control result. The control variable retrieval unit retrieves candidate control variables, i.e., alternative control variables, based on the predicted distribution. The prediction distribution calculation unit updates the prediction distribution using the candidate control variables retrieved by the control variable retrieval unit and the control result of controlling the crane using the candidate control variables. The control variable determination unit uses a function based on the updated predicted distribution to determine the optimal value of the control variable. The prediction distribution calculation unit uses the contribution of control result data of other tasks that are different from the tasks performed by the crane using the control variables in the prediction distribution calculation as the contribution corresponding to the similarity between the other tasks and the task, to calculate or update the prediction distribution.
4. A control system comprising: The information processing apparatus according to any one of claims 1 to 3; A control device that uses the control variables to control the crane; and The crane.
5. A method for determining control variables, executed by one or more information processing devices, comprising: The prediction distribution calculation step uses control variable for a garbage-handling crane and control result data obtained by associating the control result of controlling the crane with the control variable to calculate a prediction distribution of a function representing the relationship between the control variable and the control result. The prediction distribution is represented by the mean and variance of the function predicted based on the control result data. The control variable retrieval step involves retrieving, based on the predicted distribution, alternative control variables, i.e., candidate control variables, for the optimal value of the control variable. The update step involves updating the predicted distribution using the candidate control variable retrieved in the control variable retrieval step and the control result of controlling the crane using the candidate control variable; and The control variable determination step uses a function based on the mean and variance of the updated predicted distribution to determine the optimal value of the control variable.
6. A storage medium that stores a computer-readable storage medium containing a program for determining control variables. The following steps are achieved by executing the control variable determination program through the processor: The prediction distribution calculation step uses control variable for a garbage-handling crane and control result data obtained by associating the control result of controlling the crane with the control variable to calculate a prediction distribution of a function representing the relationship between the control variable and the control result. The prediction distribution is represented by the mean and variance of the function predicted based on the control result data. The control variable retrieval step involves retrieving, based on the predicted distribution, alternative control variables, i.e., candidate control variables, for the optimal value of the control variable. The update step involves updating the predicted distribution using the candidate control variable retrieved in the control variable retrieval step and the control result of controlling the crane using the candidate control variable; and The control variable determination step uses a function based on the mean and variance of the updated predicted distribution to determine the optimal value of the control variable.
Citation Information
Patent Citations
Waste agitation evaluating method, waste agitation evaluating program, and waste agitation evaluating device
JP2010275064A
Grab quantity control method for unloader
JP1993092891A