Information processing device, crane control system, learning method, and learning program
The information processing device uses a policy model and self-driven policy search to stabilize crane control in waste treatment facilities, addressing data variability and ensuring accurate action execution.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANADEVIA CO LTD
- Filing Date
- 2022-02-28
- Publication Date
- 2026-05-08
AI Technical Summary
Existing automated crane control systems in waste treatment facilities face challenges in accurately determining control settings due to the variability in observational data, leading to frequent switching of actions and difficulty in generating a highly accurate predictive model.
An information processing device that generates a policy model using observational data to determine appropriate crane actions and execution lengths, employing a self-driven policy search method with Gaussian processes to reduce variability and enhance accuracy.
Enables highly accurate automatic control of cranes in unstable environments by maintaining consistent actions for determined execution lengths, even with varying observational data.
Smart Images

Figure 0007855192000036 
Figure 0007855192000037 
Figure 0007855192000038
Abstract
Description
[Technical Field]
[0001] This invention relates to an information processing device and the like that can be used for the automatic control of a crane used to transport waste. [Background technology]
[0002] Generally, waste treatment plants that process waste are equipped with storage facilities called pits for storing the waste brought in. If the waste stored in the pit is, for example, combustible waste, it is transferred within the pit using a crane with a bucket, mixed, then put into a hopper and sent to an incinerator for incineration.
[0003] Development of technologies to automate the control of such cranes has been ongoing. For example, Patent Document 1 below discloses a technology that automatically controls a crane by calculating an evaluation value for the number of times waste is agitated at various points in the pit, selecting the position of the crane bucket based on that evaluation value, and opening and closing the bucket at that position. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Patent No. 5185197 [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] To improve the accuracy of the automated control of cranes as described above, it is necessary to accurately understand the state of the waste using observational data observed at the waste treatment plant. However, the observational data observed at the waste treatment plant has a large variation in values due to factors such as the inconsistency in the quality of the waste being treated.
[0006] Therefore, if, for example, such observational data is used as training data to generate a predictive model that predicts the actions a crane should perform, there is a possibility that the generated predictive model will frequently switch the actions the crane should perform in response to the variability in the observational data. Thus, it is difficult to generate a highly accurate trained model when the observational data has large variability. Consequently, it is also difficult to determine appropriate control settings from observational results obtained under unstable conditions.
[0007] One aspect of the present invention aims to realize an information processing device, etc., that enables the determination of appropriate control content from acquired observation results, even in unstable environments, when controlling cranes in waste treatment facilities. [Means for solving the problem]
[0008] To solve the above problems, an information processing device according to one aspect of the present invention includes a learning unit that generates a policy model using observation data observed at a waste treatment facility equipped with a crane for transporting waste, which includes an action policy for determining an action to be performed by the crane and an execution length policy for determining the execution length of the action.
[0009] Furthermore, a learning method according to one aspect of the present invention is a learning method performed by one or more information processing devices, comprising: a data acquisition step of acquiring observation data observed at a waste treatment facility equipped with a crane for transporting waste; and a learning step of generating a policy model using the observation data, which includes an action policy for determining an action to be performed by the crane and an execution length policy for determining the execution length of the action. [Effects of the Invention]
[0010] According to one aspect of the present invention, in controlling a crane in a waste treatment facility, it becomes possible to determine appropriate control settings from acquired observation results, even in an unstable environment. [Brief explanation of the drawing]
[0011] [Figure 1] This is a block diagram showing an example of the main components of an information processing device according to one embodiment of the present invention. [Figure 2] This figure shows an example configuration of a crane control system including the above-mentioned information processing device. [Figure 3] This is a graphical model of GPSTPS. [Figure 4] This diagram outlines the task of scattering waste. [Figure 5] This flowchart shows an example of the processing performed by the information processing device during the learning of action strategies and execution strategies. [Figure 6] This flowchart details the processes performed during a task trial. [Modes for carrying out the invention]
[0012] [System Configuration] The configuration of the crane control system according to this embodiment will be explained with reference to Figure 2. Figure 2 is a diagram showing an example of the configuration of the crane control system 7. The crane control system 7 shown in Figure 2 includes an information processing device 1, a crane control device 3, and a crane 5. The crane 5 is a crane that transports waste stored in a pit P, which is a waste storage facility. More specifically, the crane 5 is a bucket crane equipped with a bucket for gripping waste. The waste can be anything that can be transported by the crane 5, such as household waste or industrial waste.
[0013] The information processing device 1 uses observational data from a waste treatment facility equipped with a crane 5 to determine the action to be performed by the crane 5 and the duration of that action. The crane control device 3 then operates the crane 5 based on the action and duration determined by the information processing device 1. The duration of the action indicates how long the same action is performed. For example, if the action can be counted, the number of times the action is performed can be used as the duration of the action. Alternatively, the duration of performing the same action can also be used as the duration of the action.
[0014] Furthermore, the observational data is used to identify the conditions under which the crane 5 performs its actions. The type of observational data to be used should be determined appropriately depending on the task to be performed by the crane 5 and how that task will be evaluated. For example, if the task is to transport waste with the crane 5, the weight of the waste held in the crane's bucket could be used as observational data.
[0015] In general, observational data observed at waste treatment facilities exhibits significant variability due to factors such as inconsistent quality of the waste being processed. Therefore, if a machine learning model for controlling crane 5 is generated from such observational data, it may result in a model that frequently switches the actions performed by crane 5 in response to the variability in the observational data. Thus, generating a highly accurate machine learning model is difficult when observational data exhibits significant variability.
[0016] Therefore, the information processing device 1 uses the above observation data to generate a policy model that includes an action policy for determining the action to be performed by the crane 5, and an execution length policy for determining the execution length of the said action.
[0017] According to the above policy model, the same action is maintained for the duration of the execution length determined by the execution length policy. Therefore, even when there is a large variation in the observed data, the possibility of generating a machine learning model that frequently switches the actions to be performed by crane 5 in response to that variation can be reduced, thereby making it possible to generate a policy model, which is a highly accurate machine learning model.
[0018] The information processing device 1 then determines the action to be performed by the crane 5 using the action policy included in the machine learning-based policy model, and also determines the execution length of the action using the execution length policy included in the policy model. As a result, even in unstable environments, a reasonable control content can be determined from the acquired observation data, thereby realizing highly accurate automatic control of the crane 5.
[0019] As described above, the crane control system 7 includes a crane 5 for transporting waste, an information processing device 1 that determines an action to be performed by the crane 5 according to observed data using an action strategy and also determines the execution length of the above action according to the observed data using an execution length strategy, and a crane control device 3 that operates the crane 5 based on the action and execution length determined by the information processing device 1. As a result, even in an unstable environment, a reasonable control content can be determined from the acquired observed data, thereby realizing highly accurate automatic control of the crane 5.
[0020] [Configuration of the information processing device] The configuration of the information processing device 1 will be explained based on Figure 1. Figure 1 is a block diagram showing an example of the main components of the information processing device 1. As shown in the figure, the information processing device 1 includes a control unit 10 that controls all parts of the information processing device 1, and a storage unit 11 that stores various data used by the information processing device 1. The information processing device 1 also includes a communication unit 12 for the information processing device 1 to communicate with other devices, an input unit 13 that receives input of various data to the information processing device 1, and an output unit 14 for the information processing device 1 to output various data.
[0021] The control unit 10 also includes an action decision unit 101, an execution length decision unit 102, a crane control unit 103, an episode collection unit 104, and a learning unit 105. The memory unit 11 stores the action policy 111, the execution length policy 112, and the episode 113.
[0022] The action decision unit 101 uses an action policy 111 to determine the action to be performed by the crane 5, and determines an action based on the observation data observed at the waste treatment facility equipped with the crane 5. This process is performed both during the learning of the action policy 111 and after the learning is completed. If performed during learning, the action policy 111 being learned is used; if performed after learning is completed, the learned action policy 111 is used. The same applies to determining the execution length.
[0023] The action decision unit 101 may, for example, decide which action to have the crane 5 perform from among a predetermined set of actions. Each action can be defined arbitrarily. For example, if the action of opening the bucket of the crane 5 and the action of closing it are each defined as one action, the action decision unit 101 will decide whether to have the crane 5 perform the action of opening the bucket or the action of closing it. Alternatively, a series of actions such as opening the bucket, lowering the open bucket onto the waste surface in the pit P, closing the bucket to grasp the waste, and winding up the wire to lift the grasped waste may be defined as one action. Furthermore, even if the actions are the same or of the same type, they may be defined as different actions depending on the degree or manner in which they are performed, for example, by defining the action of opening the bucket to a certain angle and the action of opening it to a different angle as separate actions.
[0024] The execution length determination unit 102 determines the execution length according to the observed data using the execution length policy 112 for determining the execution length of the action determined by the action determination unit 101. Details of the action policy 111 and the execution length policy 112 will be described later.
[0025] The crane control unit 103 operates the crane 5 based on the action determined by the action determination unit 101 and the execution length determined by the execution length determination unit 102. More specifically, the crane control unit 103 operates the crane 5 by notifying the crane control device 3 of the action to be performed by the crane 5 and its execution length via communication through the communication unit 12. Of course, the crane control unit 103 may also control the crane 5 without going through the crane control device 3, in which case the information processing device 1 will also perform the function of the crane control device 3.
[0026] The episode collection unit 104 collects episodes 113, which are training data used for machine learning of the action policy 111 and the execution length policy 112. The collected episodes 113 are stored in the memory unit 11. As will be described in detail later, episodes 113 include observation data observed at the waste treatment facility (data showing the state when the crane 5 was made to perform an action), as well as the actions performed by the crane 5 and their execution lengths.
[0027] The learning unit 105 generates a policy model using the observation data described above. This policy model includes an action policy 111 for determining the action to be performed by the crane 5, and an execution length policy 112 for determining the execution length of that action. Details of how the action policy 111 and the execution length policy 112 are generated will be described later.
[0028] As described above, the information processing device 1 includes a learning unit 105 that generates a policy model using observation data observed at a waste treatment facility equipped with a crane 5 for transporting waste, which includes an action policy 111 for determining an action to be performed by the crane 5, and an execution length policy 112 for determining the execution length of the said action.
[0029] According to the above policy model, the same action is maintained for the duration of the execution length determined by the execution length policy. Therefore, even when there is a large variation in the observed data, the possibility of generating a machine learning model that frequently switches the actions to be performed by crane 5 in response to that variation can be reduced. Thus, with the above configuration, it becomes possible to generate a policy model, which is a highly accurate machine learning model. Furthermore, by using this policy model, it becomes possible to determine appropriate control content from the acquired observed data, even in unstable environments.
[0030] Furthermore, as described above, after the learning of the policy model (i.e., the action policy 111 and the execution length policy 112) is completed, the action decision unit 101 uses the learned action policy 111 to determine an action in accordance with the observation data observed at the waste treatment facility equipped with the crane 5. Then, the execution length decision unit 102 uses the learned execution length policy 112 to determine the execution length in accordance with the observation data. As a result, even in an unstable environment, a reasonable control content can be determined from the acquired observation data, making it possible to achieve highly accurate automatic control of the crane 5.
[0031] [Searching for solutions] Before explaining how to learn the action policy 111 and the execution length policy 112, we will first explain the underlying method called policy search. Policy search aims to acquire a policy that maximizes the expected payoff in an environment formulated by a Markov decision process. In policy search, the policy is learned by repeatedly executing an action decided according to the policy and improving the policy based on the data obtained by executing that action. In other words, policy search is a reinforcement learning method that learns the optimal policy through repeated trial and error.
[0032] In typical policy search, the actions taken during the trial-and-error process using policy π are recorded in state s. t , action a t We collect episode d, which is a series of episodes. Episode d is expressed as shown in equation (1) below. Also, the initial state probability and state transition probability of the environment in which trial and error takes place are p(s1) and p(s1) and p(s1) respectively. t+1 |s t ,a t If we assume that the policy π is used, then the probability distribution p(d|π) of episode d can be expressed as shown in equation (2) below.
[0033]
number
[0034] Also, reward r t (s t ,a t Using ), the payoff R(d), which is the sum of rewards for episode d, is expressed as follows:
[0035]
number
[0036] In policy search, the expected payoff is calculated using the above probability distribution p(d) and payoff R(d), and the policy that maximizes this payoff is calculated. This calculation can be expressed mathematically as follows.
[0037]
number
[0038] Furthermore, in the policy search described above, a Gaussian process can be applied as the policy model. Policy search using a Gaussian process as the policy model is called Gaussian process policy search. Gaussian process policy search has the advantage of being able to learn a policy with a small amount of data. On the other hand, Gaussian process policy search has the disadvantage that the computational cost of learning and prediction increases with the amount of data used for learning, because it directly uses all the data used for learning to calculate predictions for unknown inputs. For this reason, it may be difficult to apply Gaussian process policy search to tasks that require a large number of action decisions for a single task execution.
[0039] [Gaussian process self-driven policy search] Therefore, the learning unit 105 updates the action policy 111 and the action length policy 112 by a new policy search method (hereinafter referred to as GPSTPS: GAUSSIAN PROCESS SELF-TRIGGERED POLICY SERCH) that incorporates the concept of action action length into the Gaussian process policy search described above.
[0040] This reduces the number of policy decisions required relative to the actual episode length obtained. Furthermore, the amount of data used for policy learning by Gaussian processes is reduced, decreasing the computational cost required for learning and prediction. This makes it possible to explore optimal policies for tasks that require long-term episodes, which were previously difficult to learn.
[0041] More specifically, this embodiment describes an example of GPSTPS that realizes policy search based on variational inference by using Gaussian processes as the policy models for the action policy 111 and the execution length policy 112, and treating the expected gain as a gain-weighted marginal likelihood function. The details of this are described below.
[0042] (1. Formulating the problem) Expand the above-mentioned normal policy search to formulate self-driven policy search. Self-driven policy search extends policy search by introducing the above-mentioned execution length and a binary update variable o t The execution length of the action a t started at time t is represented by the execution length τ t Also, the update variable o t is used to switch whether to continue the action being executed from the previous time step, or to update the action and execution length without continuing the action, and can also be called a gate variable.
[0043] In self-driven policy search, the action and execution length at each time step are assumed to follow the following distributions that respectively encapsulate the action policy π a and the execution length policy π τ
[0044]
Equation
[0045] In these two distributions, when o t = 1, the action and execution length are updated according to the action policy π a and the execution length policy π τ respectively. On the other hand, when o t is 0, that is, in the "else" case in the above equations (5) and (6), the action from the previous time step continues, and the execution length τ t is set back to the previous time step. Therefore, the execution length τ t indicates the remaining duration (or number of continuations) of the action a t The update variable o t at time t is determined by the execution length τ t-1 at the previous time step (t - 1). Represented by an equation, it is as shown in the following equation (7).
[0046]
Equation
[0047] According to equation (7), the effective length τ t-1 When is 1, update variable o t This becomes 1. And, as mentioned above, the update variable o t When this value becomes 1, the action and execution length are updated.
[0048] Also, execution length τ t and update variable o t Episode d and its probability distribution are extended as follows:
[0049]
number
[0050] Each element included in episode d shown in the above formula (8), i.e., state s t , action a t , execution length τ t , update variable o t And, action a t Reward based on r t The relationship is represented as shown in the graphical model in Figure 3. Figure 3 is a graphical model of GPSTPS. As shown in part 31 of the graphical model in Figure 3, the update variable o at time t t The value of is the execution length τ at time t-1. t-1 It is determined by the update variable o. t Once the value is determined, action a t and its execution length τ t And then, action a t Once decided, take action a t and state s t Reward r t As this is determined, the state s at the next time t+1 is also determined. t+1 That will be decided.
[0051] In self-driven policy search, the expected payoff is calculated using the probability distribution shown in equation (9) above, and the action policy π is chosen to maximize it. a and the execution long policy π τ We will learn this. This can be expressed mathematically as follows:
[0052]
number
[0053] (2. Gaussian process policy model) As described above, the action policy π is learned through self-driven policy search. a and the execution long policy π τ GPSTPS applies a Gaussian process as the policy model. In this embodiment, we will explain an example in which a sparse Gaussian process is used as the policy model instead of a normal Gaussian process. In this case, a nonlinear function with the state as input is used as the policy model for the action.
[0054]
number
[0055] Using this, a nonlinear function with the state as input is used as the policy model for execution length.
[0056]
number
[0057] We assume the following Gaussian distribution using [the specified formula].
[0058]
number
[0059] Here,
[0060]
number
[0061] And, σ f 2 , σ g 2 is the variance of each Gaussian distribution. As shown in parts 31 and 32 of the graphical model in Figure 3, st If it is decided f t It was decided, f t If it is decided a t This is determined. Similarly, as shown in parts 31 and 33 of the graphical model in Figure 3, s t If it is decided g t It has also been decided, g t Once it is decided τ t That will be decided.
[0062] The following Gaussian processes are used as prior distributions for each nonlinear function.
[0063]
number
[0064] Here, k(·,·) is the kernel function, and m g (·) is the mean function of the function g.
[0065] To reduce the computational complexity of the Gaussian process used as the prior distribution for the nonlinear function, pseudo-input data consisting of M elements, which is significantly less than the number of data points used for policy learning, is used.
[0066]
number
[0067] and the corresponding pseudo-output of the aforementioned nonlinear function
[0068]
number
[0069] We introduce the following and set the prior distribution as follows.
[0070]
number
[0071] Here,
[0072]
number
[0073] This is a kernel gram matrix using pseudo-inputs,
[0074]
number
[0075] The output of the nonlinear function is f. t , g t The output of the nonlinear function f is shown by the following Gaussian distribution through Gaussian process regression using the distribution of pseudo-outputs. t , g t The relationship between these pseudo-outputs is represented using the kernel parameter θ as shown in parts 32 and 33 of the graphical model in Figure 3.
[0076]
number
[0077] Here,
[0078]
number
[0079] Therefore, the probability of episode d of a self-driven policy search using a Gaussian process as the policy model, i.e., GPSTPS, can be calculated as follows.
[0080]
number
[0081] Here,
[0082]
number
[0083] That is the case.
[0084] (3. Variational Learning for Policy Improvement) In this embodiment of GPSTPS, the reinforcement learning problem using policy search is formulated as a problem of maximizing the gain-weighted marginal likelihood, and the policy is improved by obtaining the posterior distribution of the nonlinear function in the two policy models using variational inference. This will be explained below.
[0085] Using the above formula (20), the expected gain J of the self-driven policy search using the Gaussian process policy model is calculated as follows.
[0086]
number
[0087] The above formula (21) is d and pseudo-output
[0088]
number
[0089] The complexity caused by makes it difficult to solve the integral analytically. Therefore, the policy used for logarithms and data samples, and the expected gain J of that policy are difficult to analyze. old and variational distribution
[0090]
number
[0091] By introducing this, the lower bound logJ of the expected payoff is obtained. L We seek this. Note that the policy used for the data sample is the policy applied to the task that was tested during training.
[0092]
number
[0093] In equation (23) above, C is a constant term that combines the policy-independent terms on the right-hand side of equation (22). For example, p(s1), p(s1) in equation (20) t+1 |s t ,a t ), and p(o t |τ t-1 ) is included in C.
[0094] As shown in equation (23), the lower bound logJ of the expected gain of GPSTPS is L This can be expressed as the sum of terms related to action policies and terms related to execution policies. The variational distribution is p old Since (d) is used, the integral of episode d can be solved (more precisely, approximated) by the Monte Carlo method using trial-and-error data. Note that the trial-and-error data is the policy π during training. old (a t |s t This is the data obtained when the task was executed using ). This trial-and-error data is used for policy π. old (a t |s t Distribution of episode data p when using ) old Using the sample from (d), we perform an approximation using the Monte Carlo method. This leads to the following equation (25).
[0095]
number
[0096] Here,
[0097]
number
[0098] Furthermore, W is a weight based on the gain,
[0099]
number
[0100] Equation (25) represents the payoff-weighted marginal likelihood. Equation (25) can also be described as weighting the actions and execution lengths selected by the policy model under training based on the sum of rewards (payoffs). Furthermore, equation (25) can be seen as a formalization of the reinforcement learning problem through policy search as a weighted supervised learning problem.
[0101] Variational learning, like the EM algorithm, seeks to maximize the variational distribution given by equation (25).
[0102]
number
[0103] Step E (Expectation step) to find the kernel parameter θ and pseudo-input
[0104]
number
[0105] The policy is optimized by repeating the M-step (Maximization step) which optimizes the distribution. The variational distribution obtained through optimization approximates the posterior distribution. The pseudo-input is the distribution in equation (25).
[0106]
number
[0107] These are the parameters that it possesses.
[0108] More specifically, in step E, one of the two variational distributions is fixed while the other is updated to its optimal value. By changing which distribution is fixed during iterations, both variational distributions can be updated to their optimal values. In step M, both the kernel parameters and pseudo-inputs can be optimized using gradient-based methods. The optimized pseudo-inputs are used to calculate the kernel gram matrix shown in equations (16) and (17) above.
[0109] By using the kernel parameters θ and posterior distribution obtained through variational learning, the unknown state s * Action a against * and execution length τ * The predictive distribution can be analytically determined. Specifically, the above predictive distribution can be expressed by the following formula.
[0110]
number
[0111] Therefore, formulas (26) to (29) above should be stored in the memory unit 11 as action strategy 111, and formulas (30) to (33) above should be stored in the memory unit 11 as execution strategy 112. As a result, the action decision unit 101 uses formulas (26) to (29) stored in the memory unit 11 as action strategy 111 to determine the unknown state s * Action a against * The execution length determination unit 102 can determine the unknown state s using the mathematical formulas (30) to (33) stored in the storage unit 11 as the execution length policy 112. * Execution length τ * It is possible to make a decision.
[0112] As described above, GPSTPS infers the optimal policy, which is an unobservable variable, using observable variables (state, action, and reward) obtained when the task is executed. By using the optimal policy inferred in this way (specifically, the action policy 111 and the execution length policy 112), even in an unstable environment, a reasonable control content can be determined from the acquired observation results, and stable control of the crane 5 can be achieved.
[0113] Furthermore, as described above, a Gaussian process may be applied to the policy model learned by the information processing device 1 (specifically, the action policy 111 and the execution length policy 112). When a Gaussian process policy model is used, high-dimensional features of the data can be implicitly handled by the kernel trick, so a nonlinear policy can be learned with a small amount of data. Also, as mentioned above, a sparse Gaussian process may be applied as the above policy model, which can further reduce the computational complexity of the Gaussian process.
[0114] [Other examples of policy models] Policy models are not limited to Gaussian processes. For example, neural networks can also be applied as policy models. As mentioned above, the reinforcement learning problem using policy search can be formulated as a weighted supervised learning problem. Therefore, when applying a neural network as a policy model, the policy can be learned by using weighted data in a supervised learning manner. The weights here determine the importance of the data in learning according to the level of the reward. In other words, by performing learning that emphasizes data with large weights, i.e., data with high rewards, it is possible to generate a policy model that selects actions with high rewards. Specifically, a policy model can be generated by training a neural network based on the cost of the least squares method multiplied by the above weights.
[0115] [Examples of behavior, performance level, and reward functions] As described above, in the information processing device 1, the action decision unit 101 determines an action according to the observed data using the action policy 111, and the execution length determination unit 102 determines the execution length of the action using the execution length policy 112. By determining the action and its execution length independently in this way, once the execution length has been determined, it is not necessary to determine the execution length again until the action of that execution length is completed. This makes it possible to minimize the number of action decisions and also reduces the amount of computation required for action decisions. Furthermore, if the observed data used to determine the policy is unstable, minimizing the number of action decisions is more appropriate overall than making frequent changes to the action in response to fluctuations in the observed data. In addition, the action policy 111 and the execution length policy 112 have the advantage of requiring less data for training.
[0116] The action predicted using the action policy 111 and the execution length predicted using the execution length policy 112 should be appropriate to the task to be performed by the crane 5. Similarly, the reward r described above t To achieve this, one should set a reward function that corresponds to the task to be performed by crane 5. Here, an example of setting the actions and execution length when crane 5 is tasked with scattering waste, and an example of a reward function to evaluate those actions and execution lengths, are explained based on Figure 4. Figure 4 is a diagram illustrating the overview of the task of scattering waste.
[0117] The task of scattering the waste (hereinafter referred to as scattering and agitation) is performed for the purpose of homogenizing the waste stored in pit P. In scattering and agitation, as shown in the upper part of Figure 4, first, the waste is picked up by crane 5 at a certain point in pit P. If the waste is not picked up sufficiently at this time, it may be released and picked up again. After the waste has been picked up, as shown in the lower part of Figure 4, the waste is scattered along the bucket's movement path by opening and closing the bucket while moving it horizontally.
[0118] Note that a series of tasks is defined from the start of the operation of the crane 5 until the end condition is reached. For example, the bucket is lowered to a predetermined position within the pit P, waste is grasped at that position, the waste is lifted, horizontal movement and opening / closing operations of the bucket are started, and when the moving distance of the crane 5 reaches a predetermined distance, the bucket is opened until no waste remains in the bucket. This can be regarded as a series of tasks.
[0119] In this task, if the amount of waste that can be grasped in the first grasping operation is small, only a small amount of waste can be scattered. By repeating the grasping operation multiple times, the possibility of grasping more waste increases, but if the number of repetitions is too large, it results in time loss. Also, if the number of opening / closing operations is too large or too small, even spreading cannot be achieved. Thus, the task of spreading waste is a highly difficult task.
[0120] The action a to be executed by the crane 5 in this task t can be classified into a grasping operation for grasping waste and an opening / closing operation for opening and closing the bucket to scatter the waste. Therefore, the action a in this task t can be represented by a binary value such as a t ={0,1}. In FIG. 4, a t for the grasping operation is set to 0, and a t for the scattering operation is set to 1. Also, the execution length τ t of the action a in this task t may be the number of times the operation is executed or the trial continuation time. For example, the execution length τ t of the grasping operation is the number of times the grasping operation is executed, and the execution length τ t of the scattering operation may be the number of repetitions of a series of opening / closing operations of opening the bucket to a predetermined angle and then closing it.
[0121] In this case, the learning unit 105 generates a policy model that includes an action policy 111 for determining whether the action to be performed by the crane 5 is a grasping operation to grasp the waste with the crane 5's bucket, or an opening and closing operation to scatter the grasped waste along the bucket's movement path, and an execution length policy for determining the number of times the grasping operation and opening and closing operation are performed. This makes it possible to generate a policy model that can accurately perform the difficult task of scattering waste.
[0122] The above example is merely one example of setting actions and execution lengths; any action and execution length can be set as the prediction target depending on the task to be performed by the crane 5. For example, a series of gripping control rules may be set as an "action," or a simple gripping operation may be set as an "action." Examples of a series of gripping controls include controls that combine multiple operations, such as winding up the bucket while closing it by a fixed value, or opening and closing the bucket while moving from one end of the pit to the other. Examples of simple gripping operations include controls consisting of a single operation, such as completely closing the bucket or closing the bucket by a fixed value.
[0123] In the above-described scattered agitation, the reward function r is, for example, as shown in equation (34) below. t You may apply this.
[0124]
number
[0125] r in equation (34) a This evaluates the performance of distributing waste, and should be defined so that it is a large value when, for example, the waste is distributed evenly or when a large amount of waste is distributed. τ This evaluates the time from the start to the end of a task, and should be defined so that a larger value is given when the task is completed in a shorter time. Note that in the above formula (34), a tWhen the value is 0, that is, when a grasping action is performed, the reward is set to zero, but it would also be possible to provide a reward for the grasping action.
[0126] The above r a and r τ , for example, can be defined as shown in equations (35) and (36) below. Note that α in equation (35) and β in equation (36) are parameters of the reward function, and the values of the reward function parameters can be set as appropriate. Also, u in equation (36) act This is the time required for scattering and mixing (the time from the start to the end of the task). Also, u min This represents the minimum time required for the scattering and agitation process (the time from the start to the end of the task).
[0127]
number
[0128] According to equation (35), the effective length τ t is state s t It is similar to the previous example, where the reward is higher when a large amount of waste is scattered. Also, according to equation (36), the shorter the time from the start to the end of the task, the higher the reward. Therefore, the reward function r in equation (34) t In this case, the r shown in equations (35) and (36) a and r τ By applying this definition, it is possible to learn action strategies and implementation strategies that can scatter large amounts of waste in a short amount of time.
[0129] Also, for example, in formula (34), r a and r τ It may also be defined as shown in the following equations (37) and (38). Note that equation (38) is the same as equation (36) above, so its explanation is omitted.
[0130]
number
[0131] In equation (37), γ is a parameter of the reward function. Also, in equation (37), w max m is the weight of the waste that was picked up (the weight before it was scattered). And in formula (37), m is the weight of the waste that was being held in the bucket of crane 5 during the scattering of the waste. I m is the ideal weight of the waste held in the bucket of crane 5 during waste scattering. As mentioned above, uniform scattering is required in scattering and mixing, I This value should be one that decreases linearly over time. Note that RMS stands for Root Mean Square. In other words, according to equation (37), the more waste is picked up, and the closer the weight change of the waste held in crane 5's bucket during scattering is to linear (the smaller the error from the ideal weight change pattern), the larger the reward value.
[0132] The above r a and r τ If we define it as shown in equations (37) and (38), we can include time-series data of the weight of the waste held in the bucket in the observation data. The weight of the waste held in the bucket can be measured, for example, by attaching a weighing scale to the wire that suspends the bucket.
[0133] Therefore, the reward function r in equation (34) t In this case, the r shown in equations (37) and (38) a and r τ By applying this definition, it is possible to learn action strategies and implementation strategies that can evenly distribute large quantities of waste in a short amount of time.
[0134] As described above, when the task in question is to scatter waste along the movement path of a bucket, the reward given for an action during learning by the learning unit 105 may be a larger value in at least one of the following cases: when a larger amount of waste is scattered, or when the waste is scattered more evenly along the movement path. In addition, the reward given for execution length during learning by the learning unit 105 may be a larger value the shorter the time it takes to complete the task. This makes it possible to determine a strategy that can scatter a larger amount of waste in a shorter time, or a strategy that can scatter waste more evenly in a shorter time.
[0135] [Processing flow (overall)] The flow of processing (learning method) performed by the information processing device 1 will be explained based on Figure 5. Figure 5 is a flowchart showing an example of the processing performed by the information processing device 1 when learning the action policy and execution policy.
[0136] In S1, a trial of the task is performed. As will be explained in detail with reference to Figure 6, in S1, the action and execution length are determined using the action policy 111 and execution length policy 112 that are being learned, and the crane 5 is controlled according to these decisions, and a predetermined task is executed. The content of the task is arbitrary; for example, a task such as scattering waste along the bucket's movement path, as shown in Figure 4, may be performed.
[0137] In S2, the episode collection unit 104 records various information such as a series of actions and execution length in the task performed in S1 as episode 113 in the storage unit 11. The specific contents of episode 113 are as shown in equation (8).
[0138] In S3, the learning unit 105 determines whether a predetermined number of episodes 113 have been recorded in the storage unit 11. Here, the predetermined number is the number of episodes 113 used to update the policy. That is, if the policy is updated using M episodes (where M is a natural number), the predetermined number is M. If the result in S3 is NO, the process returns to S1. On the other hand, if the result in S3 is YES, the process proceeds to S4.
[0139] In S4, the learning unit 105 updates the policies. Specifically, the learning unit 105 updates the action policy 111 and the execution length policy 112 using the formula (25) described above, with respect to a predetermined number (M) of collected episodes 113.
[0140] As described above, the process of updating the action policy 111 and the execution length policy 112 is a process of finding the optimal kernel parameters and variational distribution (i.e., the one that maximizes the lower bound logJL of the expected gain) by repeatedly performing an E step to find the variational distribution that maximizes the value of equation (25) and an M step to optimize the kernel parameters and pseudo-input. Episode 113 recorded in S2 is used as the initial value of the pseudo-input in this process. The learning unit 105 may also use other episodes that have obtained high gains up to the learning point for the above update.
[0141] In S5, the learning unit 105 determines whether the predetermined number of updates has been completed. If the result in S5 is NO, the process returns to S1. The predetermined number of updates can be set in advance. On the other hand, if the result in S5 is YES, the process in Figure 5 is completed, and the final updated action policy 111 and execution length policy 112 become the learned policy model.
[0142] As described above, the learning unit 105 may generate a policy model by repeatedly updating the action policy 111 and the execution length policy 112. This update is performed in such a way that the sum of the rewards given for each action determined using the action policy 111 and the rewards given for each execution length determined using the execution length policy 112 for each action is maximized. This makes it possible to generate a policy model that can determine the optimal action and its execution length for a given task.
[0143] [Processing flow (task trial)] Figure 6 is a flowchart detailing the process of S1 in Figure 5, i.e., the process performed during task trials. Note that the process of S1 in Figure 5 is performed during learning, but the determination of actions and execution length using the learned policy model is the same as the process described below. The process in Figure 6 can also be described as a control determination method that determines the control content of crane 5, or a control method that automatically controls crane 5.
[0144] In S11, the learning action policy 111 and execution length policy 112 are read. Specifically, the action decision unit 101 reads the action policy 111, and the execution length decision unit 102 reads the execution length policy 112. As described above, the action policy 111 is expressed by formulas (26) to (29), and the execution length policy 112 is expressed by formulas (30) to (33).
[0145] In S12, the update variable and state are initialized. Specifically, the action decision unit 101 sets the update variable o1 to 1, the episode collection unit 104 sets the state to the initial state s1, and the state probability to the initial state probability p(s1).
[0146] In S13, the action decision unit 101 updates the variable o t Determine whether o1 is 1 or not. If the result in S13 is YES, proceed to S14. For example, immediately after initialization, t is 1 and o1 is 1, so proceed to S14. On the other hand, if the result in S13 is NO, proceed to S19.
[0147] In S19, the action decision unit 101 decides to continue the previous action. The execution length decision unit 102 then determines the execution length τ t Decrease by 1. In other words, the execution length determination unit 102 determines τ t to τ t-1 Update to -1. This post-processing proceeds to S15.
[0148] In S14, the action and execution length are determined. Specifically, the action decision unit 101 uses the action policy 111 read in S11 to determine the current state s t The execution length determination unit 102 uses the execution length policy 112 read in S11 to determine the action to be taken in the current state s t The execution length is determined accordingly.
[0149] In S15, the action decision unit 101 updates the variable o t to o t+1 It will be updated to τ. Specifically, the action decision unit 101 will t If it is 1 then o t+1 Let τ be 1. t If it is not 1 then o t+1 Let this value be 0.
[0150] In S16, the crane control unit 103 causes the crane 5 to perform the action determined in S14. For example, if the action determined in S14 is a grasping operation to grasp waste, the crane control unit 103 causes the crane 5 to perform that operation via the crane control device 3.
[0151] In S17, the episode collection unit 104 acquires observational data after the action was performed in S16 and sets the state s t to s t+1 Update to state s. For example, the weight of the waste being gripped by the bucket of crane 5 at time t. t In this case, the episode collection unit 104 acquires the weight of the waste held by the bucket as observational data after the action in S16 is performed.
[0152] In S18, the episode collection unit 104 determines whether the task execution has finished. The task completion conditions can be predetermined. If the result in S18 is YES, the process in Figure 6 ends; if the result in S18 is NO, the process returns to S13.
[0153] As described above, the learning method executed by the information processing device 1 includes a data acquisition step (S17 in Figure 6) for acquiring observational data observed at a waste treatment facility equipped with a crane 5 for transporting waste, and a learning step (S4 in Figure 5) for generating a policy model that includes an action policy 111 for determining the action to be performed by the crane 5 and an execution length policy 112 for determining the execution length of the said action, using the observational data. Therefore, it is possible to generate a highly accurate policy model that can be used for controlling the crane 5 at the waste treatment facility. Furthermore, by using this policy model, it becomes possible to determine appropriate control content from the acquired observational results, even in an unstable environment.
[0154] Furthermore, since this learning method uses observational data observed at a waste treatment facility, it can generate a policy model that is optimal for that waste treatment facility. Moreover, according to the above learning method, it is also possible to generate a policy model that corresponds to the quality of the waste at the time the observational data used for learning was observed. For example, if learning is performed using observational data observed when transporting highly adhesive waste, a policy model suitable for controlling the transport of highly adhesive waste by crane 5 will be generated. On the other hand, if learning is performed using observational data observed when transporting less adhesive waste, a policy model suitable for controlling the transport of less adhesive waste by crane 5 will be generated. In this way, by creating multiple policy models according to the quality of the waste in advance, it is possible to apply the policy model that corresponds to the quality of the waste at that time from among these multiple policy models to determine a highly appropriate control content.
[0155] [Variation] The entities that execute each process described in the above-described embodiment are arbitrary and are not limited to the examples above. For example, in the above-described embodiment, one information processing device 1 performs the learning of the action policy 111 and the execution length policy 112, and the control of the crane 5 using the learned action policy 111 and the execution length policy 112, but these processes may be performed by separate devices.
[0156] [Examples of implementation using software] The function of the information processing device 1 (hereinafter referred to as "the device") is a learning program for causing the device to function as a computer, and this can be realized by a learning program for causing each control block of the device (particularly each part included in the control unit 10) to function as a computer.
[0157] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the learning program. By executing the learning program using this control device and storage device, the functions described in each of the embodiments are realized.
[0158] The learning program described above may be recorded on one or more computer-readable recording media, rather than on a temporary basis. These recording media may or may not be provided by the device. In the latter case, the learning program may be supplied to the device via any wired or wireless transmission medium.
[0159] Similarly, a control decision program can be used to cause the computer to function as an action decision unit 101 that determines an action to be performed by the crane 5 using a learned action policy 111, and an execution length decision unit 102 that determines the execution length of the above action using a learned execution length policy 112, thereby realizing the function of determining an action and execution length from observation data, which is one of the functions of the information processing device 1.
[0160] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.
[0161] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Explanation of Symbols]
[0162] 1. Information Processing Device 101 Decision-Making Department 102 Executive Director Decision Department 105 Learning Department 111 Action Plan 112 Implementing Long-Term Strategies 3. Crane control device 5 Cranes 7. Crane control system
Claims
1. The system includes a learning unit that generates a policy model using observational data observed at a waste treatment facility equipped with a crane for transporting waste, which includes an action policy for determining the action to be performed by the crane and an execution length policy for determining the execution length of the said action. The learning unit generates the action policy and the execution length policy by searching for a policy using an update variable to switch between continuing to execute the action determined by the action policy or updating the action and execution length based on the execution length determined by the execution length policy.
2. The information processing apparatus according to claim 1, wherein the learning unit generates the policy model by repeatedly updating the action policy and the execution length policy so as to maximize the sum of the rewards given for each action determined using the action policy and the rewards given for each execution length determined using the execution length policy for each action until a predetermined task is completed.
3. The information processing apparatus according to claim 1 or 2, wherein a Gaussian process is applied as the policy model.
4. The information processing apparatus according to any one of claims 1 to 3, wherein the learning unit generates a policy model that includes the action policy for determining whether the action to be performed by the crane is a grasping operation in which the crane grasps the waste with the crane's bucket, or an opening and closing operation in which the bucket is opened and closed in order to scatter the grasped waste along the bucket's movement path, and the execution length policy for determining the number of times the grasping operation and the opening and closing operation are performed.
5. The learning unit generates the policy model by repeatedly updating the action policy and the execution length policy so as to maximize the sum of the rewards given for each action determined using the action policy and the rewards given for each execution length determined using the execution length policy for each action, up to the completion of the task of scattering the waste along the movement path of the bucket. The information processing apparatus according to claim 4, wherein the reward given for an action is greater if a larger amount of waste is scattered, and if the waste is scattered more evenly along the movement path, and the reward given for execution length is greater the shorter the time the task is completed.
6. An action determination unit that determines the action to be performed by a crane transporting waste, using an action strategy for determining the action to be performed by the crane, in accordance with the observation data observed at the waste treatment facility equipped with the crane, The system comprises an execution length determination unit that determines the execution length according to the observed data using an execution length strategy for determining the execution length of the action determined by the action determination unit, The aforementioned action decision unit and the execution length decision unit repeatedly perform action decisions and execution length decisions during the period until a predetermined task is completed. The action decision unit is an information processing device that, when a decided action has been executed for a predetermined time but the predetermined task has not been completed, determines a new action if the action has been executed for the execution length determined by the execution length decision unit at the time the action was decided, and continues the execution of the action if it has not been executed.
7. A crane for transporting waste, The information processing apparatus according to claim 6, A crane control system including a crane control device that operates the crane based on the action and execution length determined by the information processing device.
8. A learning method performed by one or more information processing devices, A data acquisition step involves obtaining observational data observed at a waste treatment facility equipped with a crane for transporting waste, and The process includes a learning step of generating a policy model that includes an action policy for determining an action to be performed by the crane using the aforementioned observation data, and an execution length policy for determining the execution length of said action, The learning method involves generating the action policy and the execution length policy by searching for a policy using an update variable to switch between continuing to execute the action determined by the action policy or updating the action and execution length based on the execution length determined by the execution length policy.
9. A learning program for causing a computer to function as an information processing device according to claim 1, wherein the learning program causes the computer to function as the learning unit.
Citation Information
Patent Citations
Hopper feeding control method and system
CN109941886A
Chokosokuteiyopuropera
JP1976085197A
Waste agitation evaluating method, waste agitation evaluating program, and waste agitation evaluating device
JP2010275064A
Scrap grade determination system, scrap grade determination method, estimation device, learning device, learnt model generation method and program
JP2020095709A
Refuse crane control system
JP2021042873A