Estimation device, estimation method, and program

JPWO2025057366A5Pending Publication Date: 2026-05-15
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2026-02-09
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

When controlling target objects, it is difficult for the prior art to effectively estimate the characteristics of target objects, resulting in insufficient control accuracy and flexibility.

Method used

An estimation device is designed. By setting an evaluation function, the function optimizes the state difference and the next state difference under the control command, and combines the parameter values ​​calculated by simulation to search for the parameter value distribution to achieve a high degree of freedom estimation of the characteristics of the target object.

Benefits of technology

High-precision estimation of the target object characteristics is achieved, control accuracy and flexibility are improved, and the degree of freedom of the estimation process is enhanced.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This estimation device comprises: an evaluation function setting means for setting an evaluation function that demonstrates an evaluation that is more favorable the smaller the size of the difference between a subsequent state of a characteristic estimation object observed under the state of the characteristic estimation object as well as a control command for the characteristic estimation object and a subsequent state of the characteristic estimation object calculated by using a simulation under the aforementioned state and control command as well as set values for parameters indicating the characteristics of the characteristic estimation object; and a distribution calculation means for using the evaluation function to find the distribution of values of the parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Estimation device, estimation method, and recording medium

[0001] The present invention relates to an estimation device, an estimation method, and a recording medium.

[0002] Control of a control target is performed according to the state of the control target. For example, Patent Document 1 describes training an estimation model for a production facility equipped with a plurality of mechanisms including a servo motor, which estimates a correction amount related to the drive of the servo motor from the temperature of the servo motor, the environmental temperature at a position representative of the production facility, and the environmental humidity at a position representative of the production facility.

[0003] International Publication No. 2022 / 176760

[0004] When learning control of a controlled object, it is conceivable to estimate the characteristics of the controlled object and use a model that reflects the estimated characteristics to learn the control. In this way, when estimating the characteristics of the controlled object in relation to the control of the controlled object, it is preferable to have a degree of freedom regarding the estimation.

[0005] An example of an object of the present invention is to provide an estimation device, an estimation method, and a recording medium that can solve the above-mentioned problems.

[0006] According to a first aspect of the present invention, an estimation device includes: evaluation function setting means for setting an evaluation function that indicates a better evaluation the smaller the magnitude of the difference between a state of an object of characteristic estimation and a next state of the object of characteristic estimation observed under a control command for the object of characteristic estimation, and a next state of the object of characteristic estimation calculated by simulation under set values ​​of parameters indicating the state, the control command, and the characteristics of the object of characteristic estimation; and distribution calculation means for searching for a distribution of values ​​of the parameters using the evaluation function.

[0007] According to a second aspect of the present invention, the estimation method includes a computer setting an evaluation function that indicates a better evaluation the smaller the difference between a state of an object of characteristic estimation and a next state of the object of characteristic estimation observed under a control command for the object of characteristic estimation, and a next state of the object of characteristic estimation calculated by a simulation under set values ​​of parameters indicating the state, the control command, and the characteristics of the object of characteristic estimation, and searching for a distribution of values ​​of the parameters using the evaluation function.

[0008] According to a third aspect of the present invention, a recording medium is a recording medium having recorded thereon a program that causes a computer to execute the following steps: set an evaluation function that indicates a better evaluation the smaller the magnitude of the difference between a state of an object of characteristic estimation and a next state of the object of characteristic estimation observed under a control command for the object of characteristic estimation, and a next state of the object of characteristic estimation calculated by a simulation under set values ​​of parameters that indicate the characteristics of the object of characteristic estimation; and search for a distribution of values ​​of the parameters using the evaluation function.

[0009] According to the present invention, when estimating the characteristics of an object of characteristic estimation in relation to control of the object of characteristic estimation, a degree of freedom regarding the estimation can be obtained.

[0010] FIG. 1 is a diagram illustrating a first example of a configuration of an estimation device according to at least one embodiment; FIG. 2 is a diagram illustrating an example of a configuration of a robot system according to at least one embodiment; FIG. 3 is a diagram illustrating a first example of data input / output when an estimation device according to at least one embodiment estimates a probability distribution of values ​​of a parameter ξ; FIG. 4 is a diagram illustrating a first example of a processing procedure by which an estimation device according to at least one embodiment calculates a distribution of values ​​of a parameter ξ; FIG. 5 is a diagram illustrating a second example of a configuration of an estimation device according to at least one embodiment; FIG. 6 is a diagram illustrating a second example of data input / output when an estimation device according to at least one embodiment estimates a probability distribution of values ​​of a parameter ξ; FIG. 7 is a diagram illustrating a second example of a processing procedure by which an estimation device according to at least one embodiment calculates a distribution of values ​​of a parameter ξ; FIG. 8 is a diagram illustrating a third example of a configuration of an estimation device according to at least one embodiment; FIG. 9 is a diagram illustrating a third example of data input / output when an estimation device according to at least one embodiment estimates a probability distribution of values ​​of a parameter ξ; FIG. 10 is a diagram illustrating a third example of a processing procedure by which an estimation device according to at least one embodiment calculates a distribution of values ​​of a parameter ξ; FIG. 10 is a diagram showing a fourth example of a processing procedure by which an estimation device according to at least one embodiment calculates the distribution of values ​​of a parameter ξ. FIG. 11 is a diagram showing a fifth example of the configuration of an estimation device according to at least one embodiment. FIG. 12 is a diagram showing an example of a processing procedure in an estimation method according to at least one embodiment. FIG. 13 is a schematic block diagram showing the configuration of a computer according to at least one embodiment.

[0011] The following describes embodiments of the present invention, but the following embodiments do not limit the scope of the invention. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention. Below, letters with circumflexes may be indicated with a "^" after the letter. For example, a circumflex o may also be written as o^.

[0012] First Embodiment Fig. 1 is a diagram illustrating an example of the configuration of an estimation device 100 according to at least one embodiment. In the configuration illustrated in Fig. 1, the estimation device 100 includes a communication unit 110, a display unit 120, an operation input unit 130, a storage unit 180, and a processing unit 190. The processing unit 190 includes a distribution calculation unit 191, a sampling unit 192, a likelihood calculation unit 193, a posterior probability calculation unit 194, a simulation execution unit 195, and an evaluation function setting unit 196.

[0013] The estimation device 100 estimates a probability distribution of parameter values ​​that indicate the characteristics of a characteristic estimation target. The characteristic estimation target here is not limited to a specific one. The characteristic estimation target by the estimation device 100 can be any object that can be operated in response to a control command, whose state can be measured, and which can be simulated using a model that has parameters that indicate the characteristics and parameters that indicate the state.

[0014] A parameter indicating a characteristic of a characteristic estimation target is denoted as ξ. The estimation device 100 estimates a probability distribution of the value of the parameter ξ. The parameter ξ may be a combination of multiple parameters. For example, the parameter ξ may be represented by a vector.

[0015] In the following, an example will be described in which the estimation device 100 estimates a probability distribution of values ​​of physical parameters of an object to be moved in a task in which a robot moves the object, such as the mass, the position of the center of gravity, and the friction coefficient of the bottom surface of the object. In this case, the operating environment of the robot, including the robot and the object to be moved, corresponds to an example of a target for characteristic estimation. The physical parameters of the object, such as the mass, the position of the center of gravity, and the friction coefficient of the bottom surface of the object, correspond to an example of the parameter ξ.

[0016] For example, consider a task in which a robot pushes an object across a table, moving it to a desired position and orientation. The action of a robot hand pushing an object is also called a pushing action. In this case, the location and strength of the push on the object to move it as intended may differ depending on the object's mass, the position of its center of gravity, the coefficient of friction of its bottom surface, etc.

[0017] Fig. 2 is a diagram showing an example of the configuration of a robot system 910. In the configuration shown in Fig. 2, the robot system 910 includes a control device 911, a robot 912, a sensor 913, and a data generation device 914. Fig. 2 also shows an environment 921 and an object 931.

[0018] The object 931 is an object for which the probability distribution of parameter values ​​is to be estimated by the estimation device 100. The object 931 may be any object that can be moved by the robot 912, and is not limited to a specific type. Furthermore, a plurality of objects 931 may be prepared. In this case, the parameter values ​​to be estimated by the estimation device 100 may differ for each object 931. The object 931 may be configured as a part of the robot system 910, or may be configured external to the robot system 910.

[0019] The robot 912 operates in response to input of a control command. For example, in a task in which the robot 912 moves an object 931, the robot 912 moves the object 931 by operating in accordance with the control command. The robot 912 may be any robot that can move the object 931, and is not limited to a specific type of robot. For example, the robot 912 may be configured as a vertical articulated robot, or as a mobile robot equipped with a mechanism for loading and unloading the object 931 onto and from the robot 912 itself.

[0020] The environment 921 is an operating environment of the robot 912. The environment 921 is assumed to include the robot 912 and an object 931. The control device 911 controls the robot 912 by inputting a control command to the robot 912. The control device 911 may be configured using a computer. Furthermore, the control device 911 may be configured as a part of the estimation device 100, or may be configured as a device external to the estimation device 100.

[0021] The sensor 913 observes the state of the environment 921. Specifically, the sensor 913 observes values ​​related to the state of the environment 921 and outputs sensor measurement values ​​to the data generating device 914. The sensor 913 may include multiple sensors. Furthermore, the sensor 913 may include multiple types of sensors.

[0022] The data generating device 914 generates data indicating a state transition of the environment 921 under a control command to the robot 912, based on sensor measurements by the sensor 913. The data generating device 914 may be configured using a computer. Furthermore, the data generating device 914 may be configured as a part of the estimating device 100, or may be configured as a device external to the estimating device 100. The state of the environment 921 is also simply referred to as a state. Data indicating a state transition of the environment 921 under a control command to the robot 912 is also referred to as transition data.

[0023] In the following, time will be expressed in terms of time steps. The state at time t is represented as s t and the control command to the robot 912 at time t is expressed as a t Also, the state s t A control command a is sent to the robot 912 under t When input, the next state (state at time t+1) is s t+1 It is written as follows.

[0024] As the transition data o, the state s t and control command a t and the next state s t+1 and are combined (s t , a t , s t+1 ) is used. t , a t , s t+1 A set of transition data o is also called a transition data set, and D o It is written as follows.

[0025] Transition Data Set D oThe transition data included in the transition data set D may be continuous in time or may be obtained at discrete times. o The next state s in one transition data o included in t+1 The time t+1 of the transition data set D o The state s in another transition data o included in t When the time t of the transition data o is equal to the time t of the transition data o, these two transition data o are said to be continuous in time. Two transition data o that are not continuous in time are called transition data o obtained at discrete times.

[0026] The estimation device 100 determines whether the transition data o indicates a state s t and control command a t The next state obtained by the simulation is called s^ t+1 The predicted data (s t , a t , s t+1 ) and the next state s^ obtained by simulation t+1 Data including (s t , a t , s t+1 , s^ t+1 ) is also called predicted transition data and is expressed as o^. t , a t , s t+1 , s^ t+1 ) can also be written as

[0027] The set of predicted transition data o is also referred to as a predicted transition data set, D o^ The estimation device 100 calculates the next state s indicated by the transition data o. t+1 and the next state s^ obtained by simulation t+1 The probability distribution of the parameter ξ is calculated by solving an optimization problem using an evaluation function that indicates a better evaluation the smaller the difference between the two.

[0028] The communication unit 110 communicates with other devices. For example, the communication unit 110 receives the transition data set D from the data generating device 914. oThe display unit 120 may be configured to receive the parameter ξ estimated by the estimation device 100. The display unit 120 may include a display screen such as a liquid crystal panel or an LED (Light Emitting Diode) panel, and display various images. For example, the display unit 120 may display the probability distribution of the parameter ξ estimated by the estimation device 100.

[0029] The operation input unit 130 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the operation input unit 130 may receive user operations for specifying a task to be performed by the robot 912, or other user operations for making various settings related to the estimation performed by the estimation device 100.

[0030] The storage unit 180 stores various data. For example, the storage unit 180 stores a transition data set D o and predicted transition dataset D o^ The storage unit 180 may be configured using a storage device included in the estimation device 100. The processing unit 190 controls each unit of the estimation device 100 to perform various processes. The functions of the processing unit 190 may be performed by a CPU (Central Processing Unit) included in the estimation device 100 reading and executing a program from the storage unit 180.

[0031] The distribution calculation unit 191 calculates the probability distribution of the value of the parameter ξ. Specifically, the distribution calculation unit 191 calculates the probability distribution of the value of the parameter ξ using the parameter distribution model template p * θ By setting the value of the parameter θ in (ξ), the probability distribution of the value of the parameter ξ is obtained. The distribution calculation unit 191 corresponds to an example of a distribution calculation means.

[0032] Here, the parameter distribution model template p * θ (ξ) is a template of the probability distribution assumed as the probability distribution of the value of the parameter ξ. * θ The "θ" in "(ξ)" indicates the parameter of the parameter distribution model template. * θThe "ξ" in "(ξ)" indicates the parameter for which the parameter distribution model template shows the probability distribution of values. * θ The probability distribution obtained by setting a value for the parameter θ in (ξ) is called p θ This is written as (ξ).

[0033] Parameter distribution model template p * θ (ξ) is not limited to a specific one. For example, the parameter distribution model template p * θ As (ξ), a probability density function such as a Gaussian distribution, a mixed Gaussian distribution, or a Poisson distribution may be used. * θ When the probability density function of a Gaussian distribution is used as (ξ), the mean μ and variance σ of the Gaussian distribution are 2 can be used as the parameter θ.

[0034] Alternatively, the parameter distribution model template p * θ A probability distribution using a generative model such as a flow-based model may be used as (ξ). * θ (ξ) may be set in advance according to the characteristics indicated by the parameter ξ. When the parameter ξ is configured by a combination of multiple parameters, different types of parameter distribution model templates may be used for each parameter, or the same parameter distribution model template may be used for each parameter.

[0035] As described above with respect to the estimation device 100, the distribution calculation unit 191 calculates the next state s indicated by the transition data o. t+1 and the next state s^ obtained by simulation t+1The distribution calculation unit 191 solves an optimization problem using an evaluation function that indicates a better evaluation the smaller the difference between the parameter ξ and the parameter ξ, and obtains a probability distribution of the parameter ξ. Specifically, the distribution calculation unit 191 searches for a value of the parameter ξ that indicates a better evaluation as possible in the optimization problem.

[0036] The sampling unit 192 calculates the probability distribution p θ M values ​​of the parameter ξ are sampled according to (ξ), where M is an integer greater than or equal to 1. The value of M may be predetermined. The set of sampled values ​​of the parameter ξ is also denoted as Ξ. The M sampled parameter values ​​are then calculated as ξ. 1 , ξ 2 , ..., ξ M It can also be written as:

[0037] The sampling unit 192 samples M values ​​of the parameter ξ every time the distribution calculation unit 191 sets a value of the parameter θ. The set Ξ is overwritten every time the distribution calculation unit 191 sets a value of the parameter θ. In other words, the set Ξ is a set of the most recent M sampled values ​​sampled by the sampling unit 192.

[0038] The likelihood calculation unit 193 calculates the probability distribution p of the value of the parameter ξ based on the sampled value of the parameter ξ. θ Specifically, the likelihood calculation unit 193 calculates the likelihood of the M sampled parameter values ​​ξ 1 , ξ 2 , ..., ξ M For each, the probability distribution p θ Calculate the likelihood of the value of the parameter θ in (ξ). Let the likelihood of the value of the parameter θ under the value of the parameter ξ be l θ The likelihood calculation unit 193 may use a known likelihood calculation method to calculate the likelihood distribution.

[0039] The posterior probability calculation unit 194 calculates the posterior probability p(o|ξ) of the transition data o under the value of the parameter ξ. For example, the posterior probability calculation unit 194 may calculate p(o|ξ) based on Equation (1).

[0040]

[0041] exp represents the power of Napier's constant e. γ is a constant that takes a positive real value. The value of γ may be determined in advance. l(o^, ξ) is expressed as in equation (2).

[0042]

[0043] "||s t+1 -s^ t+1 ||" is the next state s indicated in the transition data. t+1 and the next state s^ calculated by simulation t+1 Indicates the magnitude of the difference between ||s t+1 -s^ t+1 ||" is the measured next state s t+1 and the next state s^ calculated by simulation t+1 This can be seen as an indication of the magnitude of the error.

[0044] s t+1 and s^ t+1 The smaller the difference between the two, the larger the value of p(o|ξ). In this respect, equation (1) can be considered to indicate the posterior probability of obtaining transition data o under the value of parameter ξ. The posterior probability p(o|ξ) is also referred to as the transition data posterior probability. The template for calculating the posterior probability p(o|ξ), as exemplified in equation (1), is also referred to as the transition data posterior probability model template, and p * (o|ξ). The transition data posterior probability model template p * (o|ξ) can be considered as a template of the probability distribution assumed as the posterior probability p(o|ξ).

[0045] The simulation execution unit 195 performs a simulation of the operation of the robot 912. Specifically, as described above with respect to the estimation device 100, the simulation execution unit 195 performs a simulation of the operation of the robot 912. t and control command a t The operation of the robot 912 is simulated under t+1 The simulation execution unit 195 calculates the next state s^ by the simulation.t is used by the posterior probability calculation unit 194 to calculate the posterior probability p(o|ξ).

[0046] The simulator used by the simulation execution unit 195 is t and control command a t and the parameter ξ, the next state s t+1 The simulator may be any of various types that can calculate the above. In particular, the simulator used by the simulation execution unit 195 may be one in which the equation of the simulation model is not analytically obtained, or may be one that is not differentiable. For example, the simulation execution unit 195 may use, as the simulator, a physical simulator such as MuJoCo or Isaac Sim, or a transition model that has been trained by a machine learning technique, but is not limited to these.

[0047] The evaluation function setting unit 196 sets an evaluation function to be used in the optimization problem for which the distribution calculation unit 191 determines the probability distribution of the value of the parameter ξ. This evaluation function is also referred to as an evaluation function for updating the parameter value. The evaluation function setting unit 196 corresponds to an example of evaluation function setting means.

[0048] The evaluation function setting unit 196 calculates the likelihood l of the value of the parameter θ under the value of the parameter ξ. θ (ξ) and the posterior probability p(o|ξ) of the transition data o under the value of the parameter ξ, an evaluation function for parameter update is set. As the evaluation function here, for example, the evaluation function L shown in equation (3) may be used.

[0049]

[0050] N data is the transition data set D o Transition data o included in i N indicates the number of data i is an integer ≧1. i That is, i is an index variable that identifies the transition data o included in the transition data set Do. i i is an index variable that identifies the dataThe transition data o is assigned an index i, and the i It is written as follows.

[0051] E ξ~pθ(ξ) [logp(o i |ξ)logl θ (ξ)] is the parameter ξ value in the probability distribution p θ (ξ) ji |ξ)logl θ (ξ) where j is the parameter value ξ j That is, j is an index variable that identifies the parameter value ξ included in the set Ξ. j j is an index variable that identifies the sampled parameter value ξ. j takes an integer value of 1≦j≦M. As mentioned above, M is the number of sampled parameter values ​​ξ j M is an integer of 1 or more, indicating the number of

[0052] Next state s indicated in the transition data t+1 and the next state s^ calculated by simulation t+1 The smaller the difference between i |ξ j ) becomes smaller and the value of the evaluation function L becomes larger.

[0053] The larger the value of the evaluation function L, the better the evaluation. t+1 and the next state s^ calculated by simulation t+1 This corresponds to an example of an evaluation function that indicates a better evaluation as the magnitude of the difference between i and the value of ξ j By setting the values ​​of θ and θ, the evaluation function L as a function of θ is set.

[0054] However, the evaluation function set by the evaluation function setting unit 196 is not limited to a specific one. t+1 and the next state s^ calculated by simulation t+1 Various evaluation functions can be used that indicate a better evaluation when the magnitude of the difference between the two is smaller.

[0055] The evaluation function is the next state s indicated in the transition data. t+1 and the next state s^ calculated by simulation t+1 In addition to the evaluation criterion that the difference between the evaluation function (i.e., the value of the evaluation function) and the evaluation function (ii.e., the value of the evaluation function) is small, other evaluation criteria may be set. For example, the evaluation function setting unit 196 may set the evaluation function based on the formula (4) instead of the formula (3).

[0056]

[0057] In equation (4), a sub-equation −p is further added to equation (3) to take entropy into consideration. θ (ξ) log p θ (ξ) is provided. The probability distribution p θ Entropy of (ξ) - p θ (ξ) log p θ The larger the value of (ξ), the larger the value of the evaluation function L, indicating a better evaluation.

[0058] The distribution calculation unit 191 calculates the probability distribution p of the value of the parameter ξ using equation (4). θ By estimating (ξ), the estimated probability distribution p θ It is expected that this can prevent excessive fitting of the probability distribution p (ξ) to the data. θ If we consider the estimation of (ξ) as machine learning, the sub-formula -p of formula (4) θ (ξ) log p θ It is expected that (ξ) will have a regularization effect that prevents overfitting.

[0059] 3 is a diagram showing an example of data input and output when the estimation device 100 estimates the probability distribution of the value of the parameter ξ. In the example of FIG. 3, the distribution calculation unit 191 calculates the probability distribution p θ In particular, the distribution calculation unit 191 searches for the value of the parameter θ by optimization calculation using the evaluation function set by the evaluation function setting unit 196, thereby obtaining the probability distribution p θ The distribution calculation unit 191 calculates the set probability distribution p θ(ξ) is output to the sampling unit 192 and the likelihood calculation unit 193 .

[0060] The sampling unit 192 calculates the probability distribution p θ The sampling unit 192 samples the value of the parameter ξ according to (ξ). The sampling unit 192 outputs the sampled value of the parameter ξ to the likelihood calculation unit 193 and the posterior probability calculation unit 194.

[0061] The likelihood calculation unit 193 calculates the probability distribution p based on the value of the parameter ξ sampled by the sampling unit 192. θ The likelihood of the value of the parameter θ in (ξ) θ The likelihood calculation unit 193 calculates the calculated likelihood l θ (ξ) is output to the evaluation function setting unit 196.

[0062] The simulation execution unit 195 calculates the value of the parameter ξ sampled by the sampling unit 192 and the state s indicated by the transition data o. t and control command a t and the simulation is performed using the next state s^ t+1 The simulation execution unit 195 calculates the next state s^ obtained as a result of the simulation. t+1 to the posterior probability calculation unit 194.

[0063] The posterior probability calculation unit 194 outputs the value of the parameter ξ sampled by the sampling unit 192 and the transition data o to the simulation execution unit 195, and calculates the next state s^ calculated by the simulation. t+1 Then, the posterior probability calculation unit 194 obtains the value of the parameter ξ sampled by the sampling unit 192 and the next state s^ calculated by the simulation. t+1 The posterior probability calculation unit 194 calculates the posterior probability p(o|ξ) of the transition data o under the value of the parameter ξ based on the above. The posterior probability calculation unit 194 outputs the calculated posterior probability p(o|ξ) to the evaluation function setting unit 196.

[0064] The evaluation function setting unit 196 calculates the likelihood l calculated by the likelihood calculation unit 193. θ(ξ) and the posterior probability p(o|ξ) calculated by the posterior probability calculation unit 194, the evaluation function setting unit 196 sets an evaluation function L shown in equation (3) or an evaluation function L shown in equation (4). The evaluation function setting unit 196 outputs the set evaluation function to the distribution calculation unit 191. As described above, this evaluation function is used by the distribution calculation unit 191 to search for the value of the parameter θ.

[0065] 4 is a diagram showing an example of a processing procedure in which the estimation device 100 calculates the distribution of the value of the parameter ξ. (Step S101) The distribution calculation unit 191 calculates the probability distribution p θ Specifically, the distribution calculation unit 191 initializes the parameter distribution model template p * θ The parameter θ of (ξ) is set to the initial value θ 0 By setting the parameter ξ, the probability distribution p θ The initial value p of (ξ) θ0 (ξ) to obtain the initial value θ of the parameter θ. 0 may be a predetermined value. After step S101, the process proceeds to step S102.

[0066] (Step S102) The sampling unit 192 calculates the probability distribution p of the parameter ξ set by the distribution calculation unit 191. θ(k-1) The value of the parameter ξ is sampled according to (ξ), where k is an index variable indicating the number of times the loop from steps S102 to S109 is repeated, and k is an integer value where k≧1.

[0067] θ (k-1) indicates the value of the parameter θ used in the k-th repetition of the loop from steps S102 to S109. When k=1, that is, in the first execution of the loop from steps S102 to S109, the value of the parameter θ is set by the distribution calculation unit 191 in step S101. 0 is used.

[0068] If k≧2, in the kth and subsequent executions of the loop from step S102 to S109, the value of the parameter θ is set to the θ obtained in the k−1th execution of the loop from step S102 to S109. (k-1) After step S102, the process proceeds to step S103.

[0069] (Step S103) The likelihood calculation unit 193 calculates the sampled parameter value ξ j For each ∈Ξ, the probability distribution p θ The likelihood p of (ξ) θ (ξ j After step S103, the process proceeds to step S104.

[0070] (Step S104) The posterior probability calculation unit 194 calculates the sampled parameter value ξ j For each ∈Ξ, each transition data o i ∈D O Specifically, the posterior probability calculation unit 194 sets the simulation settings and initial values ​​of the simulator based on the parameter value ξ j is set as a parameter value of the simulation. i The state s shown in t is set as the initial value of the simulation (state at the start of the simulation). i Control command a shown in t is set as a control command to the robot 912 in the simulation. After step S104, the process proceeds to step S105.

[0071] (Step S105) The simulation execution unit 195 executes the set simulation and generates a predicted transition data set D O^ As described above, the predicted transition data set D O^ is the predicted transition data o^ i = (s i,t , a i,t , s i,t+1 , s^ i,t+1 ) where s i,t , ai,t , s i,t+1 , s^ i,t+1 is, s t , a t , s t+1 , s^ t+1 The index i that identifies the transition data o is clearly shown in s^. i,t+1 is the transition data set D O Each transition data o included in i = (s i,t , a i,t , s i,t+1 ) state s i,t Control command a i,t The simulation execution unit 195 executes the transition data set D O All transition data o included in i All predicted transition data o^ obtained from i The predicted transition data set D O^ may be generated.

[0072] In the simulation run, the transition data i = (s i,t , a i,t , s i,t+1 ) state s i,t and control command a i,t and the sampled parameter values ​​ξ j Based on this, the transition data o i For each parameter value ξ j For each next state s^ i,t+1 The simulation performed by the simulation execution unit 195 can be expressed as in equation (5).

[0073]

[0074] In Equation 5, the function f represents the simulation performed by the simulation execution unit 195. After step S105, the process proceeds to step S106.

[0075] (Step S106) The posterior probability calculation unit 194 calculates the transition data posterior probability model template p * (o|ξ) and the predicted transition dataset D O^Based on the sampled parameter value ξ j Transition data under i The posterior probability p(o i |ξ j The posterior probability calculation unit 194 calculates the transition data o i (i.e., predicted transition data o^ i ) and parameter value ξ j For each, the posterior probability p(o i |ξ j ) is calculated.

[0076] As described above, the transition data posterior probability model template p * (o|ξ) is the posterior probability p(o i |ξ j ) is a template of a probability distribution assumed as the transition data posterior probability model template p * (o|ξ) with predicted transition data o^ i The s shown in i,t , a i,t , s i,t+1 , s^ i,t+1 , and the sampled parameter values ​​ξ j , or some of these are input, and the posterior probability p(o i |ξ j After step S106, the process proceeds to step S107.

[0077] (Step S107) The evaluation function setting unit 196 sets the parameter distribution model p θ The likelihood of (ξ) p θ (ξ) and the transition data posterior probability p(o|ξ), an evaluation function for parameter update is set. As described above, the evaluation function setting unit 196 may set the evaluation function L based on equation (3) or the evaluation function L based on equation (4). After step S107, the process proceeds to step S108.

[0078] (Step S108) The distribution calculation unit 191 calculates the probability distribution p θSpecifically, the distribution calculation unit 191 updates the parameter distribution model template p θ The distribution calculation unit 191 searches for the value of the parameter θ in (ξ). The distribution calculation unit 191 performs optimization calculations so that the evaluation function value indicates as good an evaluation as possible, and searches for the value of the parameter θ. For example, the distribution calculation unit 191 may use a gradient descent method to perform a solution search in the optimization calculations.

[0079] However, the distribution calculation unit 191 uses the parameter distribution model p θ The method used to update (ξ) is not limited to a specific one. The distribution calculation unit 191 updates the value of the parameter θ to the value obtained by the solution search, thereby obtaining the probability distribution p θ After step S108, the process proceeds to step S109.

[0080] (Step S109) The distribution calculation unit 191 determines whether or not it is necessary to continue learning. The condition for terminating the learning here is not limited to a specific one. For example, the condition for terminating the learning here may be whether the number of iterations of the loop from steps S102 to S109 has reached a predetermined number, or whether the norm of the gradient of the evaluation function L with respect to θ, ∥∇ θ L|| is the threshold ε L The present invention may be, but is not limited to, the following:

[0081] If the distribution calculation unit 191 determines that continuation of learning is necessary (step S109: YES), the process returns to step S102. In this case, the estimation device 100 executes the process in the loop from step S102 to S109 again using the updated parameter distribution model obtained in step S108. On the other hand, if the distribution calculation unit 191 determines that continuation of learning is unnecessary (step S109: NO), the estimation device 100 ends the process in FIG. 4.

[0082] As described above, the evaluation function setting unit 196 sets an evaluation function that indicates a better evaluation the smaller the difference between the next state of the environment 921 observed under the state of the environment 921 and the control command to the robot 912, and the next state of the environment 921 calculated by simulation under the set values ​​of the state of the environment 921, the control command to the robot 912, and the parameter ξ indicating the characteristics of the object 931. The distribution calculation unit 191 calculates the distribution p of the value of the parameter ξ using the evaluation function set by the evaluation function setting unit 196. θ Search for (ξ).

[0083] According to the estimation device 100, the evaluation function set by the evaluation function setting unit 196 can be provided with a sub-expression for preventing excessive fitting to data, for example, and the distribution p of the parameter ξ calculated by the distribution calculation unit 191 can be calculated. θ In this respect, the estimation device 100 provides degrees of freedom in the search for (ξ) when estimating the characteristics of the object 931 in relation to the control of the robot 912.

[0084] The evaluation function setting unit 196 also sets the distribution p of the parameter ξ values. θ The sampling value of the parameter ξ according to (ξ) j The posterior probability p(o|ξ) of the state of the environment 921, the control command for the robot 912, and the transition data o indicating the next state of the environment 921 under j The likelihood l of the probability distribution of the value of the parameter ξ based on θ (ξ) and an evaluation function based on it.

[0085] According to the estimation device 100, the posterior probability p(o|ξ) and the likelihood l θ Based on (ξ), the next state s^ calculated by simulation t+1 The value of the observed next state s t+1 According to the estimation device 100, at this point, the posterior probability p(o|ξ) is calculated and the likelihood l θ The probability distribution p of the parameter ξ, which indicates the characteristics of the object 931, can be calculated by a relatively simple calculation such as (ξ). θ(ξ) is the observed next state s t+1 The evaluation function can be set to reflect the value of

[0086] Second Embodiment Fig. 5 is a diagram showing an example of the configuration of an estimation device 200 according to at least one embodiment. In the estimation device 200 shown in Fig. 5, a processing unit 290 includes an environmental data acquisition unit 291 in addition to the units included in the processing unit 190 of the estimation device 100 shown in Fig. 1. In other respects, the estimation device 200 is similar to the estimation device 100. Furthermore, a robot system 910 in the second embodiment is similar to that in the first embodiment, and Fig. 2 will also be referred to in the second embodiment.

[0087] The environment data acquisition unit 291 acquires a transition data set Do in the environment 921. Specifically, the environment data acquisition unit 291 observes the state in the environment 921 using a sensor or the like, and controls the robot 912 to acquire the state s t and control command a t and the next state s t+1 Transition data o=(s t , a t , s t+1 The environment data acquisition unit 291 acquires a predetermined number of pieces of transition data o, and collects the acquired transition data o into a set to obtain a transition data set D o The environmental data acquisition unit 291 is an example of an environmental data acquisition unit.

[0088] The environmental data acquisition unit 291 may acquire transition data o that are continuous in time, or may acquire transition data o at discrete times. o From the state of the empty set, the transition data o is acquired and the transition data set D o Alternatively, the environmental data acquisition unit 291 may store the transition data set D o When one or more transition data o are included in the transition data set D, the acquisition of the transition data o is started, and the acquired transition data o is stored as the transition data set D. o It is also possible to add to the above.

[0089] The method of controlling the robot 912 when acquiring the transition data is not limited to a specific method. For example, the environmental data acquisition unit 291 may control the robot 912 to make the robot 912 perform a certain task based on a control rule that the environmental data acquisition unit 291 has learned to make the robot 912 perform a certain task. Alternatively, the environmental data acquisition unit 291 may randomly determine a control command for a control object and control the control object.

[0090] When the environmental data acquisition unit 291 controls the robot 912 so that the robot 912 executes a certain task, the control rules used by the environmental data acquisition unit 291 are not limited to control rules obtained by a specific method.

[0091] For example, the environmental data acquisition unit 291 may control the robot 912 using a control rule obtained by the following learning method using a level set function. Here, a vector indicating information about a task, such as the initial state of task execution, the goal state of the task, and the values ​​of task parameters (parameters set for the task), is called a condition vector, and ξ cond It is written as follows.

[0092] A parameter indicating a control command for the robot 912 is referred to as a control parameter and is denoted by α. An evaluation function indicating whether a task is successful or not when a control command indicated by the control parameter α is used is expressed as g(ξ cond , α). The evaluation function g(ξ cond The level set function that predicts the value of g^(ξ cond , α).

[0093] The level set function g^(ξ cond , α), the level set function g^(ξ cond , α) can be used. Specifically, the condition vector ξ cond and the product space ξ of the control parameter α cond ×α into a region where the task succeeds and a region where the task fails. Then, the level set function g^(ξ cond, α) in the domain where the task is successful. cond , α) < 0, and at the boundary between the area where the task is successful and the area where the task is unsuccessful, g^(ξ cond , α) = 0, and in the region where the task fails, g^(ξ cond , α)>0.

[0094] In learning using a level set function, the condition vector ξ cond Using training data that is a combination of the control parameter α and teacher data that indicates the success or failure of the task, the level set function g^(ξ cond , α) is learned. Here, the level set function g^(ξ cond , α) may be represented by a probabilistic model that allows for the calculation of predicted means and variances.

[0095] Then, in the learning using the level set function, the obtained level set function g^(ξ cond , α) to obtain the given condition vector ξ cond For the level set function g^(ξ cond A control rule π predicts the value of the control parameter α so that the value of h (ξ cond The environmental data acquisition unit 291 learns the control rule π h (ξ cond ) may be used to control the robot 912.

[0096] Alternatively, the environmental data acquisition unit 291 may control the robot 912 based on a policy obtained by reinforcement learning. In this case, the reinforcement learning is performed by using the condition vector ξ cond The value function V(ξ) predicts the future cumulative reward sum in the problem setting shown in cond ) will be studied.

[0097] In this case, the value function V(ξ cond ) to obtain the condition vector ξ cond and state s t For the input of control command a t The strategy π that outputs l(s t , ξ cond ) is learned. Here, the condition vector ξ cond and state s t We assume that the method learns a Universal Policy that treats the above as input to the policy.

[0098] The reinforcement learning method used here is not limited to a specific method. For example, Q-Learning or Soft Actor Critic may be used as the reinforcement learning method used here, but is not limited to these. Furthermore, the environmental data acquisition unit 291 may control the robot 912 using a control device configured as an external device to the estimation device 100. For example, the control device 911 in FIG. 2 may be configured as an external device to the estimation device 100, and the environmental data acquisition unit 291 may control the robot 912 using the control device 911.

[0099] 6 is a diagram showing an example of data input and output when the estimation device 200 estimates the probability distribution of the value of the parameter ξ. In FIG. 6, the environmental data acquisition unit 291 further acquires transition data o to generate a transition data set D o 6 is shown to generate or update the .times. ...

[0100] 7 is a diagram showing an example of a procedure for the estimation device 200 to calculate the distribution of the parameter ξ values. (Step S201) The environment data acquisition unit 291 acquires a transition data set D o After step S201, the process proceeds to step S202.

[0101] (Step S202) The estimating device 200 calculates the distribution of the parameter values. Specifically, in step S202, the estimating device 200 performs the process of Fig. 4. After step S202, the estimating device 200 ends the process of Fig. 7.

[0102] As described above, the environment data acquisition unit 291 acquires the state of the environment 921, the control command for the robot 912, and the transition data o indicating the next state of the environment 921. According to the estimation device 200, the transition data o can be acquired as needed, for example, when additional learning of control for the robot 912 is performed.

[0103] Third Embodiment Fig. 8 is a diagram showing an example of the configuration of an estimation device 300 according to at least one embodiment. In the estimation device 300 shown in Fig. 8, a processing unit 390 further includes a condition vector sampling unit 391, a feasibility determination unit 392, and a skill learning unit 393 in addition to the units included in the processing unit 290 of the estimation device 200 shown in Fig. 5. In other respects, the estimation device 300 is similar to the estimation device 200. Furthermore, a robot system 910 in the third embodiment is similar to that in the first embodiment, and Fig. 2 will also be referred to in the third embodiment.

[0104] 8 shows an example of the configuration of the estimation device 300 when the third embodiment is implemented based on the first and second embodiments. Alternatively, the third embodiment may be implemented based only on the first embodiment, not on the second embodiment. In this case, the estimation device 300 may not include the environmental data acquisition unit 291.

[0105] The skill here refers to an action defined as a unit of action performed by the robot 912. The skill may be defined as a combination of primitive actions performed by the robot 912. The skill may also have parameters. When the estimation device 300 determines that the task cannot be executed, the estimation device 300 performs additional learning of the skill using the distribution of learned parameter values.

[0106] The condition vector sampling unit 391 samples a task parameter ξ that indicates a task that is determined to be a task to be executed. other The distribution of values ​​of q(ξ other ) and the distribution p of the parameter ξ values ​​estimated by the distribution calculation unit 191 θ (ξ) and based on the condition vector ξ condThe value of the condition vector ξ is sampled. A task that is determined to be executed is also called a target task. cond is assumed to be expressed as in equation (6).

[0107]

[0108] In equation (6), vectors are represented by square brackets [ ]. The superscript T on a vector or matrix indicates the transpose of the vector or matrix. other The elements of the condition vector ξ cond In addition, the formula (6) shows that the condition vector ξ is composed of elements other than the parameter ξ. cond The value is the distribution p θ (ξ) and the task parameter ξ other The distribution of values ​​of q(ξ other ) and the task parameter ξ other The distribution of values ​​of q(ξ other ) may be predetermined depending on the task.

[0109] The feasibility determination unit 392 evaluates the value of the sampled condition vector based on the learned skills and determines the feasibility of the task using the learned skills. Being able to execute a task here means that the task will be successful. Being unable to execute a task means that the task will fail. The feasibility determination unit 392 is an example of a feasibility determination means.

[0110] When a control rule obtained by learning using a level set function is used, the feasibility determination unit 392 may determine the feasibility of a task based on the proportion of cases in which the evaluation index value of the level set function is equal to or less than a threshold value ε.

[0111] For example, the feasibility determination unit 392 determines the condition vector ξ cond Each sample value ξ cond,l For the level set function g^, the evaluation index value g^ is calculated based on Equation (7). l Calculate.

[0112]

[0113] where l is the condition vector ξ cond The sample value ξ cond,l is an index variable that identifies the condition vector ξ. l takes an integer value of 1≦l≦M′. M′ is the condition vector ξ cond The sample value ξ cond,l M' is an integer greater than or equal to 1, indicating the number of

[0114] μ g^ denotes the average value of the level set function g. g^ indicates the variance of the value of the level set function g^. Here, it is assumed that the value of the level set function g^ is represented by a probabilistic model that can calculate the predicted mean value and variance. β is a constant β>0. The larger the value of β, the more the variance of g^ l The value of β tends to become large, and the possibility of determining that the task is executable decreases. In this respect, the value of β can be considered to play a role similar to a confidence interval. l <ε is 95% or more, the task may be determined to be executable.

[0115] When a control rule obtained by reinforcement learning is used, the feasibility determination unit 392 may determine whether or not a task can be executed based on the proportion of cases in which the value of the value function in reinforcement learning is equal to or less than a threshold ε. For example, cond,l ) indicates a higher evaluation, the feasibility determination unit 392 calculates each sample value ξ cond,l The value function V(ξ cond,l Alternatively, the feasibility of a task may be determined based on the proportion of cases where the value of V(ξ) is equal to or greater than a threshold value ε. cond,l )<ε is 95% or more, the task may be determined to be executable.

[0116] Alternatively, the feasibility determination unit 392 may determine whether or not a task can be executed based on the ratio of cases in which the task is successful in a simulation of task execution. cond,l In the problem setting, a simulation may be performed using the learned control rule. Then, the feasibility determination unit 392 performs a simulation using M' samples ξ cond,1 , ξ cond,2 , ..., ξ cond,M’ Among them, the evaluation function value j i But, j i If the ratio of ≦ε is 95% or more, the task may be determined to be executable.

[0117] The skill learning unit 393 performs additional learning of a skill when the feasibility determination unit 392 determines that the task cannot be executed. That is, when the feasibility determination unit 392 determines that the task cannot be executed, the skill learning unit 393 performs additional learning of a control rule for the robot 912. The skill learning unit 393 is an example of a learning means.

[0118] The method by which the skill learning unit 393 performs additional skill learning is not limited to a specific method. For example, the skill learning unit 393 may perform additional skill learning using the above-mentioned learning method using a level set function. Alternatively, the skill learning unit 393 may perform additional skill learning using the above-mentioned learning method using reinforcement learning.

[0119] The skill learning unit 393 may be configured external to the estimating device 300. For example, the skill learning unit 393 may be configured as a device external to the estimating device 300.

[0120] 9 is a diagram showing an example of data input and output when the estimation device 300 estimates the probability distribution of the value of the parameter ξ. In addition to the case of FIG. 6, FIG. 9 further shows data input and output in each of the condition vector sampling unit 391, the feasibility determination unit 392, and the skill learning unit 393.

[0121] The condition vector sampling unit 391 calculates the parameter distribution (the distribution p of the parameter ξ calculated by the distribution calculation unit 191). θ (ξ)) and task distribution (task parameter ξ other The distribution of values ​​of q(ξ other )) and based on the condition vector ξ cond Then, the condition vector sampling unit 391 samples the value of the sampling result (condition vector ξ cond The sampling value ξ cond,l ) to the feasibility determination unit 392.

[0122] The feasibility determination unit 392 determines the sampling result (condition vector ξ cond The sampling value ξ cond,l The feasibility determination unit 392 determines whether the task can be executed using the learned skills based on the skill information (information indicating the learned skills) and the task execution status. The feasibility determination unit 392 then outputs an executable flag indicating the determination result of the feasibility of the task to the skill learning unit 393.

[0123] If the feasibility determination unit 392 determines that the task will fail with the learned skill, the skill learning unit 393 calculates the parameter distribution (the distribution p of the parameter ξ calculated by the distribution calculation unit 191). θ The value of the parameter ξ according to (ξ) is used to perform additional skill learning.

[0124] 10 is a diagram showing an example of a processing procedure in which the estimating device 300 calculates the distribution of the parameter ξ values. (Step S301) The estimating device 300 calculates the distribution of the parameter values. Specifically, the estimating device 300 performs the processing of FIG. 7 in step S301. Alternatively, when implementing the third embodiment based only on the first embodiment and not on the second embodiment, the estimating device 300 may perform the processing of FIG. 4 in step S301. After step S301, the process proceeds to step S302.

[0125] (Step S302) The condition vector sampling unit 191 calculates the task distribution indicating the task to be executed and the distribution p θ (ξ) and based on the condition vector ξ condAfter step S302, the process proceeds to step S303.

[0126] (Step S303) The feasibility determination unit 392 calculates the sampled value ξ of the condition vector. cond,l The feasibility of the task is evaluated based on the task information (information indicating the learned skills) and the skill information. After step S303, the process proceeds to step S304.

[0127] (Step S304) The feasibility determination unit 392 calculates the sampled value ξ of the condition vector. cond,l Based on the skill information and the learning ability, the estimating device 300 determines whether the task can be performed with the learned skills. If the feasibility determining unit 392 determines that the task can be performed with the learned skills (step S304: YES), the estimating device 300 ends the processing in Fig. 10. On the other hand, if the feasibility determining unit 392 determines that the task cannot be performed with the learned skills (step S304: NO), the processing proceeds to step S305.

[0128] (Step S305) The skill learning unit 393 calculates the distribution p of the parameter ξ value. θ (ξ) is used to perform additional skill learning. After step S305, the estimation device 300 ends the process of FIG.

[0129] As described above, the feasibility determination unit 392 determines the distribution p θ Based on (ξ), it is determined whether a task that has been determined as a task to be executed can be executed by a learned control command for the robot 912. The estimation device 300 can reflect the characteristics of the object 931 in determining the feasibility of a task, and in this respect, it is possible to determine the feasibility of a task with relatively high accuracy.

[0130] Furthermore, when it is determined that the task that is determined to be the task to be executed cannot be executed with the learned control commands for the robot 912, the skill learning unit 393 performs additional learning of control commands for the robot 912. The estimating device 300 can reduce the load of additional learning of control commands by determining whether additional learning of control commands for the robot 912 is necessary. Furthermore, the estimating device 300 can perform additional learning of control commands as necessary, which is expected to further increase the feasibility of executing the task.

[0131] Fourth Embodiment Fig. 11 is a diagram showing an example of the configuration of an estimation device 400 according to at least one embodiment. In the estimation device 400 shown in Fig. 11, a processing unit 490 further includes a condition vector sampling unit 391, a feasibility determination unit 392, and a control parameter setting unit 491 in addition to the units included in the processing unit 290 of the estimation device 200 shown in Fig. 5. In other respects, the estimation device 400 is similar to the estimation device 200. Furthermore, a robot system 910 in the fourth embodiment is similar to that in the first embodiment, and Fig. 2 will also be referred to in the fourth embodiment.

[0132] FIG. 11 shows an example of the configuration of the estimation device 400 when the fourth embodiment is implemented based on the first and second embodiments. Alternatively, the fourth embodiment may be implemented based only on the first embodiment, not on the second embodiment. In this case, the estimation device 400 may not be configured to include the environmental data acquisition unit 291. Furthermore, the third and fourth embodiments may be implemented together. In this case, the processing unit 490 of the estimation device 400 may further include the skill learning unit 393 of FIG. 8.

[0133] The condition vector sampling unit 391 and the feasibility determination unit 392 are both similar to those described for the estimation device 300. When the feasibility determination unit 392 determines that the task can be performed with the learned skills, the control parameter setting unit 491 determines the value of the parameter ξ for controlling the robot 912. The control parameter setting unit 491 corresponds to an example of a control parameter setting means.

[0134] For example, consider a case where the skill executed by the robot 912 depends on the characteristics of the object 931, and the skill is provided with a parameter that depends on the characteristics of the object 931. In this case, the control parameter setting unit 491 calculates the distribution p θ The value of the parameter ξ may be calculated based on (ξ), and the value of the skill parameter may be set based on the calculated value.

[0135] For example, the control parameter setting unit 491 determines the distribution p θ Sampling value ξ of parameter ξ according to (ξ) j The average value μ ξ may be calculated and used as the value of the parameter ξ for controlling the robot 912.

[0136]

[0137] Alternatively, the control parameter setting unit 491 may set the sampling value ξ j Alternatively, the control parameter setting unit 491 may calculate the value ξ of the parameter ξ for controlling the robot 912 based on the formula (9). * may be calculated.

[0138]

[0139] argmax is a function that outputs the value of the parameter shown below argmax so that the value of the expression shown after argmax becomes maximum. θ It is also possible to find a value of ξ that makes the value of (ξ) as large as possible.

[0140] 12 is a diagram showing an example of data input and output when the estimation device 400 estimates the probability distribution of the value of the parameter ξ. In FIG. 12, a control parameter setting unit 491 is shown instead of the skill learning unit 393 in the case of FIG. 9. When the feasibility determination unit 392 determines that the task can be executed using the learned skill, the control parameter setting unit 491 calculates the parameter distribution (the distribution p of the parameter ξ calculated by the distribution calculation unit 191) θ (ξ)), the value of the parameter ξ is determined for control of the robot 912.

[0141] 13 is a diagram showing an example of a processing procedure for calculating the distribution of the parameter ξ values ​​by the estimation device 400. Steps S401 to S404 in Fig. 13 are similar to steps S301 to S304 in Fig. 10.

[0142] If the feasibility determination unit 392 determines in step S404 that the task cannot be performed with the learned skills (step S404: NO), the estimating device 400 ends the processing in Fig. 13. On the other hand, if the feasibility determination unit 392 determines that the task can be performed with the learned skills (step S404: YES), the processing proceeds to step S405.

[0143] (Step S405) The control parameter setting unit 491 calculates the distribution p θ (ξ) is used to determine the value of the parameter ξ for controlling the robot 912. After step S405, the estimation device 400 ends the processing of FIG.

[0144] As described above, when it is determined that a task defined as a task to be executed in a learned control command for the robot 912 is executable, the control parameter setting unit 491 determines the value of the parameter ξ for controlling the robot 912. According to the estimating device 400, it is possible to reflect the characteristics of the object 931 indicated by the value of the parameter ξ in the control of the robot 912. In this respect, the estimating device 400 is able to control the robot 912 with a relatively high degree of accuracy, which is expected to increase the feasibility of executing the task.

[0145] Fifth Embodiment Fig. 14 is a diagram showing another example of the configuration of an estimation device according to at least one embodiment. In the configuration shown in Fig. 14, an estimation device 610 includes an evaluation function setting unit 611 and a distribution calculation unit 612.

[0146] With this configuration, the evaluation function setting unit 611 sets an evaluation function that indicates a better evaluation the smaller the difference between the state of the characteristic estimation object and the next state of the characteristic estimation object observed under a control command for the characteristic estimation object, and the next state of the characteristic estimation object calculated in a simulation under the set values ​​of the parameters indicating the state, the control command, and the characteristics of the characteristic estimation object.

[0147] The distribution calculation unit 612 searches for the distribution of the parameter values ​​using the evaluation function set by the evaluation function setting unit 611. The evaluation function setting unit 611 corresponds to an example of evaluation function setting means. The distribution calculation unit 612 corresponds to an example of distribution calculation means.

[0148] According to the estimation device 610, it is possible to provide the evaluation function set by the evaluation function setting unit 611 with a subexpression for preventing excessive fitting to data, for example, thereby providing a degree of freedom in searching for the distribution of parameter values ​​by the distribution calculation unit 612. In this respect, according to the estimation device 610, when estimating the characteristics of an object of characteristic estimation in relation to control of the object of characteristic estimation, a degree of freedom in estimation can be obtained.

[0149] Sixth Embodiment Fig. 15 is a diagram showing an example of a processing procedure in an estimation method according to at least one embodiment. The estimation method shown in Fig. 15 includes setting an evaluation function (step S611) and calculating a probability distribution (step S612).

[0150] In setting an evaluation function (step S611), the computer sets an evaluation function that indicates a better evaluation the smaller the difference between the state of the characteristic estimation object and the next state of the characteristic estimation object observed under a control command for the characteristic estimation object, and the next state of the characteristic estimation object calculated by simulation under the set values ​​of the parameters indicating the state, the control command, and the characteristics of the characteristic estimation object.In calculating a probability distribution (step S612), the computer searches for a distribution of parameter values ​​using the set evaluation function.

[0151] 15, the evaluation function to be set can be provided with a subexpression for preventing excessive fitting to data, for example, and thus it is possible to provide a degree of freedom in searching for the distribution of parameter values. In this respect, the estimation method shown in FIG. 15 provides a degree of freedom in estimation when estimating the characteristics of an object to be estimated in relation to control of the object to be estimated.

[0152] 16 is a schematic block diagram illustrating the configuration of a computer according to at least one embodiment. In the configuration shown in FIG. 16, a computer 700 includes a CPU 710, a main memory device 720, an auxiliary memory device 730, an interface 740, and a non-volatile recording medium 750.

[0153] Any one or more of the above-described estimation devices 100, 200, 300, 400, and 610, or a part thereof, may be implemented in a computer 700. In this case, the operation of each of the above-described processing units is stored in the auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program. The CPU 710 also allocates storage areas in the main storage device 720 corresponding to each of the above-described storage units in accordance with the program. Communication between each device and other devices is executed by an interface 740 having a communication function and performing communication under the control of the CPU 710.

[0154] When the estimation device 100 is implemented in a computer 700, the operations of the processing unit 190 and each unit thereof are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0155] Furthermore, the CPU 710 allocates a storage area for the storage unit 180 in the main storage device 720 in accordance with the program. Communication with other devices by the communication unit 110 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the display unit 120 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the operation input unit 130 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0156] When the estimation device 200 is implemented in a computer 700, the operations of the processing unit 290 and each unit thereof are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0157] Furthermore, the CPU 710 allocates a storage area for the storage unit 180 in the main storage device 720 in accordance with the program. Communication with other devices by the communication unit 110 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the display unit 120 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the operation input unit 130 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0158] When the estimation device 300 is implemented in a computer 700, the operations of the processing unit 190 and each unit thereof are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0159] Furthermore, the CPU 710 allocates a storage area for the storage unit 180 in the main storage device 720 in accordance with the program. Communication with other devices by the communication unit 110 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the display unit 120 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the operation input unit 130 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0160] When the estimation device 400 is implemented in a computer 700, the operations of the processing unit 190 and each unit thereof are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0161] Furthermore, the CPU 710 allocates a storage area for the storage unit 180 in the main storage device 720 in accordance with the program. Communication with other devices by the communication unit 110 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the display unit 120 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the operation input unit 130 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0162] When the estimation device 610 is implemented in the computer 700, the operations of the evaluation function setting unit 611 and the distribution calculation unit 612 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0163] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the estimating device 610 to perform processing in accordance with the program. Communication between the estimating device 610 and other devices is performed by the interface 740, which has a communication function and operates under the control of the CPU 710. Interaction between the estimating device 610 and a user is performed by the interface 740, which has an input device and an output device, presenting information to the user via the output device under the control of the CPU 710 and accepting user operations via the input device.

[0164] One or more of the above-described programs may be recorded on nonvolatile recording medium 750. In this case, interface 740 may read the programs from nonvolatile recording medium 750. Then, CPU 710 may directly execute the programs read by interface 740, or may temporarily store the programs in main storage device 720 or auxiliary storage device 730 and then execute them.

[0165] Alternatively, a program for executing all or part of the processing performed by the estimation devices 100, 200, 300, 400, and 610 may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed to perform the processing of each unit. The term "computer system" as used herein includes hardware such as an operating system (OS) and peripheral devices. The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, read-only memories (ROMs), and compact disc read-only memories (CD-ROMs), as well as storage devices such as hard disks built into computer systems. The program may be for implementing part of the aforementioned functions, or may be capable of implementing the aforementioned functions in combination with a program already recorded on the computer system.

[0166] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention.

[0167] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.

[0168] (Supplementary Note 1) An estimation device comprising: evaluation function setting means for setting an evaluation function that indicates a better evaluation the smaller the magnitude of the difference between a state of an object of characteristic estimation and a next state of the object of characteristic estimation observed under a control command for the object of characteristic estimation, and a next state of the object of characteristic estimation calculated by simulation under set values ​​of parameters indicating characteristics of the object of characteristic estimation, the state, the control command, and the object of characteristic estimation; and distribution calculation means for searching for a distribution of values ​​of the parameters using the evaluation function.

[0169] (Supplementary Note 2) The estimation device according to Supplementary Note 1, wherein the evaluation function setting means sets the evaluation function based on a posterior probability of transition data indicating a state of the characteristic estimation object, a control command for the characteristic estimation object, and a next state of the characteristic estimation object under sampled values ​​of the parameters according to a distribution of the parameter values, and a likelihood of a probability distribution of the parameter values ​​based on sampled values ​​of the parameters according to the distribution of the parameter values.

[0170] (Supplementary Note 3) The estimation device according to Supplementary Note 1 or Supplementary Note 2, further comprising: environmental data acquisition means for acquiring transition data indicating a state of the object of characteristic estimation, a control command for the object of characteristic estimation, and a next state of the object of characteristic estimation.

[0171] (Supplementary Note 4) The estimation device according to any one of Supplementary Notes 1 to 3, further comprising: a feasibility determination means for determining whether a task defined as a task to be executed can be executed by a learned control command for the characteristic estimation target, based on a distribution of values ​​of the parameter.

[0172] (Supplementary Note 5) The estimation device according to Supplementary Note 4, further comprising: a learning means for additionally learning the control command when it is determined that the task that is determined to be the task to be executed cannot be executed with the learned control command for the characteristic estimation target.

[0173] (Supplementary Note 6) The estimation device according to Supplementary Note 4 or Supplementary Note 5, further comprising: a control parameter setting means for determining values ​​of the parameters for controlling the characteristic estimation object when it is determined that a task defined as a task to be executed can be executed in a learned control command for the characteristic estimation object.

[0174] (Supplementary Note 7) An estimation method including: a computer sets an evaluation function that indicates a better evaluation the smaller the magnitude of the difference between a state of an object of characteristic estimation and a next state of the object of characteristic estimation observed under a control command for the object of characteristic estimation, and a next state of the object of characteristic estimation calculated by simulation under set values ​​of parameters indicating the state, the control command, and the characteristics of the object of characteristic estimation; and searching for a distribution of values ​​of the parameters using the evaluation function.

[0175] (Supplementary Note 8) The estimation method according to Supplementary Note 7, wherein setting the evaluation function includes setting the evaluation function based on a posterior probability of transition data indicating a state of the characteristic estimation object, a control command for the characteristic estimation object, and a next state of the characteristic estimation object under sampled values ​​of the parameters that follow a distribution of the parameter values, and a likelihood of a probability distribution of the parameter values ​​based on sampled values ​​of the parameters that follow a distribution of the parameter values.

[0176] (Supplementary Note 9) The estimation method according to Supplementary Note 7 or Supplementary Note 8, further comprising: acquiring transition data indicating a state of the object of characteristic estimation, a control command for the object of characteristic estimation, and a next state of the object of characteristic estimation.

[0177] (Supplementary Note 10) The estimation method according to any one of Supplementary Notes 7 to 9, further comprising: determining whether a task defined as a task to be executed can be executed by a learned control command for the characteristic estimation target, based on a distribution of the parameter values.

[0178] (Supplementary Note 11) The estimation device according to Supplementary Note 10, further comprising: when it is determined that a task that is determined to be a task to be executed cannot be executed with a learned control command for the characteristic estimation target, performing additional learning of the control command.

[0179] (Supplementary Note 12) The estimation device according to Supplementary Note 10 or Supplementary Note 11, further comprising: determining values ​​of the parameters for control of the characteristic estimation object when it is determined that a task defined as a task to be executed can be executed in a learned control command for the characteristic estimation object.

[0180] (Supplementary Note 13) A recording medium having recorded thereon a program that causes a computer to execute the following steps: setting an evaluation function that indicates a better evaluation the smaller the magnitude of the difference between a state of an object of characteristic estimation and a next state of the object of characteristic estimation observed under a control command for the object of characteristic estimation, and a next state of the object of characteristic estimation calculated by a simulation under set values ​​of parameters that indicate the state, the control command, and the characteristics of the object of characteristic estimation; and searching for a distribution of values ​​of the parameters using the evaluation function.

[0181] (Supplementary Note 14) The recording medium according to Supplementary Note 13, wherein setting the evaluation function causes the computer to set the evaluation function based on a posterior probability of transition data indicating a state of the characteristic estimation object, a control command for the characteristic estimation object, and a next state of the characteristic estimation object under sampled values ​​of the parameters according to a distribution of the parameter values, and a likelihood of a probability distribution of the parameter values ​​based on sampled values ​​of the parameters according to the distribution of the parameter values.

[0182] (Supplementary Note 15) The recording medium according to Supplementary Note 13 or Supplementary Note 14, further causing the computer to acquire transition data indicating a state of the characteristic estimation object, a control command for the characteristic estimation object, and a next state of the characteristic estimation object.

[0183] (Supplementary Note 16) The recording medium described in any one of Supplementary Notes 13 to 15, further causing the computer to execute: determining whether a task defined as a task to be executed can be executed by a learned control command for the characteristic estimation target, based on the distribution of the parameter values.

[0184] (Supplementary Note 17) The recording medium according to Supplementary Note 16, further causing the computer to perform additional learning of the control command when it is determined that the learned control command for the characteristic estimation target cannot execute a task that is determined to be a task to be executed.

[0185] (Supplementary Note 18) The recording medium described in Supplementary Note 16 or Supplementary Note 17, further causing the computer to execute: determining values ​​of the parameters for control of the characteristic estimation object when it is determined that a task defined as a task to be executed can be executed in a learned control command for the characteristic estimation object.

[0186] The present invention may be applied to an estimation device, an estimation method, and a recording medium.

[0187] DESCRIPTION OF SYMBOLS 100, 200, 300, 400, 610 Estimation device 110 Communication unit 120 Display unit 130 Operation input unit 180 Memory unit 190, 290, 390, 490 Processing unit 191, 612 Distribution calculation unit 192 Sampling unit 193 Likelihood calculation unit 194 Posterior probability calculation unit 195 Simulation execution unit 196, 611 Evaluation function setting unit 291 Environmental data acquisition unit 391 Condition vector sampling unit 392 Feasibility determination unit 393 Skill learning unit 491 Control parameter setting unit 910 Robot system 911 Control device 912 Robot 913 Sensor 914 Data generation device 921 Environment 931 Object

Claims

1. An evaluation function setting means sets an evaluation function that indicates a better evaluation the smaller the difference between the state of the characteristic estimation target and the next state of the characteristic estimation target observed under a control command for the characteristic estimation target, and the next state of the characteristic estimation target calculated by simulation under the state, the control command, and the set values ​​of the parameters indicating the characteristics of the characteristic estimation target; Distribution calculation means for searching the distribution of the parameter values ​​using the evaluation function, An estimation device equipped with the following features.

2. The evaluation function setting means sets the evaluation function based on the state of the characteristic estimation target, a control command for the characteristic estimation target, and the posterior probability of transition data indicating the next state of the characteristic estimation target, under sampling values ​​of the parameter that follow the distribution of the parameter values, and the likelihood of the probability distribution of the parameter values ​​based on sampling values ​​of the parameter that follow the distribution of the parameter values. The estimation device according to claim 1.

3. Environmental data acquisition means for acquiring the state of the characteristic estimation target, control commands for the characteristic estimation target, and transition data indicating the next state of the characteristic estimation target. The estimation device according to claim 1 or claim 2, further comprising:

4. Feasibility determination means that determines whether or not a task designated as a task to be executed by a learned control command for the characteristic estimation target is executable based on the distribution of the parameter values. The estimation device according to claim 1 or claim 2, further comprising:

5. If it is determined that the previously learned control commands for the characteristic estimation target cannot execute the task that is defined as the task to be executed, the learning means performs additional learning of the previously learned control commands for the characteristic estimation target. The estimation device according to claim 4, further comprising:

6. If it is determined that the task to be executed can be performed using the learned control commands for the characteristic estimation target, the control parameter setting means determines the value of the parameter for controlling the characteristic estimation target. The estimation device according to claim 4, further comprising:

7. Computers An evaluation function is set up in which the smaller the difference between the state of the characteristic estimation target and the next state of the characteristic estimation target observed under a control command for the characteristic estimation target, and the next state of the characteristic estimation target calculated by simulation under the state, the control command, and the set values ​​of the parameters indicating the characteristics of the characteristic estimation target, the better the evaluation. The distribution of the parameter values ​​is explored using the aforementioned evaluation function. An estimation method that includes the following.

8. On the computer, Setting an evaluation function that indicates a better evaluation the smaller the difference between the state of the characteristic estimation target and the next state of the characteristic estimation target observed under a control command for the characteristic estimation target, and the next state of the characteristic estimation target calculated by simulation under the state, the control command, and the set values ​​of the parameters indicating the characteristics of the characteristic estimation target, The distribution of the parameter values ​​is explored using the evaluation function, A program that executes the command.