System identification processing method, system control device, and program
The system identification method for humanoid robots addresses the challenge of estimating inertial parameters by iteratively selecting and refining motion data and constraints, achieving accurate model identification and control without needing a good initial model.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ATR ADVANCED TELECOMM RES INST INT
- Filing Date
- 2022-02-01
- Publication Date
- 2026-06-01
AI Technical Summary
Accurately estimating inertial parameters for humanoid robots is difficult due to the complexity of designing reference trajectories that satisfy balance constraints, making high-precision system identification challenging without a good initial model.
A system identification method involving target motion selection, generation, model learning, and selection criterion update, which selects motion data that meets predetermined conditions, generates corresponding motions, estimates model parameters, and updates constraints based on acquired measurement data, allowing for accurate model identification without requiring a good initial model.
Enables accurate model identification for complex systems like humanoid robots by iteratively improving parameter accuracy through a curriculum of motion data selection and constraint relaxation, facilitating high-precision system control.
Smart Images

Figure 0007867682000035 
Figure 0007867682000036 
Figure 0007867682000037
Abstract
Description
[Technical Field]
[0001] This invention relates to a technology for system identification, which controls systems that perform a wide variety of actions, such as humanoid robots. [Background technology]
[0002] In recent years, robots that can walk and perform other actions while maintaining posture (for example, humanoid robots) have been developed. To enable such robots to perform a variety of movements, a robot control system capable of high-speed and precise motion control of the robot (controlled object) is necessary. Such robot control systems require nonlinear control, and to achieve high-speed and precise motion control of the robot (controlled object), it is important to utilize an accurate dynamics model.
[0003] System identification is a method for estimating the model of a dynamic system. In system identification for robot control systems, the inertial parameters (mass, center of gravity, moment of inertia, etc.) of each link of the robot (controlled object) are estimated from measurement data of the motion trajectory (e.g., joint angle trajectory, joint angular velocity trajectory, and torque trajectory).
[0004] The development of identification methods for robotic manipulators has been ongoing for many years, and these methods have also been applied to more complex, multi-degree-of-freedom robots such as humanoid robots (see, for example, Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Ko Ayusawa, Gentiane Venture, and Yoshihiko Nakamura, "Identificationof humanoid robots dynamics using floating-base motion dynamics," in 2008 IEEE / RSJ International Conference on Intelligent Robots and Systems, pages 2854-2859. IEEE, 2008. [Overview of the project] [Problems that the invention aims to solve]
[0006] However, system identification of humanoid robots remains difficult. Accurately estimating inertial parameters requires various measurement data to excite the robot's dynamics. In robotic manipulators, such measurement data can be collected by first designing an appropriate reference trajectory through optimization and then generating that reference trajectory using the robot. However, designing a reference trajectory is difficult for humanoid robots. This is because designing a reference trajectory requires a good initial model of the dynamics, and the optimization becomes more complex. A good initial model is necessary to consider balance constraints so that the humanoid robot can generate the reference trajectory without falling over. Since balance constraints are imposed as nonlinear inequality constraints, the optimization problem becomes difficult to handle. In other words, when performing system identification on a controlled object that performs complex movements like a humanoid robot, it is extremely difficult to perform high-precision system identification because a good initial model is required, and the optimization problem must be solved while considering balance constraints imposed as nonlinear inequality constraints.
[0007] Therefore, in view of the above problems, the present invention aims to realize a system identification processing method, a system control device, and a program that enable accurate model identification without requiring a good initial model, even for control targets that are difficult to identify. [Means for solving the problem]
[0008] To solve the above problems, the first invention is a system identification processing method used in a system control processing system including a controlled object and a system control device for controlling the controlled object, comprising a target motion selection step, a target motion generation step, a model learning step, and a selection criterion update step.
[0009] The target movement selection step involves selecting movement data that meets predetermined conditions from a movement dataset containing multiple movement data for controlling a specific movement.
[0010] The target motion generation step uses the motion data selected in the target motion selection step to generate motion generation data that causes the controlled object to perform the motion corresponding to that motion data.
[0011] The model learning step acquires measurement data when the controlled object is operated using the motion generation data generated in the target motion generation step, and estimates the model parameters for the controlled object based on the acquired measurement data.
[0012] The selection criteria update step updates the conditions that serve as the criteria for selecting motion data in the target motion selection step, based on the parameters estimated by the model learning step.
[0013] In this system identification process, motion data that satisfies the given conditions (constraints) is selected from multiple motion data sets, motion is generated using the selected motion data, and when the controlled object executes the motion using that motion data, measurement data is acquired from the controlled object. Then, in this system identification process, learning is performed based on the acquired measurement data (the parameters of the model (the model of the controlled object) are updated), and the constraints are further updated using the updated parameters. As a result, the accuracy of the parameters gradually improves in this system identification process, allowing for appropriate relaxation of the constraints.
[0014] This allows the system identification process to obtain a curriculum (sequence of motion data) that determines what kind of motion data should be used to advance the learning process.
[0015] Furthermore, this system identification processing method allows the controlled object to gradually perform more complex movements while satisfying constraints, using the curriculum (motion data sequence) obtained as described above, thereby enabling more accurate learning processing (parameter update processing).
[0016] As a result, this system identification method enables accurate model identification (system identification of the controlled object) even for controlled objects that are difficult to identify, without requiring a good initial model.
[0017] The second invention is the first invention, the goal The motion selection step involves deriving an equation of motion for a controlled object, which is obtained by multiplying a regressor matrix φ, formed by angular data of the controlled object's rotational drive mechanism part and / or state data of the controlled object, the angular data and / or the first derivative data of the state data, and the angular data and / or the second derivative data of the state data, with a multidimensional parameter w, to obtain a torque sequence τ to be applied to the controlled object's rotational drive mechanism part. The motion data is then selected by determining whether the conditions are met based on a condition number obtained by singular value decomposition of the regressor matrix, which is obtained by singular value decomposition of the regressor matrix.
[0018] This allows this system identification process to determine whether or not a condition (constraint) is met based on the number of conditions derived from the regressor matrix.
[0019] The third invention is the second invention, wherein the measurement data includes torque series data measured in the controlled object when the controlled object is operated by the motion generation data generated by the target motion generation step.
[0020] Model The learning step involves obtaining a parameter that minimizes the difference between the data obtained by the product operation of the regressor matrix and parameter w and the torque series included in the measurement data, and then using that obtained parameter as the updated parameter w.
[0021] As a result, this system identification method acquires parameters that minimize the difference between the data obtained by the product operation of the regressor matrix and parameter w and the torque series included in the measurement data. This allows the updated parameters to be brought closer to the optimal parameters of the controlled model, enabling efficient system identification.
[0022] The fourth invention is the third invention, the goal The exercise selection steps are: Model In the learning step, the regressor matrix obtained based on the measurement data when the parameters were updated is added to the regressor matrix used to obtain the condition number. This matrix is then used to obtain the condition number.
[0023] This system identification method allows for the acquisition of condition numbers using a regressor matrix updated with actually measured data, enabling appropriate relaxation of constraints. As a result, a wider variety of motion data is selected as learning progresses, leading to the construction of a highly accurate curriculum. Therefore, this system identification method can achieve highly accurate system identification.
[0024] The fifth invention is the third or fourth invention, wherein the conditions used when selecting exercise data in the target exercise selection step are: Model The learning step involves obtaining a regression matrix based on the measurement data when the parameters are updated, and / or Model It is updated based on the parameters updated during the learning step.
[0025] This system identification method allows for updating the conditions (constraints) based on the regressor matrix obtained from measurement data when parameters are updated, and / or based on the parameters updated during the learning step. As learning progresses, a wider variety of motion data is selected, enabling the construction of a highly accurate curriculum. Therefore, this system identification method can achieve highly accurate system identification.
[0026] The sixth invention is any of the first to fifth inventions, further comprising a motion data conversion step of converting a motion data set included in a first data set into a second motion data set which is motion data for a controlled object.
[0027] The target movement selection step then involves selecting movement data from the second movement dataset to enable the controlled object to perform a predetermined movement.
[0028] This allows the system identification process to be performed using, for example, a second set of motion data obtained by transforming a first set of motion data containing a large amount of motion data. When motion data for the controlled object does not exist, this method allows for the acquisition of a large amount of motion data, making it possible to achieve a highly accurate system identification process.
[0029] The seventh invention is a system control device that controls a controlled object by performing model control using parameters obtained by a system identification processing method which is one of the first to sixth inventions.
[0030] This enables high-precision model control using parameters obtained by the system identification processing method, which is one of the first to sixth inventions, and as a result, a system control device that can control the controlled object with high precision can be realized.
[0031] The eighth invention is a program for causing a computer to execute a system identification processing method which is one of the first to sixth inventions.
[0032] This makes it possible to realize a program for execution on a computer that has the same effect as the system identification processing method, which is one of the first to sixth inventions.
[0033] The ninth invention is a system control device for controlling a controlled object and the controlled object, comprising a target motion selection processing unit, a target motion generation unit, a model learning unit, and a selection criterion update unit.
[0034] The target motion selection processing unit selects motion data that satisfies predetermined conditions from a motion dataset containing multiple motion data for performing a predetermined motion on the controlled object.
[0035] The target motion generation unit uses the motion data selected by the target motion selection processing unit to generate motion generation data that causes the controlled object to perform the motion corresponding to the selected motion data.
[0036] The model learning unit acquires measurement data when the controlled object is operated using motion generation data generated by the target motion generation unit, and estimates the model parameters for the controlled object based on the acquired measurement data.
[0037] The selection criterion update unit updates the conditions that serve as the criteria for selecting motion data in the target motion selection processing unit, based on the parameters estimated by the model learning unit.
[0038] This makes it possible to realize a system control device that produces the same effects as the first invention. [Effects of the Invention]
[0039] According to the present invention, a system identification processing method, a system control device, and a program can be realized that enable accurate model identification even for controlled objects that are difficult to identify, without requiring a good initial model. [Brief explanation of the drawing]
[0040] [Figure 1] A schematic diagram of the system control processing system 1000 according to the first embodiment. [Figure 2] A flowchart of the processes executed by the system control processing system 1000. [Figure 3] A table listing the exercise name, the number of data points, and the duration. [Figure 4] A figure showing a robot motion dataset. [Figure 5] A diagram illustrating the curriculum (the curriculum for identifying upper-body humanoid robots). [Figure 6] A diagram illustrating the relaxation of balance constraints in each iteration of the controlled object (upper body humanoid robot). [Figure 7] This figure shows the RMSE of the basis parameters for each number of repeated trials. [Figure 8] This figure shows the RMSE for joint torque at each number of repeated trials. [Figure 9] This figure shows the data on the number of conditions for each repeated trial. [Figure 10] A schematic diagram of the system control processing system 1000A. [Figure 11] A diagram showing the CPU bus configuration. [Modes for carrying out the invention]
[0041] [First Embodiment] The first embodiment will be described below with reference to the drawings.
[0042] <1.1: Configuration of the System Control Processing System> Figure 1 is a schematic diagram of the system control processing system 1000 according to the first embodiment.
[0043] As shown in Figure 1, the system control processing system 1000 comprises a data storage unit DB1, a motion data conversion processing unit PP1, a converted data storage unit DB2, a system control processing unit 100, and a controlled object Rbt1 (for example, a humanoid robot).
[0044] The data storage unit DB1 is a functional unit that can store predetermined data, and is implemented, for example, by a database. The data storage unit DB1 stores, for example, human motion data (for example, a large amount of human motion data (data including diverse human movement trajectories)), and based on a read command from the motion data conversion processing unit PP1, it reads predetermined data and outputs the read data as data Dset_hmn to the motion data conversion processing unit PP1.
[0045] The motion data conversion processing unit PP1 is a functional unit that can read predetermined data from the data storage unit DB1 by outputting a read command to the data storage unit DB1. The motion data conversion processing unit PP1 reads human motion data from the data storage unit DB1 as data Dset_hmn, and performs motion data conversion processing on the read data Dset_hmn to convert human motion data into motion data of the controlled Rbt (e.g., a humanoid robot). Then, the motion data conversion processing unit PP1 stores (writes) the data obtained by the above motion data conversion processing (data obtained by converting human motion data into motion data of the controlled Rbt (e.g., a humanoid robot)) as data Dset_trans in the converted data storage unit DB2.
[0046] The conversion data storage unit DB2 is a functional unit capable of storing predetermined data, and is implemented, for example, by a database. The conversion data storage unit DB2 receives the data Dset_trans output from the motion data conversion processing unit PP1 and stores the data Dset_trans. Furthermore, based on a read command from the target motion selection processing unit 1 of the system control processing unit 100, the conversion data storage unit DB2 reads the data it has stored that matches the read command, and outputs the read data as data Dset_rbt to the target motion selection processing unit 1 of the system control processing unit 100.
[0047] The system control processing unit 100 is a device for controlling the controlled object Rbt1 (for example, a humanoid robot), and as shown in Figure 1, it comprises a target motion selection processing unit 1, a target motion generation unit 2, a model learning processing unit 3, a storage unit 4, a selection criterion update processing unit 5, and a first regressor matrix acquisition unit 6.
[0048] The target motion selection processing unit 1 is a functional unit that can read predetermined data from the conversion data storage unit DB2 by outputting a read command to the conversion data storage unit DB2. The target motion selection processing unit 1 outputs a read command to the conversion data storage unit DB2 and reads the motion data after conversion processing by the motion data conversion processing unit PP1 (for example, motion data for the controlled object Rbt (for example, a humanoid robot)) from the conversion data storage unit DB2 as data Dset_rbt. The target motion selection processing unit 1 also reads data Ds_Φ (regressor matrix data) from the storage unit 4. The target motion selection processing unit 1 also receives data D_h, which includes inequality constraint data output from the selection criterion update processing unit 5. The target motion selection processing unit 1 performs target motion selection processing using the motion data Dset_rbt, data Ds_Φ (regressor matrix data), and inequality constraint data contained in data D_h. The target motion selection processing unit 1 then selects the target motion data Q by the target motion selection processing. * (Reference orbit Q * The data including ) is output as data Dt1 to the second regressor matrix acquisition unit 21 of the target motion generation unit 2. In addition, the target motion selection processing unit 1 outputs target motion data Q for PD control. PD * (Reference orbit Q PD * The data including ) is output as data Dt2 to the PD control unit 22 of the target motion generation unit 2.
[0049] As shown in Figure 1, the target motion generation unit 2 comprises a second regressor matrix acquisition unit 21, a PD control unit 22, and a target motion generation processing unit 23.
[0050] The second regressor matrix acquisition unit 21 receives data Dt1 output from the target motion selection processing unit 1 and performs the process of acquiring a regressor matrix from data Dt1. Then, it outputs the acquired regressor matrix, including the data, as data D1 to the target motion generation processing unit 23.
[0051] The PD control unit 22 receives data Dt2 output from the target motion selection processing unit 1 and measurement data D4_msr output from the controlled object Rbt1. Based on data Dt2 and measurement data D4_msr, the PD control unit 22 generates PD control data D2 and outputs the generated PD control data D2 to the target motion generation processing unit 23.
[0052] The target motion generation processing unit 23 receives data D1 output from the second regressor matrix acquisition unit 21, PD control data D2 output from the PD control unit 22, and data D_prm output from the model learning processing unit 3. The target motion generation processing unit 23 uses data D1, PD control data D2, and data D_prm to perform target motion generation processing. The target motion generation processing unit 23 then outputs data D3 to the controlled object Rbt1, which is data obtained through the target motion generation processing to cause the controlled object Rbt1 to perform a predetermined motion (for example, torque data (torque vector) to be applied to multiple actuators to move the controlled object Rbt1).
[0053] The model learning processing unit 3 receives the measurement data D5_msr output from the controlled object Rbt1 and the measurement data D6_msr output from the controlled object Rbt1, and the output data (regressor matrix data) obtained by outputting these to the first regressor matrix acquisition unit 6. The model learning processing unit 3 also reads data necessary for model learning processing from the memory unit 4 (for example, past regressor matrix data and past torque data).
[0054] The model learning processing unit 3 uses the measurement data D5_msr, the output data of the first regressor matrix acquisition unit 6 (regressor matrix data), and the data read from the storage unit 4 (for example, past regressor matrix data and past torque data) to perform model learning processing and obtain the model parameters w (estimated model parameters w). The model learning processing unit 3 then outputs the data containing the obtained model parameters w as data D_prm to the target motion generation processing unit 23 and the selection criterion update processing unit 5. The model learning processing unit 3 also stores the data contained in the measurement data D5_msr (for example, measured torque data), the obtained parameters w, and the regressor matrix data used in the model learning processing in the storage unit 4.
[0055] The memory unit 4 is a functional unit that stores and holds predetermined data, and performs data reading / writing processing based on commands from the model learning processing unit 3, the selection criterion update processing unit 5, and the target motion selection processing unit 1.
[0056] The selection criterion update processing unit 5 receives data D_prm output from the model learning processing unit 3 and data Ds_Φ read from the storage unit 4. The selection criterion update processing unit 5 uses data D_prm and data Ds_Φ to perform a selection criterion update process and obtains the updated selection criterion data. The selection criterion update processing unit 5 then outputs the data containing the obtained updated selection criterion data as data D_h to the target motion selection processing unit 1. The selection criterion update processing unit 5 is assumed to be a functional unit capable of storing and retaining the obtained selection criterion data. The first regressor matrix 6 receives measurement data D6_msr output from the controlled object Rbt1, obtains regressor matrix data from the measurement data, and outputs the obtained regressor matrix data to the model learning processing unit 3.
[0057] The controlled Rbt1 (for example, a humanoid robot) is equipped with, for example, one or more actuators, and can perform a predetermined operation when a predetermined torque is applied to the actuators. The controlled Rbt is also equipped with, for example, one or more sensors, and these sensors can acquire state data (for example, joint angles), first derivative data of the state (state data), and second derivative data of the state (state data) for each part of the controlled Rbt (for example, each joint). The controlled Rbt1 performs a predetermined motion by being driven and controlled by data D3 (for example, torque data) output from the system control processing unit 100. Furthermore, when the controlled Rbt1 is being driven and controlled (motion generated) by the system control processing unit 100, it outputs state data (e.g., joint angles), first derivative data of the state (state data), and second derivative data of the state (state data) (data acquired by sensors) of each part (e.g., each joint) of the controlled Rbt to the system control processing unit 100 as data D4_msr, data D5_msr, and data D6_msr.
[0058] <1.2: Operation of the System Control Processing System> The operation of the system control processing system 1000, configured as described above, will be explained below.
[0059] Figure 2 is a flowchart of the processes performed by the system control processing system 1000.
[0060] The operation of the system control processing system 1000 will be explained below with reference to the flowchart in Figure 2. The operation of the system control processing system 1000 will be explained in two parts: (1) motion retargeting processing and (2) system identification processing.
[0061] (1.2.1: Motion retargeting processing) First, the motion retargeting process will be described. The motion retargeting process refers to a process of converting the motion data of a living body or an object different from the control target into the motion data of the control target. In the present embodiment, for the sake of convenience of explanation, the motion retargeting process in the case (an example) of converting the motion data of a human into the motion data of a robot (humanoid robot) will be described.
[0062] (Step S11): In step S11, initialization processing of a variable k (a variable for counting the number of executions of the loop process) (k: natural number) used in the loop process (loop 1) is performed. Specifically, the motion data conversion processing unit PP1 sets k = 1.
[0063] (Step S12): In step S12, the loop process (loop 1) is started.
[0064] (Step S13): In step S13, the motion data conversion process is executed. Specifically, the following processes are executed.
[0065] The motion data conversion processing unit PP1 reads the motion data P of a human for a predetermined time (time for T time steps (T: natural number) (one time step is set to, for example, 1 second)) from the data storage unit DB1. h Note that the motion data P of a human for a period of T time steps m is a series of 3D positions of markers measured using a motion capture system as p m and is expressed by the following mathematical formula.
Equation
number
[0066] The N control points, which divide the time (time step) from t=1 to T equally, and the basis functions of the B-spline are expressed as follows.
number
number
number
[0067] Furthermore, the penalty function e is a function that takes into account the joint angle limit, and the value of function e increases when the joint exceeds a specified angle.
[0068] Furthermore, the time derivative of the angular trajectory is calculated using the finite difference method, as shown in the following formula.
number
[0069] The motion data conversion processing unit PP1 converts human motion data P h The robot's motion data Q obtained through the above process r The converted data is stored (written) in the DB2 data storage unit.
[0070] (Step S14): In step S14, the motion data conversion processing unit PP1 increments the variable k by +1 and proceeds to step S15.
[0071] (Step S15): In step S15, it is determined whether the termination condition (k>K) of the loop process (loop 1) is met. If the termination condition of the loop process (loop 1) is not met (k≦K), the process returns to step S12 and the loop process (processing in steps S13 and S14) is repeatedly executed. On the other hand, if the termination condition of the loop process (loop 1) is met (k>K), the motion retargeting process is terminated and the process proceeds to step S21.
[0072] By processing in this way, the system control processing system 1000 executes the motion data conversion process K times, and the motion data Q of K robots are generated. r (Human movement data P) h It is possible to obtain robot motion data (obtained through conversion processing).
[0073] (1.2.2: System Identification Process) Next, I will explain the system identification process.
[0074] To explain the system identification process, we will first describe the necessary background knowledge for system identification.
[0075] Prior Knowledge for System Identification In this embodiment, the controlled object Rbt1 of the system control processing system 1000 is assumed to be a humanoid robot (one example), and the dynamics model of the controlled object Rbt1 (humanoid robot) can be expressed by the following equations of motion.
number
[0076] In system identification, the minimum set of identifiable parameters from the robot's (uncontrolled object Rbt) inertia parameters, called the basis parameters, are estimated. The basis parameters include the mass of each link of the robot (uncontrolled object Rbt), the position of its center of gravity, and the coefficients of the inertia tensor. The basis parameters are determined by the equation of motion shown in (Equation 7), where the parameter w∈R DThis is estimated by taking advantage of the fact that it can be rewritten linearly (as a linear combination) with respect to (D-dimensional real numbers). In other words, the equation of motion shown in (Equation 7) can be expressed as follows:
number
number
number
number
number
[0077] In reality, the parameter estimate w * The accuracy of the measurement is affected by noise in the measurement data. For example, the accuracy of the parameters is affected by the noise δf of the joint torque.
[0078] Estimated value of parameter w * (Solution lol * The sensitivity of the regressor matrix can be measured using the condition number, which is defined as follows:
number
number
number
number
number
number
number
number
[0079] Furthermore, the imposition of nonlinear inequality constraints (Equation 16) makes the optimization (Equation (14)) more complex. The numerous variables to be optimized and the nonlinear balance constraints make the optimization problem of finding the reference trajectory difficult to handle.
[0080] In the system control processing system 1000 of this embodiment, it is not necessary to solve the complex optimization problem for determining the reference trajectory of the robot (controlled object Rbt1) as described above, and a system identification process can be performed without the need for a good initial model.
[0081] The system identification process in the system control processing system 1000 will be explained below with reference to the flowchart in Figure 2.
[0082] (Step S21): In step S21, the variable i (a variable used to count the number of loop executions) (i: a natural number) used in the loop process (loop 2) is initialized. Specifically, the variable i is set to i=1.
[0083] (Step S22): In step S22, the loop process (loop 2) is started.
[0084] (Step S23): In step S23, the target movement selection process is executed. Specifically, the following processes are performed.
[0085] The target motion selection processing unit 1 outputs a read command to the conversion data storage unit DB2, and reads the motion data after conversion processing by the motion data conversion processing unit PP1 (for example, motion data for the controlled target Rbt (for example, a humanoid robot)) from the conversion data storage unit DB2 as data Dset_rbt. Here, the target motion selection processing unit 1 reads the motion data Q of K robots stored in the conversion data storage unit DB2. r (Human movement data P) h The robot motion data (obtained through conversion processing) is read out. This is the motion data Q of the K robots. r Q1 r Q2 r , , , Q K r Let Q be the motion data of the kth robot. r Q k r This is denoted as (k: a natural number, 1 ≤ k ≤ K).
[0086] The target movement selection processing unit 1 performs the following processing to determine the target movement Q * Select this option. (1) The target motion selection processing unit 1 obtains the regressor matrix φ from the motion data Q k r (k: natural number, 1 ≤ k ≤ K). (Obtained by transforming the left side of Equation (7).) (2) The target motion selection processing unit 1 obtains the regressor matrix Φ from time step 1 to T ((obtains the regressor matrix Φ(Q k r ) in Equation (10)). (3) The target motion selection processing unit 1 performs singular value decomposition on the regressor matrix Φ, obtains the largest singular value σ max (Φ) of the regressor matrix Φ, and the smallest singular value σ min (Φ) of the regressor matrix Φ, and further obtains the condition number cond(Φ) of the regressor matrix Φ by performing the processing corresponding to the following formula.
Equation
Equation
Equation
number
[0087] In the initial stages of learning (when i=1 in the loop processing of loop 1), the parameter w 1 The initial model defined by is inaccurate. Therefore, the initial constraint data h 1 (inequality constraint h 1 It is preferable to restrict the reference trajectory so that even inaccurate dynamics models can be generated. Therefore, in the system control processing system 1000, it is preferable to select a quasi-static motion trajectory, for example, in the initial stages of learning.
[0088] The target exercise selection processing unit 1 processes the target exercise data Q selected by the above target exercise selection process. * (Reference orbit Q * The data including ) is output as data Dt1 to the second regressor matrix acquisition unit 21 of the target motion generation unit. In addition, the target motion selection processing unit 1 outputs target motion data Q for PD control. PD * (Reference orbit Q PD * )(Target motion data Q for PD control) PD * This is the target exercise data Q * Data containing the data necessary for PD control (which is extracted from the data) is output as data Dt2 to the PD control unit 22 of the target motion generation unit 2.
[0089] (Step S24): In step S24, the target motion generation process is executed. Specifically, the following processes are performed.
[0090] The second regressor matrix acquisition unit 21 receives data Dt1 (target motion data Q) output from the target motion selection processing unit 1. * (Reference orbit Q * The system takes data (including ) as input and performs a process to obtain the regressor matrix φ from data Dt1. Specifically, the regressor matrix is obtained from the target motion data Q of the controlled object Rbt1. * (Reference orbit Q * ) From this, the regressor matrix φ(θ) at time step t (time t) (t: natural number, 1 ≤ t ≤ T) * t ,dot_θ * t ddot_θ * t )(dot_θ * t :θ * t The first derivative of ddot_θ t :θ * tThe second derivative of is obtained. Also, the second regressor matrix acquisition unit 21, at the first time of loop processing of loop 2 (when i=1), obtains the parameter w based on the previously obtained information about the model of the controlled object Rbt1 (for example, drawing data from the time of manufacture). 1 The parameter w when i=1 is obtained. Then, the second regressor matrix acquisition unit 21 obtains the regressor matrix φ(θ) obtained above. * t ,dot_θ * t ddot_θ * t ) and parameter w 1 The data including this is output to the target motion generation processing unit 23 as data D1.
[0091] The target motion generation processing unit 23 executes the target motion generation process using the data D1 output from the second regressor matrix acquisition unit 21. Specifically, the target motion generation processing unit 23 executes the process corresponding to the following formula to generate the torque τ at time t (time step t). t Obtain the torque sequence (torque vector). Note that if i=1, w=w 1 Therefore, at t=1, the target motion generation processing unit 23 processes the second and third terms on the right-hand side of the following equation (data acquired by the PD control unit 22) as zero (none).
number
[0092] The controlled Rbt1 (humanoid robot) then performs a predetermined motion by being driven and controlled by data D3 (torque data) output from the system control processing unit 100 (target motion generation processing unit 23).
[0093] Then, when the controlled Rbt1 is being driven and controlled (motion generated) by the system control processing unit 100, data such as the mass, center of gravity position, and inertia tensor of each part of the controlled Rbt (for example, each link) (data acquired by sensors) are output to the PD control unit 22 as data D4_msr, and data D5_msr and data D6_msr are output to the model learning processing unit 3 and the first regressor matrix acquisition unit 6, respectively.
[0094] The PD control unit 22, from t=2 onward, processes the data Dt2 (target motion data Q) output from the target motion selection processing unit 1. PD * (Reference orbit Q PD * The PD control unit 22 takes data including ) and measurement data D4_msr output from the controlled Rbt1 as input and generates PD control data D2 based on data Dt2 and measurement data D4_msr. Specifically, the PD control unit 22 acquires data corresponding to the following formula.
number
[0095] The target motion generation processing unit 23 executes the process corresponding to (equation 25) from t=2 onward, thereby generating the torque τ at time t (time step t). t The torque series (torque vector) is obtained. Note that the second and third terms on the right-hand side of (Equation 25) use the data included in the data D2 output from the PD control unit 22.
[0096] Then, the target motion generation processing unit 23 assigns a predetermined motion (target motion data Q) to the controlled object Rbt1, which was obtained through the target motion generation process (processing by (equation 25)). * (Reference orbit Q * Data for executing motion defined by (for example, torque data (torque vector) τ applied to multiple actuators to move the controlled object Rbt1) t This is output as data D3 to the controlled Rbt1.
[0097] Then, the controlled Rbt1 (humanoid robot) is driven and controlled by the data D3 (torque data) output from the system control processing unit 100 (target motion generation processing unit 23), performing a predetermined motion. The controlled Rbt1 outputs the measured data D4_msr to the PD control unit 22, and the measured data D5_msr and D6_msr to the model learning processing unit 3 and the first regressor matrix acquisition unit 6, respectively.
[0098] In the system control processing system 1000, the above process is executed sequentially and repeatedly from t=3 to t=T. As a result, in the system control processing system 1000, at i=1 of the loop processing of loop 1, measurement data of the controlled object Rbt1 from time step t=1 to t=T is acquired.
[0099] (Step S25): In step S25, the model training process is executed. Specifically, the following processes are performed.
[0100] The model learning processing unit 3 uses the measurement data D5_msr, the data read from the memory unit 4 (for example, past regressor matrix data and past torque data), and the output data of the first regressor matrix acquisition unit 6 to perform model learning processing and obtain (update) the model parameters w (estimated model parameters w). Specifically, the model learning processing unit 3 obtains (updates) the model parameters w (estimated model parameters w) by executing a process corresponding to the following mathematical formula.
number
[0101] Furthermore, the torque data from time (time step) t=1 to t=T (the torque data actually observed by the controlled object Rbt1) (corresponding to f in the above formula) is included in the data D5_msr, and the model learning processing unit 3 stores this torque data in the storage unit 4. The model learning processing unit 3 then reads the torque data from time (time step) t=1 to t=T (the torque data actually observed by the controlled object Rbt1) from the storage unit 4 and performs the above processing.
[0102] Furthermore, the parameters obtained through the above process (updated parameters w) * ) is substituted (updated) for parameter w (w←w * The updated parameter w is then used in the i=2 process of loop 2.
[0103] Furthermore, the model learning processing unit 3 processes the parameters w (= updated parameters w) obtained through the above process. * The data including ) is output as data D_prm to the target motion generation processing unit 23 and the selection criterion update processing unit 5.
[0104] (Step S26): In step S26, the selection criteria update process is executed. Specifically, the following processes are performed.
[0105] The selection criterion update processing unit 5 receives the data D_prm output from the model learning processing unit 3 and the data Ds_Φ read from the storage unit 4. The selection criterion update processing unit 5 uses the data D_prm and the data Ds_Φ to perform a selection criterion update process and obtains the updated selection criterion data. Specifically, in (Equation 22), the selection criterion update processing unit 5 replaces the parameter w with the parameter w (=w) updated by the model learning processing unit 3. * (Parameter w included in data D_prm) and constraint data h is obtained by performing the processing corresponding to (Equation 22) and (Equation 23). Note that this obtained constraint data h is the constraint data used in the loop processing of loop 2 at i=2, so constraint data h 2 This is written as (the constraint data h used in the i-th process of the loop processing of loop 2 is h i (This is how it is written.)
[0106] (Step S27): In step S27, the variable i is incremented by +1, and the process proceeds to step S28.
[0107] (Step S28): In step S28, it is determined whether the termination condition (i>I) of the loop process (loop 2) is met. If the termination condition of the loop process (loop 2) is not met (i≦I), the process returns to step S22, and the loop process (processing in steps S23 to S27) is repeatedly executed. On the other hand, if the termination condition of the loop process (loop 2) is met (i>I), the system identification process is terminated.
[0108] The following explanation will focus on the differences between the i-th step (i≧2) of the loop process (loop 2) and the step when i=1.
[0109] In step S23, the regressor matrix Φ used in (Equation 21), (Equation 23), and (Equation 24) is replaced with the regressor matrix Φ in the following equation. That is, in the i-th (i≧2) process, the motion data Q of the controlled object Rbt1 that is the target of the optimization process (processing in (Equation 24)) is replaced. r The regressor matrix Φ(Q r ) and the regressor matrix [Φ(Q) for the measurement data obtained up to the i-1th process (loop processing (loop 2)). 1 ),···,Φ(Q i-1 )] T (The regression matrix of this past measurement data is stored in the memory unit 4, and the target motion selection processing unit 1 can obtain it from the memory unit as data Ds_Φ.) The matrix Φ(Q) (formula below) formed by combining this matrix and the data is used as the regression matrix Φ in (formula 21), (formula 23), and (formula 24). Then, the target motion selection processing unit 1 uses the newly set regression matrix Φ (the regression matrix in the formula below) to perform the same processing as when i=1, thereby obtaining the target motion data Q * (Reference orbit Q * Select (confirm) ).
number
[0110] Then, in step S24, the system control processing system 1000 processes the target motion data Q acquired as described above. * (Reference orbit Q * Using this, the target motion generation process is executed, and the measured data Q is taken from the controlled object Rbt1. i ,f i This is obtained.
[0111] Then, in step S25, the model learning processing unit 3 creates a regressor matrix and torque sequence using newly acquired measurement data from the controlled object Rbt1, and integrates them with the previously acquired regressor matrix and torque sequence. In other words, the model learning processing unit 3 obtains the regressor matrix Φ and torque sequence f by performing a process corresponding to the following formula.
number
[0112] Then, the model learning processing unit 3 uses the regression matrix Φ and torque sequence f set above to perform a process corresponding to the following formula, thereby generating new parameters w (= updated parameters w). * =w i ) obtain.
number
[0113] The above process is then repeatedly executed until i = I, and when i > I, the system identification process ends.
[0114] The system control processing system 1000 then obtains the parameter w at the end of the system identification process as the estimated parameter of the controlled object Rbt1 (estimated parameter after system identification process).
[0115] Summary As described above, the system control processing system 1000 can accurately identify a model without having to solve a complex optimization problem using a good initial model, as in conventional methods. In other words, the system control processing system 1000 prepares motion data (multiple motion data sets) converted from a large amount of motion data (e.g., human motion data) for the controlled object Rbt1, selects motion data that satisfies the constraints from the multiple motion data sets, generates motion using the selected motion data, and acquires measurement data from the controlled object Rbt1 when the motion using that motion data is executed. Then, the system control processing system 1000 performs a learning process based on the acquired measurement data (updating the parameters of the model (model of the controlled object Rbt1)), and further updates the constraints using the updated parameters. In the system control processing system 1000, the accuracy of the parameters gradually improves, so the constraints can be appropriately relaxed.
[0116] This allows the system control processing system 1000 to acquire a curriculum (sequence of motion data) that determines what kind of motion data should be used to advance the learning process.
[0117] Furthermore, the system control processing system 1000 can use the curriculum (motion data sequence) obtained as described above to gradually make the controlled object Rbt1 perform more complex movements while satisfying the constraints, thereby enabling more accurate learning processing (parameter update processing).
[0118] As a result, the system control processing system 1000 can accurately identify a model (system identification for the controlled object Rbt1) without requiring a good initial model, even for controlled objects that are difficult to identify.
[0119] In other words, in the system control processing system 1000, in the initial state, learning processing (parameter update processing) is performed under strict constraints (for example, a state in which the controlled object Rbt1 is made to perform motion with a quasi-static motion trajectory), and the accuracy of the estimated parameters can be improved by updating the estimated parameters based on the measurement data. Then, in the system control processing system 1000, the constraints can be relaxed with the improved accuracy of the estimated parameters, so as learning progresses, the controlled object Rbt1 can be made to move (motion generated) with complex motion data while satisfying the constraints.
[0120] Therefore, the system control processing system 1000 does not require a good initial model and can accurately identify a model even for controlled objects that are difficult to identify (controlled objects with complex configurations (structures) that make system identification difficult).
[0121] (1.3: Experiment) To evaluate the effectiveness of the system identification process (a process executed by the system control processing system 1000) using the curriculum of this embodiment, experiments were conducted in a simulation environment, and these will be described below.
[0122] (1.3.1: Evaluation Method) System identification experiments were performed 200 times. In each experiment, different sequences of Gaussian noise δf were added to the calculated torque (torque for driving the controlled object Rbt1) so that errors would be introduced into the identification results. The standard deviation of the noise was set to 5% of the torque limit. The maximum number of iterations (number of repetitions of the system identification method) in the system identification process using the curriculum of the present invention (processing performed by the system control processing system 1000) was set to I=5. An inaccurate initial model was prepared, and it was verified whether the estimation error of the basis parameters gradually decreased as learning progressed. The estimation error was calculated as follows.
number
number
number
number
[0123] (1.3.2: Experimental Setup) ≪Robot Model≫ In this experiment, a simulation was conducted using a model of a humanoid upper body robot. Assuming a risk of the robot falling over, balance constraints using ZMP were considered. The robot has a total of 18 degrees of freedom: 7 degrees of freedom in each arm, 2 degrees of freedom in the torso, and 2 degrees of freedom in the head. Dynamics simulations were performed using the MuJoCo physics engine. Sensor signals (joint angle, joint angular velocity, and torque) were measured at 10-millisecond intervals. The gains Kp and Kd of the proportional-derivative (PD) controller in (Equation 25) were set manually. Because the gains were set to small values, the robot could not follow the reference trajectory with only the PD controller.
[0124] The robot has D = 132 basis parameters. In the experiment, the true basis parameter w true w' is the inertia parameter used to calculate the true CAD parameters were used as the initial parameter w. 1 w' is the inertia parameter used to calculate the 1 It was randomly generated from a uniform distribution. If w' true If the j-th value is greater than 0, the interval for the uniform distribution is [0, 2w' true j ], otherwise the interval is [2w' true j I set it to ,0].
[0125] ≪Relaxation of Inequality Constraints≫ In the method of the present invention, in each iterative trial, the inequality constraint (h i (Q r , w i ) < 0) was relaxed. Specifically, in each iterative trial, the reduction coefficient of the support polygon was calculated as in (Equation 23). The reduction coefficient α was determined according to the criteria of (Equation 23).
[0126] In the initial iteration, α max = 0.74 was set so that 5% of the reference trajectories could be selected from the robot motion dataset. The scaling coefficient a and the bias coefficient b were respectively a = 2.5×10 -4 b = 1.0 was set.
[0127] Since the condition number cond(Φ) is considered to decrease as the number of iterations increases, the reduction coefficient also decreases. Therefore, the size of the support polygon gradually expands, and in the method of the present invention, as learning progresses, more dynamic motions can be selected.
[0128] ≪Human Motion Database and Retargeting Settings≫ In the motion retargeting process, the KIT Bimanual Manipulation Dataset was used. The dataset contains a lot of measurement data related to daily household chores performed by human hands, such as opening and closing lids and wiping dishes. Each measurement data has the three-dimensional coordinates for each part of the body saved at a time interval of Δt = 10 milliseconds and is encoded in the c3d format. From the dataset, 110 motions were selected, 98 of which were used as the human motion trajectory database, and the remaining 12 were used as reference trajectories for cross-validation.
[0129] The motion retargeting process was performed using a computer environment, specifically in MATLAB using a quasi-Newton method. The weights of the penalty function were w g =100 and w c The value was set to =50 (see Equation 5). The joint angle trajectory was represented using a B-spline curve with control points at 200 millisecond intervals. After motion retargeting, each motion data was labeled according to the name of the motion shown in Figure 3.
[0130] <<Experimental Results>> In the method of the present invention, a series of reference trajectories were selected from a robot motion dataset to create a curriculum. The robot motion dataset is shown in Figure 4. In the method of the present invention, the robot motion dataset was constructed using a large-scale human motion database, so the dataset contains a variety of motion trajectories. Several examples of motion are shown in Figure 4 (in Figure 4, each rectangle shown within the thick rectangle represents motion data (one rectangle represents one motion data)). System identification experiments were performed 200 times with various noise settings, so 200 curricula were constructed by selecting appropriate reference trajectories from the motion dataset.
[0131] Figure 5 shows a typical curriculum (an identification curriculum for an upper-body humanoid robot) (in Figure 5, each rectangle shown within the thick-lined rectangle represents motion data (one rectangle represents one motion data)). In Figure 5, each block within the rectangle enclosed by the thick line represents a reference trajectory that can be generated in each iterative trial. The shade of each block represents the relative magnitude of the condition number in each iterative trial. Darker colored blocks have a smaller condition number than lighter colored blocks. The arrows connecting the blocks with the smallest condition number in each iterative trial (shown as dots in Figure 5) indicate the curriculum.
[0132] As shown in Figure 5, in the method of the present invention, in each iterative trial, the reference trajectory with the smallest number of conditions is selected from the entire dataset, demonstrating that an appropriate curriculum can be constructed.
[0133] Figure 6 shows the relaxation of the balance constraint in each iterative trial in the method of the present invention. In Figure 6, the circle approximately below the robot's body shows the ZMP calculated using the reference trajectory, the larger (outer) polygon approximately below the robot's body represents the upper and lower limits of the support polygon, and the smaller (inner) polygon approximately below the robot's body represents the upper and lower limits of the support polygon reduced by the reduction coefficient α of the support polygon (see Equation 23).
[0134] As shown in Figure 6, in the method of the present invention, the upper and lower limits of ZMP were gradually widened by the processing of (Equation 23), so the diversity of reference trajectories increased with the number of iterations. For example, in the first and second trials, motions within Sweep were not available, but these motions became available in the third trial and showed the smallest number of conditions. As shown in Figure 6, in the method of the present invention, the support polygon expands with the number of trials, so it can be seen that as the number of trials increases, it becomes possible to make the robot perform a variety of motions (more complex motions). Figure 7 shows the RMSE (Equation 31) of the true and estimated parameters for each iteration. The error of the initial model is represented by the blue solid line. As shown in Figure 7, the estimation error gradually decreased as the iterations progressed. Finally, in the fifth trial, the median error was 0.243. Furthermore, the reliability of the parameters was evaluated using the relative standard deviation (Equation 32). In the method of the present invention, 75.8% of the parameters (100 out of 132 parameters) could be estimated with high reliability in the fifth trial. In the initial model, only 57.6% of the parameters (76 out of 132 parameters) had high reliability. Therefore, the method of the present invention was able to learn an accurate model even when a good initial model was not available.
[0135] Figure 8 shows the RMSE (Equation 34) for joint torque. The RMSE was calculated using a dataset for cross-validation, and these datasets were not used for training. As shown in Figure 8, the error was small in the dataset for cross-validation. Therefore, the method of the present invention was able to estimate the basis parameters with high reliability. Figure 9 shows the results of comparing the identification method using the curriculum of the present invention with a method without a curriculum. In Figure 9, the data on the left for each trial count is the condition count data for the method without a curriculum, and the data on the right is the condition count data for the identification method using the curriculum of the present invention. As shown in Figure 9, when the reference trajectory was randomly selected from the robot motion dataset, the condition count became a considerably large value in each iteration. On the other hand, the identification method using the curriculum of the present invention made it possible to efficiently learn accurate basis parameters, indicating that it was important to guide the data collection process using a curriculum.
[0136] As described above, the above experiments demonstrate the remarkable effectiveness of the method of the present invention (system control processing system 1000 using the system control processing system 1000).
[0137] [Other embodiments] In the above embodiment, it is assumed that the Rbt1 controlled by the system control processing system 1000 (or system control processing device 100) is a humanoid robot or an upper-body humanoid robot, but it is not limited to this, and the Rbt controlled by the system control processing system 1000 (or system control processing device 100) is not limited to a humanoid robot or an upper-body humanoid robot. For example, the Rbt controlled by the system control processing system 1000 (or system control processing device 100) may be a system, device, etc., for which high-precision system identification is difficult using conventional methods (for example, a robot (for example, a robot other than a humanoid robot), a multi-legged walking robot, a machine with a complex structure, etc.).
[0138] Furthermore, the system control processing system 1000 (or system control processing device 100) of the above embodiment may be used to realize a system control processing system or system control processing device fixed with parameters after system identification processing. In this case, for example, as shown in Figure 10, this can be realized by replacing the system control processing device 100 with a system control processing device 100A in the system control processing system 1000. That is, as shown in Figure 10, target motion data (target motion data Q * The parameters w of the target motion generation unit 2 of the system control processing device 100A are input from an external source to the second regressor matrix acquisition unit 21 and the PD control unit 22, and the parameters w of the target motion generation processing unit 23 are fixed to the parameters w (optimal parameters w) acquired by the system identification process, and motion generation processing is performed in the same way as the system control processing device 100. As a result, the system control processing device 100A can perform model predictive control using the parameters (optimal parameters) w after the system identification process, making it possible to control the target Rbt1 with high accuracy.
[0139] Furthermore, in the system control processing system 1000 and system control processing device 100 described in the above embodiments, each block may be individually integrated into a single chip using semiconductor devices such as LSIs, or it may be integrated into a single chip including part or all of the blocks.
[0140] Although we have used the term LSI here, depending on the degree of integration, they may also be called IC, system LSI, super LSI, or ultra LSI.
[0141] Furthermore, the method of integrated circuit implementation is not limited to LSIs; it may also be implemented using dedicated circuits or general-purpose processors. After LSI manufacturing, FPGAs (Field Programmable Gate Arrays) that can be programmed, or reconfigurable processors that allow for the reconfiguration of the connections and settings of circuit cells inside the LSI, may also be used.
[0142] Furthermore, some or all of the processing of each functional block in each of the above embodiments may be implemented by a program. And some or all of the processing of each functional block in each of the above embodiments is performed by the central processing unit (CPU) in a computer. The programs for each of these processes are stored in a storage device such as a hard disk or ROM, and are read from the ROM or RAM and executed.
[0143] Furthermore, each of the processes in the above embodiments may be implemented by hardware, or by software (including cases where it is implemented together with an OS (operating system), middleware, or a predetermined library). Moreover, it may be implemented by a hybrid process of software and hardware.
[0144] For example, if each functional part of the above embodiment is implemented by software, the hardware configuration shown in Figure 11 (for example, a hardware configuration in which a CPU (including a GPU), ROM, RAM, input unit, output unit, etc. are connected by a bus) may be used to implement each functional part by software processing.
[0145] Furthermore, when each of the functional units of the above embodiment is implemented by software, the software may be implemented using a single computer having the hardware configuration shown in Figure 11, or it may be implemented by distributed processing using multiple computers.
[0146] In addition, in the descriptions in this specification and the claims, "optimization" (or "optimal") means to achieve the best state, and the parameters for "optimizing" a system (model) refer to the parameters when the value of the objective function of the system reaches the optimal value. The "optimal value" is the maximum value when the system is in a better state as the value of the objective function of the system increases, and the minimum value when the system is in a better state as the value of the objective function of the system decreases. Also, the "optimal value" may be an extreme value. Further, the "optimal value" may allow a predetermined error (measurement error, quantization error, etc.), or may be a value included in a predetermined range (a range where it can be considered to have converged sufficiently).
[0147] In addition, the execution order of the processing method in the above embodiment is not necessarily limited to the description in the above embodiment, and the execution order can be interchanged as long as the gist of the invention is not deviated from. Also, in the processing method in the above embodiment, some steps may be executed in parallel with other steps as long as the gist of the invention is not deviated from.
[0148] A computer program for causing a computer to execute the above-described method and a computer-readable recording medium recording the program are included in the scope of the present invention. Here, examples of the computer-readable recording medium include a flexible disk, a hard disk, a CD-ROM, a MO, a DVD, a DVD-ROM, a DVD-RAM, a large-capacity DVD, a next-generation DVD, and a semiconductor memory.
[0149] The above computer program is not limited to being recorded on the above recording medium, and may be transmitted via an electric communication line, a wireless or wired communication line, a network typified by the Internet, etc.
[0150] Note that the specific configuration of the present invention is not limited to the above-described embodiment, and various changes and modifications are possible without departing from the gist of the invention.
Explanation of Reference Numerals
[0151] 1000, 1000A System Control Processing System 100, 100A System Control Processing Unit 1. Target exercise selection processing unit 2 Target motion generation section 3. Model Learning Processing Unit 5. Selection Criteria Update Processing Unit 6. First Regreser Matrix Acquisition Unit Rbt1 Controlled object PP1 Motion Data Conversion Processing Unit
Claims
1. A system identification processing method used in a system control processing system that includes a controlled object and a system control device for controlling the controlled object, A target movement selection step involves selecting movement data that satisfies predetermined conditions from a movement dataset containing multiple movement data for causing the controlled object to perform a predetermined movement, A target motion generation step generates motion generation data to cause the controlled object to perform the motion corresponding to the motion data using the motion data selected in the target motion selection step, A model learning step which involves acquiring measurement data when the controlled object is operated using the motion generation data generated in the target motion generation step, and estimating the parameters of the model for the controlled object based on the acquired measurement data, A selection criterion update step updates the conditions that serve as the basis for selecting motion data in the target motion selection step, based on the parameters estimated by the model learning step. A method for identifying a system comprising the following components.
2. The aforementioned target exercise selection step is: When the equation of motion for the controlled object is derived, the torque sequence τ applied to the rotational drive mechanism of the controlled object is obtained by multiplying a regressor matrix φ, which is formed by angular data of the rotational drive mechanism part of the controlled object and / or state data of the controlled object, the angular data and / or first derivative data of the state data, and the angular data and / or second derivative data of the state data, with a multidimensional parameter w, the motion data is selected by determining whether the conditions are met based on a condition number obtained by singular value decomposition of the regressor matrix, which is obtained by singular value decomposition of the regressor matrix, and determining whether the conditions are met. The system identification processing method according to claim 1.
3. The measurement data includes torque series data measured in the controlled object when the controlled object is operated using the motion generation data generated by the target motion generation step. The model learning step involves obtaining a parameter that minimizes the difference between the data obtained by the product operation of the regressor matrix and the parameter w and the torque sequence included in the measurement data, and setting the obtained parameter as the updated parameter w. The system identification processing method according to claim 2.
4. The aforementioned target exercise selection step is: In the model learning step described above, the regressor matrix obtained based on the measurement data when the parameters were updated is added to the regressor matrix used to obtain the condition number, and this matrix is called the updated regressor matrix, and the condition number is obtained using this updated regressor matrix. The system identification processing method according to claim 3.
5. The conditions used when selecting motion data in the target motion selection step are updated based on the regressor matrix obtained based on the measurement data when the parameters were updated by the model learning step, and / or based on the parameters updated by the model learning step. The system identification processing method according to claim 3 or 4.
6. A motion data conversion step is performed to convert the motion data set included in the first data set into a second motion data set which is the motion data for the controlled object. Furthermore, The aforementioned target exercise selection step is: Motion data for causing the controlled object to perform a predetermined movement is selected from the second motion dataset. A system identification processing method according to any one of claims 1 to 5.
7. A system control device that controls the target of control by performing model control using parameters obtained by the system identification processing method described in any one of claims 1 to 6.
8. A program for causing a computer to execute the system identification processing method described in any one of claims 1 to 6.
9. A system control device for controlling a controlled object, A target motion selection processing unit selects motion data that satisfies predetermined conditions from a motion dataset containing multiple motion data for causing the controlled object to perform a predetermined motion, A target motion generation unit generates motion generation data to cause the controlled object to perform a motion corresponding to the motion data selected by the target motion selection processing unit, A model learning unit acquires measurement data when the controlled object is operated using the motion generation data generated by the target motion generation unit, and estimates the model parameters for the controlled object based on the acquired measurement data. A selection criterion update unit updates the conditions that serve as the basis for selecting motion data in the target motion selection processing unit, based on the parameters estimated by the model learning unit. A system control device equipped with the following features.