Machine learning device and machine learning method
The machine learning device and method efficiently estimate object pose by training on behavioral data, reducing computational and sensor reliance, thus enhancing accuracy in dynamic environments.
Patent Information
- Application Number
- PCT/JP2025/000436
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2025-01-09
- Publication Date
- 2025-08-14
AI Technical Summary
Existing technologies face challenges in efficiently estimating the pose of objects from images, particularly in dynamic environments, due to high computational demands and reliance on sensors that can affect accuracy.
A machine learning device and method that utilizes behavioral data to train an estimation model by capturing multiple frames of an object's changing posture, reducing the need for depth sensors and enhancing estimation accuracy through object-centric learning and Newtonian transition models.
The approach enables efficient and accurate estimation of object pose from images, reducing computational load and sensor dependency, while improving estimation accuracy in dynamic conditions.
Smart Images

Figure JP2025000436_14082025_PF_FP_ABST
Abstract
Description
Machine learning device and machine learning method
[0001] The present disclosure relates to a machine learning device and a machine learning method for an estimation model that estimates the pose of an object from an image.
[0002] Non-Patent Document 1 discloses a technology for estimating six-dimensional object pose from an RGB-D image containing color information and depth information through machine learning using a neural network (NN). This technology is applied to robots and the like that grasp and manipulate objects based on the estimated pose. The technology in Non-Patent Document 1 reduces calculation time by introducing an NN into pose estimation using ICP (Iterative Closest Point), and also extracts color information and depth information through machine learning to achieve pose estimation that is robust against occlusion, etc.
[0003] Chen Wang, Danfei Xu, Yuke Zhu, Roberto Martin-Martin, Cewu Lu, Li Fei-Fei, and Silvio Savarese, "DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion", Computer Vision and Pattern Recognition (CVPR), 2019.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever, . "Learning transferable visual models from natural language supervision," International conference on machine learning, PMLR, 2021.F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf, “Object-Centric learning with slot attention,” in NeurIPS, 2020.A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Doll´ar, and R. Girshick, “Segment anything,” in arXiv, 2023.M. Jaques, M. Burke, and T. M.Hospedales, “NewtonianVAE: Proportional control and goal identification from pixels via physical latent spaces,” in CVPR, 2021. M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, X. Wang, X. Zhai, T. Kipf, and N. Houlsby, “Simple Open-Vocabulary object detection with vision transformers,” in ECCV, 2022.
[0004] The present disclosure provides a machine learning device and a machine learning method that can efficiently perform machine learning of an estimation model that estimates the pose of an object from an image.
[0005] A machine learning device according to one aspect of the present disclosure includes a storage unit that stores image data representing an image of an object captured by an imaging device, and a control unit that acquires behavioral data representing a change in the posture of the object as seen from the imaging device. The image data includes multiple frames of images in which the posture of the object changes according to the change in the behavioral data. The control unit acquires additional information associated with at least one frame of images in the image data, representing the posture of the object in the image. The control unit performs machine learning on the posture of the object in each of the multiple frames based on the behavioral data, the image data, and the additional information, to generate a trained estimation model. The estimation model calculates an estimate of the posture of the object from the image of the object.
[0006] A machine learning method according to one aspect of the present disclosure is executed by a computer control unit and includes steps of acquiring image data representing an image of an object captured by an imaging device and acquiring behavioral data representing a change in the posture of the object as seen from the imaging device. The image data includes multiple frames of images in which the posture of the object changes according to the change in the behavioral data. In this method, the control unit acquires additional information associated with at least one frame of the image data and representing the posture of the object in that image, and performs machine learning on the posture of the object in each of the multiple frames based on the behavioral data, image data, and additional information to generate a trained estimation model. The estimation model calculates an estimate of the posture of the object from the image of the object.
[0007] These general and specific aspects may be realized by a system, a method, and a computer program, as well as combinations thereof.
[0008] The machine learning device and machine learning method of the present disclosure can efficiently perform machine learning of an estimation model that estimates the pose of an object from an image.
[0009] FIG. 1 is a diagram illustrating an overview of a control system according to a first embodiment; FIG. 2 is a block diagram illustrating the configurations of a robot, a camera, a control device, and a terminal device in the control system of the first embodiment; FIG. 3 is a diagram illustrating machine learning of a posture estimation model in the control system; FIG. 4 is a sequence diagram illustrating the operation during training of a posture estimation model in the control system of the first embodiment; FIG. 5 is a diagram showing an example of a GUI display in the terminal device; FIG. 6 is a diagram illustrating a machine learning model used for machine learning of a posture estimation model; FIG. 7 is a sequence diagram illustrating the operation after training of a posture estimation model in the control system;
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. However, more detailed explanation than necessary may be omitted. For example, detailed explanation of well-known matters or redundant explanation of substantially the same configuration may be omitted. This is to avoid unnecessary redundancy in the following explanation and to facilitate understanding by those skilled in the art.
[0011] The inventors have provided the accompanying drawings and the following description to enable those skilled in the art to fully understand the present disclosure, and do not intend for them to limit the subject matter described in the claims.
[0012] First Embodiment In a first embodiment, a robot control system will be described as an example of using a machine learning device according to the present disclosure.
[0013] 1. Configuration The configuration of the control system according to the first embodiment will be described with reference to FIGS. 1 and 2.
[0014] 1-1. Overview of the Control System FIG. 1 is a diagram illustrating an overview of a control system 1 according to this embodiment. The control system 1 includes a robot 2, a camera 3, a control device 4, and a terminal device 5. In the control system 1, the robot 2, the camera 3, the control device 4, and the terminal device 5 are connected to each other so that data can be communicated with each other. The control device 4 is an example of a machine learning device in this embodiment.
[0015] The control system 1 is applied to an application in which the operation of a robot 2 is controlled by a control command transmitted from, for example, a control device 4. For example, the robot 2 includes a manipulator 2b including multiple joints and an end effector 2c for manipulating various objects 6. The end effector 2c is connected to the tip of the manipulator 2b so as to come into contact with the object 6 in, for example, an operation to grasp the object 6. The camera 3 is attached to the tip of the end effector 2c.
[0016] In the control system 1, the robot 2 moves the end effector 2c by driving the manipulator 2b in response to a control command. In the example of Fig. 1, the end effector 2c performs an action of grasping an object 6, such as a hammer, placed on a workbench 10. The control system 1 estimates the posture of the object 6 from an image of the object 6 captured by the camera 3, and causes the robot 2 to perform an action such as grasping the object 6. In this embodiment, the object posture obtained by such posture estimation includes the position and angle of the object 6 relative to the camera 3.
[0017] In the control system 1, for example, the control device 4 executes machine learning of a posture estimation model that estimates the posture of the object 6 from an image. The control system 1 also accepts, for example, a user operation on the terminal device 5 to input various information used in the machine learning of the posture estimation model.
[0018] 1-2. Detailed Configuration The detailed configuration of the robot 2, camera 3, control device 4, and terminal device 5 in the control system 1 will be described using Fig. 2. Fig. 2 is a block diagram illustrating the configuration of the robot 2, camera 3, control device 4, and terminal device 5 in the control system 1 of this embodiment.
[0019] 2, the robot 2 includes a control unit 20 and a communication interface 22. Hereinafter, the interface will be abbreviated as "I / F." The control unit 20 is configured with various processors such as a CPU, and controls the overall operation of the robot 2.
[0020] The communication I / F 22 is a circuit for performing data communication in accordance with a predetermined communication standard. For example, the predetermined communication standard may include USB, IEEE 1395, IEEE 802.3, IEEE 802.11a / 11b / 11g / 11ac, Wi-Fi (registered trademark), Bluetooth (registered trademark), etc. That is, the communication I / F 22 may be configured to include a connection terminal for connecting to an external device, and may communicate with the external device via a communication network or directly.
[0021] The control unit 20 of the robot 2 receives control commands from the control device 4 via the communication I / F 22 and drives the manipulator 2b based on the control commands. For example, the robot 2 includes a drive device such as a motor that drives each joint of the manipulator 2b. Based on the control commands, the control unit 20 changes the posture of the tip of the manipulator 2b, i.e., the hand posture, using the drive device.
[0022] The camera 3 includes an imaging unit 31 that captures an image and generates image data, and a communication I / F 32 that transmits the image data to an external device such as the control device 4. The imaging unit 31 is realized by a CCD image sensor, a CMOS image sensor, or the like. The communication I / F 32 is a circuit for performing data communication in accordance with a predetermined communication standard, similar to the communication I / F 22 of the robot 2, for example. The camera 3 is an example of an imaging device in the control system 1.
[0023] The control device 4 includes a control unit 40, a storage unit 41, and a communication I / F 42. The control device 4 is configured by various information processing devices such as computers.
[0024] The control unit 40 includes, for example, a CPU, and executes a program (software) to realize predetermined functions in the control device 4. Instead of a CPU, the control unit 40 may include a processor configured with a dedicated electronic circuit designed to realize the predetermined functions. That is, the control unit 50 can be realized by various processors such as a CPU, MPU, GPU, DSU, FPGA, ASIC, etc. The control unit 40 may be configured with one or more processors.
[0025] The storage unit 41 is a storage medium that stores programs and data necessary to realize the functions of the control device 4, and is configured, for example, as an HDD or SSD. The above programs may be provided via a communication network such as the Internet, or may be stored on a portable recording medium. In addition to the above programs, the storage unit 41 also stores image data D1, behavioral data D2, object identification information D3, annotation data D4, and a machine learning model 7, etc., which are used for machine learning of the posture estimation model.
[0026] The image data D1 is acquired from the camera 3 via the communication I / F 42. The behavior data D2 is generated, for example, by the control unit 40, as data indicating behavior at each time point for calculating control commands to the robot 2. The object identification information D3 is information for identifying the object 6. The annotation data D4 is an example of additional information that is associated with an image indicated by the image data D1 and indicates the posture of the object 6 in the image. The object identification information D3 and the annotation data D4 are input, for example, by a user operation on the terminal device 5, and acquired via the communication I / F 42. The machine learning model 7 includes various programs, parameters, etc. for generating a posture estimation model that has been trained by machine learning.
[0027] Furthermore, for example, the machine learning model 7 includes a program and parameters that function as the posture estimation model 70 in the control system 1. The storage unit 41 stores the trained posture estimation model 70 after the control system 1 executes machine learning.
[0028] The storage unit 41 may include a RAM such as a DRAM or an SRAM, and may function as a buffer memory that temporarily stores (i.e., holds) data. The storage unit 41 may function as a working area for the control unit 40, and may include a storage area in the internal memory of the control unit 40.
[0029] The communication I / F 42 is a circuit for performing data communication in accordance with a predetermined communication standard, similar to the communication I / F 22 of the robot 2, for example. The communication I / F 42 is an example of an input unit that inputs various information to the control device 4. In the control device 4, the communication I / F 42 may constitute an acquisition unit that receives various information through communication with an external device, or an output unit that transmits various information.
[0030] The terminal device 5 includes a control unit 50, a storage unit 51, a communication I / F 52, an operation unit 53, and a display unit 54. The terminal device 5 is configured as an information processing device such as a personal computer (PC).
[0031] The control unit 50 is composed of various processors such as a CPU, and realizes predetermined functions such as controlling each part of the terminal device 5 by executing a program (software) stored in the storage unit 51, for example. The storage unit 51 is a storage medium that stores the above-mentioned programs, and is composed of, for example, a non-volatile semiconductor memory. The storage unit 51 may include a RAM such as a DRAM or an SRAM, and may function as a work memory for the control unit 50. The communication I / F 52 is, for example, a circuit for performing data communication in accordance with a predetermined communication standard, similar to the communication I / F 22 of the robot 2.
[0032] The operation unit 53 is a general term for operation members operated by the user. The operation unit 53 is configured by, for example, any one of a keyboard, a mouse, a trackpad, a touchpad, buttons, switches, etc., or a combination thereof. The operation unit 53 may configure a touch panel together with the display unit 54. The operation unit 53 acquires various information input by user operations. The operation unit 53 may also include a connection terminal for connecting an external operation member.
[0033] Display unit 54 is configured by, for example, a liquid crystal display or an organic EL display, and displays a GUI used for machine learning of posture estimation model 70. Display unit 54 may display various icons for operating operation unit 53, and various information such as information input from operation unit 53.
[0034] In the above description, an example has been described in which the control system 1 includes the control device 4 and the terminal device 5, but the configuration of the control system 1 is not limited to the above example. For example, the control device 4 may be configured integrally with the robot 2, or may be configured integrally with the terminal device 5.
[0035] 2. Operation The operation of the control system 1 and the control device 4 configured as above will be described below.
[0036] In order to enable the robot 2 illustrated in FIG. 1 to grasp an object 6, the control system 1 performs machine learning of the posture estimation model 70 in the control device 4 to calculate an estimated value indicating the estimated posture of the object 6 from an image captured by the camera 3.
[0037] In FIG. 1, the coordinate system seen from the camera 3, i.e., the camera coordinate system (X C , Y C , Z C ) and a coordinate system seen from the object 6, i.e., an object coordinate system (X1, Y1, Z1). The pose estimation model 70 of this embodiment calculates an estimated value of the object pose as a relative pose between the camera coordinate system and the object coordinate system. The control system 1 collects image data D1, behavioral data D2, etc. as learning data used in machine learning of the pose estimation model 70, and stores the data in the storage unit 41 of the control device 4, for example, as shown in FIG. 2.
[0038] In collecting learning data, the control device 4 of the control system 1 drives the manipulator 2b to which the camera 3 is attached based on the behavioral data D2. The control device 4 generates the behavioral data D2 so that the relative posture of the object 6 changes between multiple frame images sequentially captured by the camera 3. The control device 4 collects image data D1 including the captured multiple frame images.
[0039] The control device 4 performs machine learning of the posture estimation model 70 using a machine learning model 7 constructed to sequentially transition the object posture as the relative posture. The machine learning model 7 can generate a trained posture estimation model 70 using annotation data D4 associated with some of the frames based on image data D1 and behavioral data D2 collected so that the object posture changes over multiple frames. This can reduce various loads related to annotations for machine learning, for example.
[0040] 2-1. Overview of Machine Learning Fig. 3 is a diagram for explaining machine learning of the posture estimation model 70 in the control system 1 of this embodiment. Fig. 3(A) shows a graphical model of the machine learning model 7. Fig. 3(B) shows examples of image data D1 and annotation data D4.
[0041] FIG. 3 shows the image I at each time in the image data D1 illustrated in FIG. 3B. 0 , I 1 , I 2 1 shows an example in which a pair of scissors 61, a hammer 62, and a clamp 63 are shown as the plurality of objects 6. The control device 4 of the control system 1 may take, for example, an image I 0 From the above, a latent representation x indicating the estimated object pose for each of the objects 61 to 63 is 0 and obtain the latent representation x 0 The posture estimation model 70 is trained by using the machine learning model 7 that sequentially transitions the image I. Each of the objects 61 to 63 is identified by, for example, the object identification information D3 (FIG. 2). 0 Similarly, the subsequent images I 1 , I 2 From the above, the latent representation x 1 , x 2 is obtained.
[0042] 3 , the machine learning model 7 acquires four (K=4) latent representations from the image at each time, one corresponding to each of the objects 61-63 and the background in the image. In this way, the machine learning model 7 of this embodiment acquires K (K is an integer equal to or greater than 2) latent representations corresponding to each of the one or more objects 6 identified by the object identification information D3 and the background. The number K can be set to any number in response to user operation of the terminal device 5, for example, both when the machine learning model 7 is training the posture estimation model 70 and when making inference using the trained posture estimation model 70.
[0043] The control device 4 uses the machine learning model 7 to generate, for example, the first frame image I in the image data D1. 0 The latent representation x obtained from 0 , and the behavior u corresponding to the time of the frame in the behavior data D2 0 From the latent representation x at the time of the next frame 1 The action u at each time indicates, for example, the speed at which the hand posture of the manipulator 2b is changed as a speed control command to the robot 2 at that time. In this example, the action u corresponds to the speed at which the object posture as seen from the camera 3 attached to the tip of the manipulator 2b changes as the camera 3 moves relative to the object 6 stationary on the workbench 10.
[0044] The control device 4 receives, for example, the image I of the initial frame described above. 0 Latent representation x 0 and behavior 0 The latent representation of the next frame x is calculated from 1 and the next frame image I 1 The latent representation x obtained from 1 The machine learning model 7 performs machine learning so that the behavior data D1 and the behavior data D2 match. 1 By this, the latent expression x 1 The state of the object posture indicated by is transitioned to the state at the next time. Then, the latent representation x 2 and the image I of the frame 2 The latent representation x obtained from 2The pose estimation model 70 is trained so that
[0045] The control device 4 generates a trained posture estimation model 70 by performing machine learning on the object posture in each frame of the image data D1 using the behavior data D2 with the machine learning model 7 as described above. The control device 4 uses annotation data D4 as a teacher signal in this machine learning. In the example of FIG. 3 , the annotation data D4 is the latent representation x 0 Annotation coordinate y with respect to 0 is given to the machine learning model 7. The image I of the initial frame shown in FIG. 0 In the example, two predetermined points P1 and P2 representing the representative points of the scissors 61 are annotated, and their coordinates (p 1 , q 1 ), (p 2 , q 2 ) is the annotation coordinate y of the scissors 61 0 is obtained as:
[0046] In the control system 1, the control device 4 calculates the annotation coordinates y for each of the objects 61 to 63 in response to a user operation on the terminal device 5, for example. 0 In the machine learning model 7 of this embodiment, the annotation coordinate y 0 The initial frame image I 0 The information about the object pose added to the image I 1 , I 2 In this way, for example, information about the object pose added to only some of the frames can be efficiently used to perform machine learning of the pose estimation model 70, thereby reducing the annotation load.
[0047] 2-2. Operation During Learning The operation during learning, in which the control system 1 of this embodiment collects learning data for the posture estimation model 70 and performs machine learning, will be described with reference to FIGS.
[0048] FIG. 4 is a sequence diagram illustrating an example of the operation during learning of the posture estimation model 70 in the control system 1 of this embodiment.
[0049] The control system 1 accepts a user operation to input object identification information D3, for example, at the operation unit 53 of the terminal device 5 (S51). The object identification information D3 is input for each object 6 so as to identify multiple objects 6, such as scissors 61, a hammer 62, and a clamp 63, as shown in FIG. 3B. The control unit 50 of the terminal device 5 transmits the input object identification information D3 to the control device 4 via the communication I / F 52. The control unit 40 of the control device 4 acquires the object identification information D3 from the terminal device 5 via the communication I / F 42 (S11).
[0050] The control unit 40 generates behavior data D2 for driving the manipulator 2b so as to change the relative attitude between the camera 3 and the object 6 (S12). The control unit 40 stores the generated behavior data D2 in the storage unit 41, for example. The control unit 40 generates behavior data D2 for randomly driving the manipulator 2b, for example, each time step S12 is executed. Alternatively, the control unit 40 may generate new behavior data D2 for changing the relative attitude by referring to past behavior data D2 stored in the storage unit 41, for example. Alternatively, the control unit 40 may generate behavior data D2 so as to follow a predetermined reference trajectory, or so as to follow a trajectory obtained by randomly perturbing the reference trajectory.
[0051] The control unit 40 transmits a control command to the robot 2 via the communication I / F 42 to cause the robot 2 to change the hand posture based on the generated behavior data D2 (S13). Such a posture change control command is, for example, a speed control command to the robot 2 in the behavior data D2. Based on the received control command, the control unit 20 of the robot 2 calculates a position command that instructs the drive device of each joint of the manipulator 2b to each drive position, or a torque command to each drive device. The calculation of the position command etc. may be performed by the control unit 40 of the control device 4, and a control command including the calculated position command etc. may be transmitted to the robot 2.
[0052] The control unit 20 of the robot 2 changes the hand posture by driving the manipulator 2b based on the control command received from the control device 4 via the communication I / F 22 (S20).
[0053] In the control system 1, the change in the hand posture changes the posture of the camera 3 fixed to the tip of the manipulator 2b. Then, the control system 1 uses the camera 3 whose posture has been changed to capture an image of the object 6 placed on the workbench 10 (S30). The camera 3 transmits image data representing the captured image to the control device 4 via the communication I / F 32. Note that in step S20, for example, data representing the hand posture before and after the change may be transmitted from the robot 2 to the camera 3. The camera 3 may include various processors such as a CPU, and the control system 1 may stop the imaging operation by the camera 3 and proceed to the next data collection operation (S1) if the trajectory or speed of the change in the hand posture deviates from a predetermined reference value.
[0054] The control unit 40 of the control device 4 acquires image data from the camera 3 via the communication I / F 42 (S14). The control unit 40 stores the acquired image data in the storage unit 41, for example, in association with the time in the behavior data D2.
[0055] As described above, the control system 1 acquires image data captured by the camera 3 while changing the hand posture of the manipulator 2b based on the behavior data D2 (S12 to S14, S20, S30), thereby collecting learning data for the posture estimation model 70 (S1). The control system 1 repeats this data collection operation (S1) multiple times, changing the relative posture between the camera 3 and the object 6 in accordance with changes in the hand posture, and repeats collection of learning data until, for example, a predetermined number of frames of image data D1 are acquired. As a result, learning data including the image data D1 and the behavior data D2 is obtained corresponding to one trajectory along which the manipulator 2b is driven during the period of the relevant number of frames.
[0056] After acquiring image data D1, the control system 1 accepts a user operation to input annotation data D4 for the initial frame of the image data D1 at the operation unit 53 of the terminal device 5 (S55). The control unit 40 of the control device 4 acquires the annotation data D4 input from the terminal device 5 via the communication I / F 42 (S15).
[0057] 5 is a diagram showing a display example of a graphical user interface (GUI) in the terminal device 5. The control unit 50 of the terminal device 5 causes the display unit 54 to display a GUI for inputting, for example, object identification information D3 and annotation data D4. In the example of FIG. 5, the display unit 54 displays an image I of an initial frame in image data D1 similar to the example of FIG. 3(B). 0 is displayed.
[0058] The operation unit 53 inputs the annotation coordinate y of the annotation data D4. 0 Then, by clicking the mouse, etc., you can 0 The display unit 54 accepts a user operation to designate two predetermined points P1 and P2 on the image I. For example, the two predetermined points for the scissors 61 as the object 6 are set in advance to correspond to the tip and hinge of the scissors 61, respectively. In the example of FIG. 5, the display unit 54 displays the image I 0 The coordinates of the two specified points above are displayed.
[0059] In this embodiment, an example will be described in which the pose estimation model 70 calculates an estimated value of the object pose for an object 6 placed on the plane of the workbench 10 as shown in Fig. 1. In this example, as the pose on such a two-dimensional plane, the estimated value of the object pose is calculated in three degrees of freedom: coordinates indicating a position according to translation in the X and Y directions in the camera coordinate system, and an angle according to rotation on the XY plane. In this case, in training the pose estimation model 70, for example, the image I shown in Fig. 3B and Fig. 5 is used to give the origin and orientation of the object coordinate system as seen from the camera coordinate system. 0 Annotation coordinate y of the two points P1 and P2 above 0 can be used.
[0060] Furthermore, the operation unit 53 accepts a user operation to input, for example, text information specifying a target object of the pose estimation model 70. The target object indicates the object 6 for which an estimated value of the object pose is to be calculated by the pose estimation model 70. In the example of FIG. 5 , text information "scissors" specifying scissors 61 as the target object is input. For example, in step S51, the control unit 50 may generate object identification information D3 from the input text information using a model capable of converting text information into features, such as that disclosed in Non-Patent Document 2.
[0061] The control system 1 associates, for example, one trajectory of the manipulator 2b with learning data collected by repeating the data collection operation (S1), and acquires object identification information D3 and annotation data D4 input from the terminal device 5 as described above (S11, S15).
[0062] In this embodiment, the control unit 40 of the control device 4 performs machine learning of a posture estimation model 70 using the machine learning model 7 based on, for example, image data D1 and behavior data D2 collected for each of a plurality of trajectories, as well as annotation data D4 for each trajectory (S16). Furthermore, in step S16 of this embodiment, the posture estimation model 70 is trained to calculate an estimated value of an object posture using the object 6 specified by the object identification information D3 as the target object. The control unit 40 generates the trained posture estimation model 70 as described above and stores the generated posture estimation model 70 in the storage unit 41.
[0063] According to the above-described operation of the control system 1, the control device 4 can acquire image data D1 including multiple images in which the object posture changes in response to changes in the hand posture based on the behavior data D2 (S1). Furthermore, object identification information D3 and annotation data D4 are input through user operation on the terminal device 5 (S51, S55). The control device 4 performs machine learning of the posture estimation model 70 using the data D1, D2, D4, and the object identification information D3 (S16). As a result, the posture estimation model 70 can learn images in multiple frames of the image data D1 in which the object posture changes in response to the behavior data D2, using, for example, annotation data D4 added to an initial frame as a teacher signal to propagate to other frames.
[0064] Furthermore, by using the object identification information D3 that identifies the object 6, it is possible to generate a trained pose estimation model 70 that calculates an estimated value of an object pose, for example, using a specific object 6 specified by the object identification information D3 as a target object. The object identification information D3 is not limited to the above example, and may be generated using various algorithms that convert text information into a numeric string or the like, or a predetermined conversion table, etc. Furthermore, the object identification information D3 is not limited to text information, and may also be input as a numeric value.
[0065] The annotation data D4 is not limited to being added to the initial frame as a part of the image data D1, but may be added to, for example, an intermediate frame or to two or more frames.
[0066] The pose estimation model 70 is not limited to the above example, and may calculate an estimated value of the object pose with six degrees of freedom (three degrees of freedom each for translation and rotation). In this case, the position of the object 6 is expressed, for example, by coordinates corresponding to translation in three directions in a Cartesian coordinate system. The angle of the object 6 is expressed, for example, by Euler angles corresponding to rotation around each of the mutually orthogonal axes of the coordinate system. The angle may be expressed by a rotation matrix, a quaternion, or the like, and may be expressed by various expressions that specify the orientation of the object 6. For example, in such machine learning for estimating the object pose in three-dimensional space, in step S55, three or more annotation coordinates are input as annotation data D4 for each object 6. For example, by using the origin of the Cartesian coordinate system and two points serving as references for two mutually orthogonal coordinate axes as the three annotation coordinates, the coordinate axis for the remaining direction can be automatically determined. Furthermore, for example, multiple camera coordinate systems corresponding to multiple cameras may be used, and annotation may be performed for each image from each camera.
[0067] In the control system 1, an estimated value of an object pose can be calculated from a two-dimensional image without, for example, a sensor for acquiring depth information used in Non-Patent Document 1, or reference point cloud data of the object referenced in ICP. Furthermore, since the control system 1 does not require the use of such sensors, it is possible to avoid, for example, the influence of the accuracy of the sensor on the calculation accuracy of the estimated value. Furthermore, in conventional techniques such as feature matching that do not use machine learning for pose estimation, the estimation accuracy depends on the design of the features, but in the control system 1, features for calculating an estimated value with high accuracy can be acquired through machine learning.
[0068] 2-3. Details of the Machine Learning Model Details of the machine learning model 7, which performs machine learning of the posture estimation model 70 in step S16 of Fig. 4, will be described using Fig. 6 and Fig. 7. The various functions of the machine learning model 7 are realized, for example, by the control unit 40 of the control device 4 executing a predetermined program.
[0069] 6A and 6B are diagrams illustrating a machine learning model 7 used in machine learning of a posture estimation model 70. Fig. 6A is a diagram illustrating the operation during learning by the machine learning model 7. For example, the machine learning model 7 shown in Fig. 6A includes a convolutional neural network (CNN) 71 that functions as the posture estimation model 70, a slot initializer 72, and a slot attention module 73.
[0070] 6A further includes an object-centric Newton transition model 74 and a slot decoder 75. The machine learning model 7 learns the parameters of each part by error backpropagation so as to minimize the error in the estimated value of the object pose from the input image and the reconstruction error of the image by the slot decoder 75.
[0071] (1) Slot Attention In the machine learning model 7, the control unit 40 first learns slots corresponding to the latent representations, for example, by a process similar to the slot attention disclosed in Non-Patent Document 3, so as to obtain a latent representation including information about the object 6 in the image. The slots store information about the object 6 acquired by learning.
[0072] The control unit 40 transmits to the CNN 71 an image I of a frame at time t in the image data D1. t is input to output W × H × D dimensional feature quantities, where W, H, and D are positive integers, W and H respectively represent the width and height of the feature map output by the convolution process of the CNN 71, and D represents the number of feature maps.
[0073] The control unit 40 generates positional embeddings that indicate the position on a plane of each feature, and linearly maps the generated W×H×4-dimensional positional embeddings to the same W×H×D dimensions as the feature. The parameters of the linear mapping can be learned together with minimizing the overall error of the machine learning model 7. The control unit 40 adds the linearly mapped positional embeddings to the feature output from the CNN 71, and flattens the resulting W×H×D-dimensional feature to N×D dimensions (N=W×H).
[0074] Furthermore, the control unit 40 functions as a slot initializer 72, which samples an initial slot including an initial value of the slot from a D-dimensional Gaussian distribution. The slot initializer 72 determines parameters (i.e., mean and variance) of the Gaussian distribution for a target slot corresponding to a target object of the pose estimation model 70, for example, using a multi-layer perceptron (MLP) such as that shown in FIG. 6A. The MLP is conditioned by a target object ID that identifies the object 6 as the target object. The target object ID is an example of object identification information D3. The slot initializer 72 may determine the parameters using other neural networks instead of the MLP.
[0075] The slot initializer 72 of this embodiment samples K slots corresponding to the target object specified by the target object ID and the background in the image. For example, in image I in FIG. 0 6A shows an image I of a pair of scissors 61, a hammer 62, and a clamp 63 similar to the image I of FIG. t In the example, three target objects and one background (i.e., K=4) are sampled. The background is, for example, image I t 6A, the slot initializer 72 samples the background slots corresponding to the background from a standard Gaussian distribution without conditioning on identification information such as the target object ID.
[0076] Next, the control unit 40 initializes the sampled initial slot as a slot initializer 72 to the input image I t In the slot attention module 73, the control unit 40 first calculates the attention attn as shown in the following equation (1). In the following description, the query q is a D-dimensional slot, and the key k and value v are the input image I t is an N×D-dimensional feature (inputs) obtained by flattening output values from the above after adding positional embedding. In equation (1), the query q is expressed as a K×D-dimensional matrix including K D-dimensional slots.
[0077] Here, i = 1, 2, ... N, j = 1, 2, ... K, l = 1, 2, ... K. Equation (1) calculates N-dimensional attention for each of the K slots. Attention attn can express the probability that each of the N features belongs to each of the K slots by applying a softmax function to the N x K-dimensional matrix M in the K-dimensional slot direction.
[0078] In the slot attention module 73, the control unit 40 then calculates a K×D-dimensional weighted average "updates" used to update the slot value based on the attention "attn" and the value "v" using the following equation (2). The weight "W" in equation (2) is calculated by normalizing the N×K-dimensional matrix "attn" in the N-dimensional feature direction. The control unit 40 calculates the weighted average "updates" by weighting the value "v" with the weight "W" using the inner product calculation of equation (2).
[0079] The slot attention module 73 includes, for example, a Gated Recurrent Unit (GRU) as a recursive function that is learned to update the slot value. The control unit 40 inputs a weighted average "updates" to the GRU, using the current slot as the hidden state of the GRU, and causes the GRU to output a slot residual. The control unit 40 adds the slot residual output from the GRU to the current slot to update the slot value. The slot attention module 73 performs this update process on the image I t By repeating a predetermined number of times (eg, three times) in , the updated slot is obtained.
[0080] For example, the control unit 40 may t The target slots initialized by the slot initializer 72 for each target object in the image I and the background slots are updated by the slot attention module 73. t From the above, K slots can be obtained, each storing information about an object and a background.
[0081] A feature representation that extracts information about a specific object, such as a target object, as a latent representation obtained separately from information about the background or other objects, such as the K slots described above, is also called an object-centric feature representation. Furthermore, machine learning that acquires such object-centric feature representations is called object-centric learning.
[0082] The number of slots is not limited to the above example, and for example, multiple slots may be sampled for one target object, or there may be multiple background slots. Also, the background slots may be conditioned by, for example, identification information that identifies multiple types of background. t Even if scissors 61, hammer 62, and clamp 63 are shown in the image, if the target object whose pose is to be estimated is only scissors 61, the number of slots to be sampled may be just two: the slot for scissors 61 and the background slot. In this case, hammer 62 and clamp 63 are also treated as background. t If multiple pairs of scissors 61 are shown in the image, slots may be sampled according to the number of pairs of scissors 61 shown. In this case, information about each pair of scissors 61 is stored in each slot. If the number of pairs of scissors 61 shown is unknown, a number of slots that is sufficiently larger than the expected number of scissors 61 may be sampled. In this case, information about each pair of scissors 61 is stored in some slots, and no information is stored in other slots.
[0083] Furthermore, the calculation of feature amounts from an image is not limited to the CNN 71, and various feature extraction algorithms such as SIFT (Scale-Invariant Feature Transform) may be used.
[0084] (2) Object-centric Newtonian Transition Model The slots obtained as object-centric feature representations by the slot attention module 73 include, in addition to information about the object pose, information about the appearance of the object 6 in the image, such as the color and shape. Therefore, the machine learning model 7 of this embodiment uses the object-centric Newtonian transition model 74 to further use the behavioral data D2 to learn the latent representations by constraining a portion of the latent representation in the slot to acquire a higher correlation with the object pose than other portions.
[0085] The object-centric Newtonian transition model 74 is a state transition model that transitions the state of the latent representation so that a part of the latent representation corresponding to the object posture is constrained to change over time according to Newton's equation of motion. For example, the object-centric Newtonian transition model 74 is t The state of the latent expression in the slot obtained from t This causes a transition to the state at the next time t+1.
[0086] 6B is a diagram for explaining the object-centered Newton transition model 74 in the machine learning model 7. First, the control unit 40 calculates the time-varying feature x t and a time-invariant feature value f t For example, when the object posture on a plane is estimated using the posture estimation model 70 with three degrees of freedom, specific three dimensions (for example, the first three dimensions) among the D dimensions of the slot are divided into time-varying feature quantities x t and the remaining dimensions are time-invariant features f t is assigned to.
[0087] Next, the control unit 40 calculates the time-varying feature x at time t. t and behavior t Then, the object-centered Newton transition model 74 calculates the time-varying feature x t+1 Calculate x t = (X t , Y t , θ t ), u t= (v x , v y , ω), then in the object-centered Newton transition model 74, x t+1 = (X t+1 , Y t+1 , θ t+1 ) is calculated according to the following equations (3) and (4).
[0088] where X t and Y t are coordinates indicating the position of the target object in the X and Y directions of the camera coordinate system, respectively, and θ t is the angle corresponding to the rotation of the target object on the XY plane. R is the rotation matrix indicating the rotation on the XY plane. x , v y , ω are the velocity of the hand posture in the X direction, the velocity in the Y direction, and the angular velocity due to rotation on the XY plane, respectively. dt indicates the period from time t to time t+1.
[0089] FIG. 7 is a diagram for explaining the geometric relationship between the camera coordinate system and the object coordinate system. t indicates the camera coordinate system at time t, and O indicates the object coordinate system. In this example, the camera 3 is moved relative to the stationary object 6 to change the object orientation, so C t changes with time, while O is fixed. For example, as shown in FIG. t is the camera coordinate system C t represents the object posture in t It changes depending on
[0090] As described above, the object-centric Newtonian transition model 74 calculates the behavior u corresponding to the change in the relative posture between the camera 3 and the target object. t Depending on the time-varying feature x t is made to follow Newton's equation of motion. t is the camera coordinate system C tBy learning such a latent representation, the latent representation output from the slot attention module 73 of the trained pose estimation model 70 can be trained to represent the object pose of the target object in the time-varying feature x t The value of a specific dimension corresponding to can be interpreted as an estimate of the object pose, and object pose estimation can be realized.
[0091] Furthermore, by conditioning the object-centric Newtonian transition model 74 with the behavioral data D2, it is believed that an accurate estimate of the object's orientation can be obtained even if the object's orientation changes during irregular movements such as a random walk.
[0092] (3) Slot Decoder The machine learning model 7 of this embodiment further uses a slot decoder 75 to generate a latent representation at time t+1 obtained via the object-centric Newton transition model 74, and then generates an image I t+1 For example, the slot decoder 75 synthesizes each reconstructed image obtained from the latent representation by a CNN or the like for each of the K slots, similar to the slot attention technique in Non-Patent Document 3, and reconstructs the reconstructed image I t+1 The time-invariant feature f t For example, the output value from the slot attention module 73 is input to the slot decoder 75 .
[0093] (4) Objective Function The control unit 40 minimizes the loss function L shown in the following equation (5) for each slot, for example, as the objective function of machine learning using the machine learning model 7. T in equation (5) indicates the period during which the image data D1 and the like are collected.
[0094] In equation (5), p represents a probability distribution, q represents an approximate probability distribution in the machine learning model 7, and KL represents the KL divergence. The KL divergence is an example of the distance between probability distributions. The first and third terms of equation (5) represent the distance between the image I t , I 0 The second term of equation (5) represents the time-varying feature x t For the image I at time t,t and the feature value x at the previous time t-1. t-1 and behavior t-1 10 shows the distance from the prior distribution transitioned by the object-centered Newton transition model 74.
[0095] The fourth term of equation (5) is the feature x related to the object orientation in the initial frame of image data D1. 0 About Image I 0 and the annotation coordinate y of the annotation data D4. 0 For example, the control unit 40 calculates the distance between the annotation coordinates y 0 After converting the coordinates of one point into a posture in the camera coordinate system, the prior distribution is calculated using a Gaussian distribution with the converted posture as a parameter. In this case, the prior distribution p(x 0 |y 0 The dispersion parameter of (5) is set to a value corresponding to the expected annotation error (for example, 0.001 meters [m] if the expected annotation error is about 1 mm). The angle θ can be calculated, for example, from the coordinates of two points using the arctangent function. For slots other than the target slot, the fourth term of Equation (5) may be set to "0."
[0096] According to the above-described machine learning using machine learning model 7, a prior distribution of features at a next time point, which is transitioned by object-centric Newton transition model 74, is obtained from the behavior at a time point and the features estimated as the object posture at that time point. Furthermore, a posterior distribution of features at that next time point is obtained from the image at that next time point by slot attention module 73 or the like. Pose estimation model 70 is trained so as to reduce the KL divergence, which is the distance between the prior distribution and the posterior distribution.
[0097] For example, by sequentially minimizing the distance between the two distributions obtained for the object posture for each time in the behavior data D2, information added to some frames of the image data D1 by the annotation data D4 can be propagated to other frames. This allows the posture estimation model 70 to be efficiently trained from the collected training data, improving sample efficiency, for example, without annotating all frames of the image data D1 with respect to the object posture.
[0098] In addition, in the machine learning model 7, for example, the object posture calculated from the annotation data D4 and the annotated image I 0 The pose estimation model 70 is trained so as to reduce the distance between the two distributions corresponding to the object poses estimated from the respective frames. In this way, the annotation data D4 added to only some of the frames can be used as a teaching signal for weakly supervised learning.
[0099] In the above example, pose estimation model 70 is trained to minimize the image reconstruction error by slot decoder 75 in addition to the error related to the object pose estimate. For example, machine learning model 7 does not need to include slot decoder 75. In this case, pose estimation model 70 can be trained using the error related to the object pose estimate without using the image reconstruction error.
[0100] 2-4. Post-Learning Operation In the control system 1, post-learning operation using the trained posture estimation model 70 generated by the machine learning model 7 as described above will be described with reference to FIGS. 8 and 9.
[0101] Fig. 8 is a sequence diagram illustrating an example of post-learning operation of posture estimation model 70 in control system 1. Fig. 9 is a diagram for explaining the operation of posture estimation model 70 after training.
[0102] 4, the control system 1 accepts a user operation to input object identification information D3 at the terminal device 5 (S61). The control device 4 acquires the object identification information D3 input from the terminal device 5 (S62).
[0103] Next, the control system 1 causes the camera 3 to capture an image of the object 6 (S63), and the camera 3 transmits image data representing the captured image to the control device 4.
[0104] The control unit 40 of the control device 4 acquires image data from the camera 3 via the communication I / F 42 (S64). In the example of Fig. 9, image data of an image I in which a hammer 62 or the like is captured as the object 6 is acquired, similar to the example of Fig. 3(B).
[0105] Based on the acquired image data, the control unit 40 calculates an estimated object pose for the target object specified in the object identification information D3 using the trained pose estimation model 70 (S65). For example, as shown in FIG. 9 , in the trained pose estimation model 70, the control unit 40 inputs image data of image I to the CNN 71 and causes the slot attention module 73 to output a latent representation x corresponding to the estimated object pose for the target object. For example, if the target object is a hammer 62, the control unit 40 initializes the target slot using the target object ID that identifies the hammer 62 using the slot initializer 72, and updates the target slot using the slot attention module 73.
[0106] The control unit 40 generates behavior data for driving the manipulator 2b of the robot 2 based on the estimated value of the object posture (S66). By using the estimated value of the object posture, behavior data can be generated so that, for example, the manipulator 2b is driven to grasp the target object with the end effector 2c. In this example, the behavior data is generated as a speed control command for the robot 2 at each time, similar to when the posture estimation model 70 was learned.
[0107] 4, for example, the control unit 40 calculates a control command for changing the posture of the robot 2 based on the generated behavior data, and transmits it to the robot 2 (S67). The control unit 20 of the robot 2 changes the hand posture by driving the manipulator 2b based on the control command received from the control device 4 (S68). In the control system 1, the above processing may be repeatedly executed so that the robot 2 performs an operation according to the object posture.
[0108] According to the above operation of the control system 1, the hand posture of the robot 2 is changed based on the estimated object posture calculated from the image of the object 6 by the trained posture estimation model 70 for the specific target object specified by the object identification information D3 (S61 to S68). In this way, by using the estimated object posture to change the hand posture, it is possible for the robot 2 to grasp the target object with high accuracy, for example.
[0109] 2-5 Verification Experiment A verification experiment related to the first embodiment will be described with reference to Fig. 10. Fig. 10 is a diagram showing the results of a performance evaluation of the trained posture estimation model 70.
[0110] In this performance evaluation, the correlation coefficient between the correct object pose and the estimated value by the trained pose estimation model 70 was evaluated for each target object in six types of evaluation datasets. Each evaluation dataset contains an image of the object and the correct object pose in association with each other, and contains different combinations of two types of background and three types of disturbance objects in the image. Disturbance objects are objects that are not included in the object identification information D3, unlike the target object.
[0111] (A) Calculation Method First, as a method for calculating estimated values of object pose, a method (this method) using a pose estimation model 70 trained by the machine learning model 7 of this embodiment was compared with two baseline methods. FIG. 10(A) shows the correlation coefficient between the estimated value calculated by each method and the correct answer for each of the three degrees of freedom object poses (X, Y, θ) in terms of the average and standard deviation for all target objects and all types of evaluation datasets. In this performance evaluation, "SAM-NVAE" and "OWL-ViT" were used as the two baseline methods.
[0112] In SAM-NVAE, slot attention is not used, and a region corresponding to the target object is cropped from the image using the segmentation model disclosed in Non-Patent Document 4, and the model disclosed in Non-Patent Document 5 is trained. The model in Non-Patent Document 5 is designed to acquire a latent space corresponding to a mapping from an image to physical coordinates and to be used for self-localization of a robot or the like, so this baseline shows correlation coefficients with the sign reversed.
[0113] In OWL-ViT, a bounding box surrounding an object in an image is estimated using the object detection method disclosed in Non-Patent Document 6, and the object pose in the camera coordinate system is calculated from the coordinates on the image. Because the bounding box does not include information on the angle θ, the results are shown only for the X and Y coordinates.
[0114] As shown in Figure 10(A), our method (OC-NVAE) achieved the best performance compared to the two baseline methods. Thus, the latent representation obtained as an object pose estimate by our method appears to correlate well with the actual object pose.
[0115] (B) Evaluation Datasets Next, we evaluated the robustness of our method against disturbances caused by disturbing objects and / or unknown backgrounds using six evaluation datasets. Figure 10(B) shows the correlation coefficient between the estimated object pose calculated by our method and the correct answer for each evaluation dataset, expressed as the average and standard deviation across all target objects. The backgrounds in each image of the evaluation dataset are divided into two types: backgrounds included in the training data images (seenBG) and unknown backgrounds not included in the training data images (unseenBG). Furthermore, the disturbing objects in each image are divided into three types: no disturbing object (target), disturbing objects included in the training data images (known), and disturbing objects not included in the training data images (unknown).
[0116] In Fig. 10(B), a t-test was performed between the evaluation data set without disturbance (seenBG, target) and the other evaluation data sets with disturbance, and items for which there was a significant difference (p value < 0.05) depending on whether or not there was disturbance are underlined. As shown in Fig. 10(B), according to this method, even in the presence of disturbance, the correlation coefficient between the estimated object pose and the correct answer does not decrease significantly from the case without disturbance, and it is believed that object pose estimation that is robust to disturbances is possible.
[0117] 3. Effects, etc. As described above, the control device 4, which is an example of a machine learning device in this embodiment, includes a storage unit 41 and a control unit 40. The storage unit 41 stores image data D1 indicating an image of an object 6 captured by a camera 3, which is an example of an imaging device. The control unit 40 acquires behavioral data D2 indicating the amount of change in object posture, which is an example of the posture of the object 6 as seen from the camera 3 (S12). The image data D1 includes images of multiple frames in which the object posture changes according to the amount of change in the behavioral data D2. The control unit 40 selects an image I of an initial frame from the images of the multiple frames in the image data D1. 0 (an example of an image of at least one frame) 0 The control unit 40 acquires annotation data D4, which is an example of additional information indicating the posture of the object 6 in each of the frames, based on the behavioral data D2, the image data D1, and the annotation data D4 (S15). The control unit 40 performs machine learning on the object posture in each of the multiple frames, based on the behavioral data D2, the image data D1, and the annotation data D4, to generate a trained posture estimation model 70 (an example of an estimation model) (S16). The posture estimation model 70 calculates an estimate of the object posture from the image of the object 6.
[0118] According to the control device 4 described above, the posture estimation model 70 is trained based on the image data D1 including multiple frames of images in which the object posture changes, the behavior data D2 indicating the amount of change in the object posture, and the annotation data D4 related to the object posture (S16). This allows efficient machine learning using the annotation data D4 associated with images of some frames, such as the initial frame, among the multiple frames in the image data D1.
[0119] In this embodiment, the control device 4 further includes a communication I / F 42, which is an example of an input unit that inputs annotation data D4. As illustrated in FIG. 3B, the annotation data D4 is an image I 0 In the example, annotation coordinates y 0 The control unit 40 receives the image I of the initial frame via the communication I / F 42. 0The annotation data D4 is acquired in response to a user operation of annotating two points P1 and P2 of the annotation data D4 (S55, S15, FIG. 5). As a result, the annotation data D4, which is an example of additional information acquired by annotation, can be provided as a teacher signal regarding the posture of the object 6 in the machine learning of the posture estimation model 70.
[0120] In this embodiment, the camera 3 is attached to the manipulator 2b, which is driven based on the behavior data D2 (FIG. 1). As a result, the camera 3 is driven by the manipulator 2b. For example, the hand posture of the manipulator 2b is changed (S20) in accordance with a control command for changing the posture calculated based on the behavior data D2 (S13), and image data D1 showing the change in the posture of the object 6 as seen from the camera 3 as the object posture can be acquired (S1, S30, S14).
[0121] In this embodiment, the control device 4 further includes a communication I / F 42, which is an example of an input unit that inputs object identification information D3 as an example of identification information that identifies at least one object 6. The control unit 40 generates a trained pose estimation model 70 (S16) so as to calculate an estimated value of the object pose of a target object as an example of the object 6 identified by the object identification information D3. This makes it possible to specify a specific object 6 as the target object using the object identification information D3, for example, and calculate an estimated value of the object pose of the target object using the trained pose estimation model 70.
[0122] In this embodiment, the storage unit 41 stores a machine learning model 7 ( FIG. 2 ), which is an example of a learner used in machine learning of the posture estimation model 70. The machine learning model 7 includes an object-centered Newtonian transition model 74 ( FIG. 6 ), which is an example of a state transition model that transitions an object posture at a first time to an object posture at a second time in accordance with the amount of change in the object posture indicated by the behavior data D2. The machine learning model 7 enables efficient machine learning of the object posture in each of the multiple frames of the image data D1.
[0123] In this embodiment, as shown in Fig. 6, the control unit 40 generates a trained posture estimation model 70 based on the multiple frame images of the image data D1 and the machine learning model 7 so as to reduce the distance defined between the prior distribution and the posterior distribution (S16). The prior distribution is determined based on the image I captured at a first time among the multiple frame images. t , the latent representation x corresponding to the estimated value calculated by the posture estimation model 70 during training. t is a latent representation x corresponding to the estimated value at the second time point, which is transitioned by the object-centric Newton transition model 74 in the machine learning model 7. t+1 The posterior distribution is the probability distribution of the image I captured at the second time among the images of the plurality of frames. t+1 , the latent representation x corresponding to the estimated value calculated by the posture estimation model 70 during training. t+1 This shows the probability distribution of . This allows information about the object posture based on the annotation data D4, which is assigned to a frame at a certain time, such as an initial frame in the image data D1, to be propagated to frames at other times, improving the sample efficiency of machine learning. This reduces the annotation load, for example, when inputting the annotation data D4 (S55).
[0124] In this embodiment, the object-centric Newtonian transition model 74 transitions the object posture according to Newton's equation of motion. As a result, for example, the latent representation transitioned by the object-centric Newtonian transition model 74 can be acquired to represent the object posture. Furthermore, such rule-based state transitions can facilitate rapid propagation of the teacher signal from the annotation data D4 and make it less susceptible to attenuation throughout the entire trajectory in the behavior data D2 corresponding to multiple frames of the image data D1.
[0125] In this embodiment, the control unit 40 calculates a prior distribution of the object posture calculated based on the annotation data D4 and an image I of an initial frame, which is an example of at least one frame. 0A trained posture estimation model 70 is generated so as to reduce the distance between the posterior distribution of the estimated value calculated by posture estimation model 70 during training and annotation data D4 (S16). This allows annotation data D4 to be used as a teacher signal in the machine learning of posture estimation model 70.
[0126] In this embodiment, the pose estimation model 70 has a slot, which is an example of an information storage unit, that stores information including the object pose of the target object (an example of the object 6 identified by the object identification information D3), and the slot is conditioned by the object identification information D3 (S16, S65). In the examples of FIGS. 6A and 9 , the pose estimation model 70 has a target slot acquired by the slot attention module 73, and the target slot is conditioned by the target object ID, which is an example of the object identification information D3, in the slot initializer 72. By using the object identification information D3 in this way, for example, the target slot can be acquired as a latent representation that stores information including the object pose of the target object.
[0127] In this embodiment, the object orientation includes the relative position and relative angle between the camera 3 and the object 6. Such an object orientation is an example of the relative orientation between the camera 3 and the object 6.
[0128] In this embodiment, the control device 4, which is an example of an estimation device, includes a storage unit 41 and a control unit 40. The storage unit 41 stores a trained pose estimation model 70 generated by the control device 4, which is an example of a machine learning device. The control unit 40 inputs an image I of the object 6 into the trained pose estimation model 70, and causes the trained pose estimation model 70 to output a latent representation x corresponding to an estimated value of the object pose (S65, FIG. 9 ). For example, by operating the robot 2 using the estimated value of the object pose (S66 to S68), the robot 2 can accurately manipulate the object 6.
[0129] The machine learning method (S11 to S16) in this embodiment is executed by the control unit 40 of the control device 4, which is an example of a computer control unit. This method includes a step of acquiring image data D1 showing an image of an object 6 captured by a camera 3, which is an example of an imaging device (S14), and a step of acquiring behavior data D2 showing an amount of change in object posture, which is an example of the posture of the object 6 as seen from the camera 3 (S12). The image data D1 includes images of multiple frames in which the object posture changes according to the amount of change in the behavior data D2. The control unit 40 acquires an image I of an initial frame as an example of an image of at least one frame of the images of the multiple frames in the image data D1. 0 In association with the image I 0 The control unit 40 acquires annotation data D4, which is an example of additional information indicating the posture of the object 6 in each of the frames, based on the behavioral data D2, the image data D1, and the annotation data D4 (S15). The control unit 40 performs machine learning on the object posture in each of the multiple frames, based on the behavioral data D2, the image data D1, and the annotation data D4, to generate a trained posture estimation model 70 (an example of an estimation model) (S16). The posture estimation model 70 calculates an estimate of the object posture from the image of the object 6.
[0130] In this embodiment, a program for causing a control unit of a computer to execute the above-described machine learning method is provided. The machine learning method of this embodiment allows efficient machine learning of the pose estimation model 70 that estimates the pose of the object 6 from an image.
[0131] 11 to 15, a second embodiment of the present disclosure will be described below. In the first embodiment, a control system 1 was described in which annotation data D4 generated by a user operation of a terminal device 5 is used for machine learning of a posture estimation model 70. In the second embodiment, a control system 1A will be described in which the posture of a marker attached to an object 6 is used instead of the annotation data D4.
[0132] Hereinafter, the description of the configuration and operation similar to those of the control system 1 of the first embodiment will be omitted as appropriate, and the control system 1A of this embodiment will be described.
[0133] 1. Configuration FIG. 11 is a diagram illustrating an overview of a control system 1A according to a second embodiment. In the control system 1A, a marker 8 is attached to an object 6 before collecting learning data for the pose estimation model 70. For example, the marker 8 is an AR (Augumented Reality) marker. In the control system 1A, a marker pose measurement unit 9 measures the pose of the marker 8 relative to the marker pose measurement unit 9, and from the measurement result, a marker pose corresponding to the object pose is obtained as the pose of the marker 8 as seen from the camera 3. In the control system 1A, the marker pose measurement unit 9 is calibrated with respect to the camera 3 in advance.
[0134] FIG. 12 is a block diagram illustrating the configuration of the robot 2, the camera 3, the marker attitude measurement unit 9, the control device 4, and the terminal device 5 in the control system 1A of the second embodiment.
[0135] The control system 1A of this embodiment includes a marker orientation measurement unit 9 in addition to the same configuration as the control system 1 of the first embodiment illustrated in Fig. 1. The marker orientation measurement unit 9 includes a communication I / F 92 and is capable of data communication with external devices such as the control device 4 and the terminal device 5. The communication I / F 92 is a circuit that performs data communication in accordance with a predetermined communication standard, similar to, for example, the communication I / F 32 of the robot 2. The marker orientation measurement unit 9 includes, for example, a camera that captures an image of the marker 8. The marker orientation measurement unit 9 may include, for example, a processor that applies image recognition processing to the image.
[0136] Furthermore, in the control system 1A, instead of the annotation data D4 of the first embodiment, marker posture data D5 acquired from, for example, the marker posture measurement unit 9 is stored in the storage unit 41 of the control device 4. The marker posture data D5 indicates, for example, the measurement result of the marker posture, and is an example of additional information in this embodiment.
[0137] 2. Operation FIG. 13 is a sequence diagram illustrating an example of the operation of the control system 1A of the second embodiment when learning the posture estimation model 70.
[0138] The control system 1A of this embodiment executes processes related to acquisition of marker postures (S50, S90, S40) instead of the process related to acquisition of annotation data D4 (S55) in the learning operation ( FIG. 4 ) of embodiment 1. Furthermore, the control device 4 of the control system 1A executes machine learning of a posture estimation model 70 using marker posture data D5 (S16A) instead of the machine learning using annotation data D4 (S16).
[0139] For example, the control unit 50 of the terminal device 5 accepts a user operation to input a marker orientation reading command on the operation unit 53 (S50). The reading command is a command to instruct the marker orientation measurement unit 9 to read the marker orientation. The control unit 50 transmits the input reading command to the marker orientation measurement unit 9 via the communication I / F 52.
[0140] When the marker orientation measurement unit 9 receives a read command from the terminal device 5 via the communication I / F 92, it captures an image of the marker 8 to read the marker orientation (S90). The marker orientation measurement unit 9 calculates the marker orientation from the captured image of the marker 8 by image recognition processing, for example, and generates marker orientation data D5.
[0141] The control unit 40 of the control device 4 acquires the marker attitude data D5 from the marker attitude measurement unit 9 via the communication I / F 42 (S40).
[0142] The control system 1A, similar to the control system 1 of the first embodiment, acquires object identification information D3 input in response to a user operation on the terminal device 5 (S51). The control system 1A also repeats the data collection operation (S1) of image data D1 and behavioral data D2.
[0143] In the data collection operation (S1) of the control system 1A, the markers 8 are removed from the object 6, an image is captured without the markers 8 in view (S30), and image data D1 is acquired (S14). For example, when capturing an initial frame of image data D1 (S1, S30), the control unit 40 generates behavior data D2 (S12) so that the frame is captured in the same object orientation as when the marker orientation was read (S90). The control unit 40 stores the acquired image data of the initial frame in the storage unit 41, for example, in association with marker orientation data D5.
[0144] In the control system 1A, the control unit 40 performs machine learning of the posture estimation model 70 using the collected image data D1 and behavior data D2 as well as the marker posture data D5 (S16A). The control unit 40 may calculate the posture with respect to the camera 3 from the marker posture with respect to the marker posture measurement unit 9, for example, based on the marker posture data D5. The control unit 40 calculates the annotation coordinate y 0 Instead of the above, for example, based on the orientation calculated from the marker orientation, a feature amount x 0 In this way, in this embodiment, the prior distribution of the annotation coordinate y 0 Instead, for example, a marker pose corresponding to the object pose at the time of capturing the initial frame is used as a training signal for machine learning.
[0145] According to the above-described operation of control system 1A, marker posture data D5 is acquired by measuring the marker postures by reading markers 8 attached to object 6 (S50, S90, S40). Control system 1A performs machine learning of posture estimation model 70 based on image data D1, behavioral data D2, and marker posture data D5, using the marker postures as teacher signals (S16A). As a result, posture estimation model 70 can be efficiently trained using marker postures corresponding to object postures, for example, without requiring user operation to annotate the posture of the target object of posture estimation model 70 on the image.
[0146] Furthermore, in the control system 1A, when collecting image data D1, for example, the markers 8 are removed from the object 6, and image data representing an image captured without the markers 8 is acquired (S30, S14). This allows the pose estimation model 70 to be trained from image data D1 that does not include the markers 8 (S16A), and, similar to the operation after training in the first embodiment, for example, an estimated value of the object pose can be calculated from the image that does not include the markers 8 (S65 in FIG. 8).
[0147] Furthermore, the control system 1A acquires marker poses corresponding to the object poses at the time of capturing a part of the image data D1, such as the initial frame, in association with the image (S90, S40, S1). In this way, the pose estimation model 70 can be trained without acquiring marker poses associated with each image of the image data D1. This avoids the inefficiency of capturing an image including the marker 8 and an image not including the marker 8 each time the hand pose is changed when collecting the image data D1, and allows the pose estimation model 70 to calculate an estimated value of the object pose from an image not including the marker 8.
[0148] The marker orientation measurement unit 9 is not limited to the above example, and may, for example, transmit image data indicating a captured image of the marker 8 to the control device 4. In this case, in step S40 of Fig. 13, the control unit 40 of the control device 4 may calculate the marker orientation from the received image data. Also, the camera 3 may be used as the marker orientation measurement unit 9. For example, instead of step S50, a command to read the marker orientation may be transmitted from the terminal device 5 to the camera 3, and instead of step S90, image data from the camera 3 may be transmitted to the control device 4.
[0149] The marker 8 and the marker orientation measurement unit 9 are not limited to the above examples, and may be implemented as various optical motion capture or tracking systems. For example, the marker 8 is not limited to an AR marker, but may be an OptiTrack (registered trademark) marker or the like, and the marker orientation measurement unit 9 may include multiple cameras. Furthermore, the marker 8 and the marker orientation measurement unit 9 may be integrally configured, and for example, a VIVE (registered trademark) tracker used for VR (Virtual Reality) applications may be attached to the object 6, and the tracking result may be acquired as the marker orientation.
[0150] As described above, the control device 4 of this embodiment further includes a communication I / F 42, which is an example of an acquisition unit that acquires the posture of the marker 8 attached to the object 6 as seen from the camera 3. In this embodiment, the annotation data D4, which is an example of additional information, is an image I of an initial frame in the image data D1, which is an example of at least one image. 0 This includes the orientation of the marker 8 at the time of capturing the image. As a result, for example, instead of the annotation data D4 in the first embodiment, the orientation of the marker 8 acquired as the marker orientation data D5 can be used to provide information about the object orientation to the machine learning model 7 (S41, S16A). This, for example, can further reduce the annotation load.
[0151] In this embodiment, multiple frames of image data D1 are captured without capturing markers 8 (S1, S30), and trained pose estimation model 70 calculates an estimated value of the object pose from the images captured without capturing markers 8 (S65). In this way, even when the poses of markers 8 are used in the machine learning of pose estimation model 70, trained pose estimation model 70 can be generated so that an estimated value of the object pose is calculated without using markers 8.
[0152] In the above example, the control system 1A acquires object identification information D3 input by a user operation on the terminal device 5 (S51). The object identification information D3 may be acquired, for example, using a marker 8 attached to the object 6 for acquiring the marker orientation. Such a first modified example of the second embodiment will be described with reference to FIG. 14 .
[0153] 14 is a sequence diagram illustrating an example of an operation during learning of the posture estimation model 70 in the control system 1A according to the first modification of embodiment 2. The control system 1A of this modification performs a process (S41) of acquiring object identification information D3 from the marker 8 in the control device 4, instead of a process (S51) of inputting object identification information D3 by a user operation.
[0154] In this modification, for example, when reading the posture of the marker 8 (S90), the marker posture measurement unit 9 acquires marker information such as an ID that identifies the marker 8. The marker information is recognized by, for example, image recognition processing on an image of the marker 8.
[0155] The control unit 40 of the control device 4 receives the marker information of the marker 8, for example, from the marker attitude measurement unit 9, via the communication I / F 42, and acquires the object identification information D3 based on the marker information (S41). The control unit 40 may use the received marker information as the object identification information D3, or may generate the object identification information D3 from the marker information using a predetermined conversion algorithm or the like.
[0156] According to the operation of control system 1A of this modification, for example, even without inputting object identification information D3 by a user operation, posture estimation model 70 can be trained to identify a target object using object identification information D3 obtained from marker information and calculate an estimated value of the object posture (S16A). In this way, for example, marker 8 used to acquire marker posture data D5 can also be used to acquire object identification information D3 (S40, S41), thereby enabling efficient machine learning of posture estimation model 70.
[0157] In the above example, when acquiring object identification information D3, the control unit 40 acquires recognized marker information from the marker orientation measurement unit 9 (S41). The control unit 40 may, for example, acquire image data of the marker 8 from the marker orientation measurement unit 9 and recognize the marker information from the image data.
[0158] As described above, in this modification, the control unit 40 acquires object identification information D3 (an example of identification information for identifying at least one object 6) based on marker information such as the ID of the marker 8 via the communication I / F 42 (S41). The marker information is an example of an identifier possessed by the marker 8. The control unit 40 generates a trained estimation model to calculate an estimated object posture of the target object identified by the object identification information D3 (S16A). Such a trained posture estimation model 70 can calculate an estimated object posture for a specific object 6 designated as the target object by the object identification information D3 using, for example, marker information.
[0159] In the above example, the control system 1A acquires the orientation of the marker 8 directly attached to the object 6 as the marker orientation (S90, S40). The marker 8 may be attached to, for example, a simulated effector that simulates the end effector 2c that comes into contact with the object 6, and the orientation of the marker 8 on the simulated effector may be acquired as the marker orientation. Such a second modification of the second embodiment will be described with reference to FIG. 15 .
[0160] 15 is a diagram illustrating an overview of a control system 1A according to a second modification of the second embodiment. In this modification, when acquiring the marker posture, for example, a marker 8 is attached to the simulated effector 12 in a state in which the simulated effector 12 is gripping an object 6. The simulated effector 12 has the same structure and / or shape as, for example, the end effector 2c, and is an example of a jig having a structure in common with the end effector 2c. The simulated effector 12 may be a clamp that grips the object 6, or the like.
[0161] The term "structure" here refers to the mechanical structure of the main components that act on an object with the end effector 2c. For example, if the end effector 2c is a parallel gripper, the structure includes the shape, size, and opening width of the claws. If the end effector 2c is an underactuated gripper, the structure also includes the relative positions of the knuckles. If the end effector 2c is a suction gripper, the structure includes the shape, size, arrangement, and number of suction pads. If the robot 2 is a welding robot, the structure includes the shape and size of the torch tip.
[0162] In the control system 1A of this modified example, the marker posture measurement unit 9 reads the marker posture from the marker 8 attached to such a simulated effector 12 (S90), and the control device 4 acquires marker posture data D5 indicating the marker posture (S40).
[0163] For example, by using the simulated effector 12, the marker orientation when the object 6 is grasped can be obtained. By using such a marker orientation as a teacher signal for the object orientation in learning the orientation estimation model 70, when the robot 2 is made to grasp the object 6 by the end effector 2c based on the estimated value of the object orientation, it becomes easier to calculate an estimated value suitable for grasping the object 6. Furthermore, for example, when the target object is scissors, the user can teach the object orientation that he or she wants the robot 2 to grasp by attaching a marker 8 to the handle of the scissors if the scissors grasped by the robot 2 are to be used for cutting, or to the blade of the scissors if the scissors are to be handed over to a user.
[0164] As described above, by attaching the markers 8 to the simulated effector 12, it is possible to acquire marker postures according to tasks that use estimated values of object postures, and use these for machine learning of the posture estimation model 70. For example, a trained posture estimation model 70 can be obtained that calculates estimated values of object postures that reflect the user's intentions in various tasks, depending on the positions at which the markers 8 are attached to the object 6 by the simulated effector 12.
[0165] As described above, in this modification, the manipulator 2b is connected to the end effector 2c. The communication I / F 42 of the control device 4 is an example of an acquisition unit that acquires the attitude of the marker 8, as seen from the camera 3, attached to the dummy effector 12, which is an example of a jig having a structure common to the end effector 2c. The marker attitude data D5, which is an example of additional information, is acquired from the initial frame image I, which is an example of an image of at least one frame. 0 This makes it possible to acquire the marker orientation data D5 so that an estimated value of the object orientation according to a task using the end effector 2c can be easily obtained.
[0166] (Other Embodiments) As described above, embodiments 1 and 2 have been described as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited to these and can be applied to embodiments in which modifications, substitutions, additions, omissions, etc. are made as appropriate. Furthermore, it is also possible to combine the components described in each of the above embodiments to create a new embodiment. Therefore, other embodiments will be described below as examples.
[0167] In the first embodiment, the annotation data D4 is stored in the image I 0 Annotation coordinate y indicates a predetermined representative point of the object 6 in 0 In this embodiment, the annotation data D4 is acquired as, for example, the image I 0 Alternatively, the annotation data D4 may be acquired as coordinates of a plurality of points indicating a partial region such as a bounding box surrounding the object 6 in the image I. 0 In the above, the annotation data D4 may be acquired as a set of coordinates indicating a region corresponding to the object 6, such as a segmentation mask of the object 6. The annotation data D4 may be acquired by combining two or more of the representative points, the bounding box, and the segmentation mask.
[0168] As described above, in each of the above embodiments, the annotation data D4, which is an example of additional information, is an image I, which is an example of an image of at least one frame. 0 The additional information includes at least one of coordinates indicating two points P1 and P2 that are an example of representative points of the object 6, coordinates of a plurality of points that indicate a bounding box as an example of a partial region of a predetermined shape that surrounds the object 6, and a set of coordinates that indicate a segmentation mask that is an example of a region that corresponds to the object 6. In this way, various types of annotation data can be used as additional information.
[0169] In each of the above embodiments, the pose estimation model 70 is trained to calculate, as the object pose, an estimate of the position and angle of the object 6 relative to the camera 3. The pose estimation model may also be trained to calculate, as the object pose, an estimate of, for example, one of the position and angle relative to the camera 3.
[0170] As described above, in each of the above embodiments, the object posture includes at least one of the relative position and the relative angle between the camera 3, which is an example of an imaging device, and the object 6. For example, even when the object posture is one of the relative position and the relative angle, machine learning of a posture estimation model can be performed using annotation data D4 or marker posture data D5 according to the position or angle. Furthermore, for an object that includes a joint, the object posture may be the posture of a part of the object that has changed due to the movement of the joint.
[0171] In each of the above embodiments, the control system 1, 1A generated common behavior data D2 for K slots so as to move the camera 3 relative to the stationary object 6 (S12). In this embodiment, for example, behavior data may be generated for each of multiple target objects of the pose estimation model 70. The machine learning model 7 may use the behavior data for each target object to acquire, for example, a latent representation of a slot corresponding to each target object.
[0172] In each of the above embodiments, the machine learning model 7 acquires a latent representation as an object-centric feature representation by using the slot attention module 73, etc. The object-centric feature representation is not limited to a process similar to the above-mentioned slot attention, but may also be acquired by, for example, cropping a region corresponding to a target object in an image using a trained model disclosed in Non-Patent Document 4, and calculating features from the cropped image.
[0173] In each of the above embodiments, the behavior data D2 was generated so that the behavior at each time indicates the velocity of the robot 2 corresponding to the amount of change in the object posture (S12). In this embodiment, the behavior data D2 may include, as the behavior at each time, not only the velocity at which the robot 2 is operated but also the acceleration. In this case, the posture estimation model 70 can be trained by using, for example, a state transition model as the object-centered Newtonian transition model 74 in which the object posture, including the acceleration, transitions in accordance with Newton's equation of motion.
[0174] In the above embodiments, an example has been described in which the robot 2 grasps an object 6 using an object pose estimated by the pose estimation model 70. The robot 2's task using the object pose estimate may be, for example, the adsorption of the object 6 by the end effector 2c, or an insertion operation, such as changing the pose of one grasped or adsorbed object 6 to match the pose of another object 6. Furthermore, for example, the calculated object pose estimate may be visualized on the display unit 54 of the terminal device 5. The control system 1, 1A is not limited to the robot 2 illustrated in FIG. 1 , but may be applicable to the control of various robots that perform various operations. For example, the control system of the present disclosure may be applied to the control of a robot whose end effector does not come into contact with the object 6, such as for visual inspection of a product. For example, the present disclosure may be applied to a system that recognizes the pose of the object 6, moves the camera 3 to an appropriate imaging position for the visual inspection, and captures an image for the visual inspection.
[0175] In the above embodiments, an example has been described in which the trained attitude estimation model 70 is used in the control system 1, 1A to control the robot 2. The attitude estimation model of the present disclosure may be trained in a control system that controls various moving objects and may be used to control the moving objects. For example, the attitude estimation model may be used to estimate the attitude of surrounding objects as seen from a camera mounted on various types of autonomous vehicles or ships, or may be used to estimate the attitude of landing facilities as seen from a camera on an unmanned aerial vehicle such as a drone.
[0176] As described above, the embodiments have been described as examples of the technology in the present disclosure, and for that purpose, the accompanying drawings and detailed description have been provided.
[0177] Therefore, the components shown in the accompanying drawings and detailed description may include not only essential components for solving the problem, but also components that are not essential for solving the problem in order to illustrate the above technology. Therefore, the fact that these non-essential components are shown in the accompanying drawings or detailed description should not be interpreted as immediately indicating that these non-essential components are essential.
[0178] Furthermore, since the above-described embodiments are intended to illustrate the technology of the present disclosure, various modifications, substitutions, additions, omissions, etc. may be made within the scope of the claims or their equivalents.
[0179] (Summary of Aspects) Various aspects of the present disclosure are listed below.
[0180] A first aspect of the present disclosure is a machine learning device including a storage unit that stores image data representing an image of an object captured by an imaging device, and a control unit that acquires behavioral data indicating a change in the posture of the object as seen from the imaging device. The image data includes multiple frames of images in which the posture of the object changes according to the change in the behavioral data. The control unit acquires additional information associated with at least one frame of the image data and indicating the posture of the object in the image, and performs machine learning on the posture of the object in each of the multiple frames based on the behavioral data, the image data, and the additional information to generate a trained estimation model. The estimation model calculates an estimate of the posture of the object from the image of the object.
[0181] In a second aspect, the machine learning device of the first aspect further includes an input unit that inputs additional information. The additional information includes, in at least one frame of an image, at least one of coordinates indicating a representative point of an object, coordinates of a plurality of points indicating a partial region of a predetermined shape surrounding the object, and a set of coordinates indicating a region corresponding to the object. The control unit acquires the additional information via the input unit in response to a user operation to annotate the additional information to at least one frame of the image.
[0182] In a third aspect, the machine learning device of the first aspect further includes an acquisition unit that acquires a posture of a marker attached to the object as seen from the imaging device, wherein the additional information includes a posture of the marker at the time of capturing at least one image.
[0183] In a fourth aspect, in the machine learning device of the third aspect, the multiple frames of images are captured in a state in which no markers are visible, and the trained estimation model calculates an estimate of the pose of the object from the images captured in a state in which no markers are visible.
[0184] In a fifth aspect, in the machine learning device of the third or fourth aspect, the control unit acquires, via the acquisition unit, identification information for identifying at least one object based on an identifier held by the marker, and the control unit generates a trained estimation model to calculate an estimated value of the posture of the object identified by the identification information.
[0185] In a sixth aspect, in the machine learning device of any one of the first to fifth aspects, the imaging device is attached to a manipulator that is driven based on behavioral data.
[0186] In a seventh aspect, in the machine learning device of the sixth aspect, an end effector is connected to the manipulator. The machine learning device further includes an acquisition unit that acquires the orientation of a marker attached to a jig having a structure common to that of the end effector, as viewed from the imaging device. The additional information includes the orientation of the marker when capturing at least one frame of image.
[0187] In an eighth aspect, the machine learning device according to any one of the first to seventh aspects further includes an input unit configured to input identification information for identifying at least one object, and the control unit generates a trained estimation model to calculate an estimate of the pose of the object identified by the identification information.
[0188] In a ninth aspect, in the machine learning device of any one of the first to eighth aspects, the storage unit stores a learning device used for machine learning of the estimation model, and the learning device includes a state transition model that transitions a posture of the object at a first time to a posture of the object at a second time in accordance with an amount of change in the posture of the object indicated by the behavioral data.
[0189] In a tenth aspect, in the machine learning device of the ninth aspect, the control unit generates a trained estimation model based on images of multiple frames and the learning device so as to reduce a distance defined between a prior distribution and a posterior distribution. The prior distribution represents a probability distribution of estimated values calculated by the estimation model under training from an image captured at a first time among the images of the multiple frames, and transitions to a second time according to a state transition model in the learning device. The posterior distribution represents a probability distribution of estimated values calculated by the estimation model under training from an image captured at a second time among the images of the multiple frames.
[0190] In an eleventh aspect, in the machine learning device according to the ninth or tenth aspect, the state transition model transitions the posture of the object in accordance with Newton's equation of motion.
[0191] In a twelfth aspect, in the machine learning device of any of the first to eleventh aspects, the control unit generates a trained estimation model so that the distance between the object posture calculated based on the additional information and the estimated value calculated by the estimation model being trained from at least one frame of image is small.
[0192] In a thirteenth aspect, in the machine learning device of the fifth or eighth aspect, the estimation model has an information storage unit that stores information including the posture of an object identified by identification information, and the information storage unit is conditioned by the identification information.
[0193] In a fourteenth aspect, in the machine learning device of any one of the first to thirteenth aspects, the orientation of the object includes at least one of a relative position and a relative angle between the imaging device and the object.
[0194] A fifteenth aspect is an estimation device including: a memory unit that stores a trained estimation model generated by the machine learning device of any one of the first to fourteenth aspects; and a control unit that inputs an image of an object into the trained estimation model and outputs an estimate of the object's posture.
[0195] A sixteenth aspect is a machine learning method executed by a computer control unit, including the steps of acquiring image data representing an image of an object captured by an imaging device and acquiring behavioral data representing a change in the posture of the object as seen from the imaging device. The image data includes multiple frames of images in which the posture of the object changes according to the change in the behavioral data. In this method, the control unit acquires additional information associated with at least one of the multiple frames of images in the image data, representing the posture of the object in the image, and performs machine learning on the posture of the object in each of the multiple frames based on the behavioral data, image data, and additional information to generate a trained estimation model. The estimation model calculates an estimate of the posture of the object from the image of the object.
[0196] A seventeenth aspect is a program for causing a control unit of a computer to execute the machine learning method of the sixteenth aspect.
[0197] The present disclosure is applicable to various control systems that control machine learning of estimation models for estimating the pose of an object from an image.
Claims
1. A machine learning device comprising: a memory unit that stores image data showing an image of an object captured by an imaging device; and a control unit that acquires behavioral data showing an amount of change in the posture of the object as seen from the imaging device, wherein the image data includes multiple frames of images in which the posture of the object changes according to the amount of change in the behavioral data, and the control unit acquires additional information that indicates the posture of the object in at least one frame of images in the image data, by associating the additional information with the image of the multiple frames, and performs machine learning on the posture of the object in each of the multiple frames based on the behavioral data, the image data, and the additional information to generate a trained estimation model, and the estimation model calculates an estimate of the posture of the object from the image of the object.
2. The machine learning device of claim 1, further comprising an input unit for inputting the additional information, wherein the additional information includes at least one of coordinates indicating a representative point of the object in the at least one frame of image, coordinates of a plurality of points indicating a partial area of a predetermined shape surrounding the object, and a set of coordinates indicating an area corresponding to the object, and the control unit acquires the additional information via the input unit in response to a user operation to annotate the additional information to the at least one frame of image.
3. The machine learning device according to claim 1, further comprising an acquisition unit that acquires the orientation of a marker attached to the object as seen from the imaging device, wherein the additional information includes the orientation of the marker at the time of capturing the at least one image.
4. The machine learning device according to claim 3, wherein the multiple frames of images are captured in a state in which the markers are not visible, and the trained estimation model calculates an estimated value of the pose of the object from the images captured in a state in which the markers are not visible.
5. The machine learning device according to claim 3, wherein the control unit acquires, via the acquisition unit, identification information that identifies at least one object based on an identifier held by the marker, and generates the trained estimation model so as to calculate an estimated value of the posture of the object identified by the identification information.
6. The machine learning device according to claim 1, wherein the imaging device is attached to a manipulator that is driven based on the behavioral data.
7. The machine learning device according to claim 6, wherein an end effector is connected to the manipulator, the machine learning device further comprises an acquisition unit that acquires the orientation of a marker attached to a jig having a structure common to that of the end effector, as viewed from the imaging device, and the additional information includes the orientation of the marker at the time of capturing the at least one frame of image.
8. The machine learning device according to claim 1, further comprising an input unit for inputting identification information for identifying at least one object, wherein the control unit generates the trained estimation model so as to calculate an estimated value of the posture of the object identified by the identification information.
9. The machine learning device of claim 1, wherein the memory unit stores a learning device used for machine learning of the estimation model, and the learning device includes a state transition model that transitions the posture of the object at a first time to the posture of the object at a second time in accordance with the amount of change in the posture of the object indicated by the behavioral data.
10. The machine learning device described in claim 9, wherein the control unit generates the trained estimation model based on the images of the multiple frames and the learning device so that the distance defined between the prior distribution and the posterior distribution is small, the prior distribution indicates a probability distribution of estimated values calculated by the estimation model under training from an image of the multiple frame images captured at the first time, and transitions to the second time by the state transition model in the learning device, and the posterior distribution indicates a probability distribution of estimated values calculated by the estimation model under training from an image of the multiple frame images captured at the second time.
11. The machine learning device according to claim 9, wherein the state transition model transitions the posture of the object according to Newton's equation of motion.
12. The machine learning device of claim 1, wherein the control unit generates the trained estimation model so that the distance between the posture of the object calculated based on the additional information and the estimated value calculated by the estimation model being trained from the at least one frame of image is small.
13. The machine learning device according to claim 8, wherein the estimation model has an information storage unit that stores information including the posture of an object identified by the identification information, and the information storage unit is conditioned by the identification information.
14. The machine learning device according to claim 1, wherein the orientation of the object includes at least one of a relative position and a relative angle between the imaging device and the object.
15. An estimation device comprising: a memory unit that stores the trained estimation model generated by the machine learning device according to any one of claims 1 to 14; and a control unit that inputs an image of the object into the trained estimation model and outputs an estimated value of the posture of the object.
16. A machine learning method executed by a control unit of a computer, comprising: steps of acquiring image data showing an image of an object captured by an imaging device; and acquiring behavioral data showing an amount of change in the posture of the object as seen from the imaging device, wherein the image data includes images of multiple frames in which the posture of the object changes according to the amount of change in the behavioral data, and the control unit acquires additional information showing the posture of the object in at least one frame of images in the image data, associated with the image of the multiple frames, and performing machine learning on the posture of the object in each of the multiple frames based on the behavioral data, the image data, and the additional information, to generate a trained estimation model, and the estimation model calculates an estimate of the posture of the object from the image of the object.
17. A program for causing a control unit of a computer to execute the machine learning method according to claim 16.
Citation Information
Patent Citations
End effector control method
JP2023131026A
Learning data generation apparatus, learning system, and learning data generation method
JP2023169481A
Data collection device and control device
JP2024005599A