Machine learning device and method, and robot control device

WO2026204425A1PCT designated stage Publication Date: 2026-10-01PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/009654
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-12
Publication Date
2026-10-01

Smart Images

  • Figure JP2026009654_01102026_PF_FP_ABST
    Figure JP2026009654_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This machine learning device comprises: a communication unit that is communicably connected to an external device including an imaging device that captures an image of a target object; and a control unit that controls the external device via the communication unit. The target object includes a first portion having a physical property that causes a change in appearance different from a change in posture, and a second portion in which a change in appearance due to the physical property is less likely to occur as compared with the first portion. The control unit causes, via the communication unit, the imaging device to capture a first captured image including the target object in a first imaging posture of the imaging device when viewing the target object. The control unit causes, via the communication unit, the imaging device to capture a second captured image in which the change in appearance due to the physical property has occurred in the first portion of the target object, as compared with the first captured image, in a second imaging posture different from the first imaging posture. The control unit generates, on the basis of the first and second captured images and the first and second imaging postures, at least one of a training data set and a trained model in machine learning in which the image of the second portion of the target object is recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning apparatus and method, and robot control device

[0001] This disclosure relates to a machine learning apparatus and method for machine learning in image recognition, as well as a robot control device.

[0002] Patent Document 1 discloses an inference system for virtual production that outputs a mask image extracted from a captured image, based on a non-extraction target image such as a rendered image of 3D background data and a captured image. The inference system undergoes pre-training using machine learning techniques to infer the foreground region excluding the fixed background region in an image captured by a fixed camera, and calibration training. In calibration, the inference system learns to output a mask image based on the input non-extraction target image and a second captured image of the said non-extraction target image. The second captured image is taken in a situation where no physical objects such as performers are placed. Thus, Patent Document 1 achieves the objective of absorbing changes in the lighting environment and inferring the extraction target region with high accuracy.

[0003] Non-Patent Document 1 discloses a novel model for image segmentation called SAM (Segment Anything Model). Non-Patent Document 1 constructs the model using a segmentation dataset (SA-1B) containing 11 million images and more than 1 billion masks. SAM is designed and trained in a promptable manner, allowing it to transition to new image distributions and tasks in zero shots.

[0004] International Publication No. 2024 / 042893

[0005] Alexander Kirillov, et al., "Segment Anything", arXiv: 2304.02643 [cs.CV], 2023.F. Locatello, et al., "Object-Centric learning with slot attention", in NeurIPS, 2020.

[0006] This disclosure provides a machine learning apparatus and method, as well as a robot control device, that can facilitate machine learning for image recognition of target objects.

[0007] In this disclosure, the machine learning apparatus comprises a communication unit that is communicably connected to an external device including an imaging device that captures an image of a target object, and a control unit that controls the external device via the communication unit. The target object includes a first part having physical properties that cause a change in appearance separate from a change in posture, and a second part that is less susceptible to changes in appearance due to the above-mentioned physical properties than the first part. The control unit, via the communication unit, causes the imaging device to capture a first image including the target object in a first imaging posture as viewed from the imaging device, and via the communication unit, causes the imaging device to capture a second image in a second imaging posture different from the first imaging posture, in which the first part of the target object has undergone a change in appearance due to the above-mentioned physical properties from the first image, and generates at least one of a training dataset and a trained model in machine learning in which the second part of the target object is image recognized, based on the first and second image and the first and second imaging postures.

[0008] In this disclosure, the robot control device comprises a communication unit that is communicatively connected to a manipulator on which an imaging device for capturing images of a target object is located, and a control unit that controls the manipulator via the communication unit. The target object includes a first part having physical properties that cause changes in appearance separate from changes in posture, and a second part that is less susceptible to changes in appearance due to the above-mentioned physical properties than the first part. The control unit operates the manipulator via the communication unit based on the image of the target object captured by the imaging device and a trained model in which machine learning is performed to image recognize the second part of the target object.

[0009] These general and specific embodiments may be implemented by systems, methods, and computer programs, or combinations thereof.

[0010] The machine learning apparatus and method, as well as the robot control device described herein, make it easier to perform machine learning for image recognition of target objects.

[0011] Figure illustrating the machine learning system in Embodiment 1 of this disclosure Figure illustrating the operating state of the robot in the machine learning system Block diagram illustrating the configuration of the machine learning system Flowchart illustrating the operation of the machine learning system Figure for explaining the operation of the machine learning system Functional block diagram illustrating the machine learning process in the machine learning system of Embodiment 1 Flowchart illustrating the operation of a trained robot control device by the machine learning system Flowchart illustrating the training data capture process in the machine learning system Figure illustrating the data structure of the background dataset in the machine learning system Figure illustrating the data structure of the training dataset in the machine learning system Flowchart illustrating the machine learning process in the machine learning system Functional block diagram illustrating the pose estimation model in the machine learning system of Embodiment 1 Figure for explaining the processing of the pose estimation model in the machine learning system of Embodiment 1 Functional block diagram illustrating the pose estimation model in the machine learning system of Embodiment 2 Functional block diagram showing an example of the configuration of the prompt generation unit in the machine learning system of Embodiment 2 Figure for explaining keypoint transfer in the pose estimation model of the machine learning system Functional block diagram illustrating the pose estimation model in the machine learning system of Embodiment 3 Block diagram illustrating a modified configuration of the machine learning system Flowchart illustrating the training data capture process in the machine learning system of the modified configuration of Figure 17

[0012] The embodiments will be described in detail below, with reference to the drawings as appropriate. However, unnecessarily detailed explanations may be omitted. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding for those skilled in the art.

[0013] The applicant provides the accompanying drawings and the following description so that a person skilled in the art can fully understand the disclosure, and not intends to limit the subject matter described in the claims.

[0014] (Embodiment 1) Hereinafter, Embodiment 1 of the present disclosure will be described with reference to the drawings.

[0015] 1. The machine learning apparatus and method according to Embodiment 1, as well as the configuration of the system using the robot control device, will be described below.

[0016] 1.1. System Overview The overview of the machine learning system according to this embodiment will be explained using Figures 1 and 2.

[0017] Figure 1 illustrates a machine learning system 1 in this embodiment. This system 1 comprises, for example, a robot 20, a machine learning device 5, and a display device 10, as shown in Figure 1. This system 1 is applied to machine learning for image recognition, for example, to enable the robot 20 to recognize a target object 11. The target object 11 is an object to be recognized by image recognition and can be appropriately set to various objects.

[0018] In this system 1, the target object 11 includes a part 12 whose appearance is relatively easy to change other than by changes in posture (hereinafter referred to as the "unsteady part 12") and a part 13 whose appearance is difficult to change (hereinafter referred to as the "steady part 13"). The unsteady part 12 and the steady part 13 are examples of the first and second parts in this embodiment, respectively.

[0019] The transient portion 12 of the object 11 has physical properties that cause changes in appearance due to, for example, deformation or the influence of texture such as transparency or metallic luster. The physical properties of the transient portion 12 include, for example, flexibility such as plasticity, light transmittance such as transparency, or specular reflectivity such as metallic luster. The steady portion 13 changes in appearance with changes in posture, but such changes in appearance are less likely to occur than in the transient portion 12. The steady portion 13 has, for example, higher rigidity or light diffusion than the transient portion 12, and the physical properties that cause the aforementioned changes in appearance are lower than those of the transient portion 12.

[0020] For example, if the object 11 is a pouch product, the deformable pouch container is the transient part 12, and the cap is the transient part 13. If the object 11 is a PET bottle beverage, the transparent bottle is the transient part 12, and the cap or label is the transient part 13. If the object 11 is a device with a cable, the cable is the transient part 12, and the housing is the transient part 13. The object 11 may also be various products in which the parts to be designated as transient parts 12 are covered with packaging such as metallic luster or plasticity.

[0021] In such target objects 11, it is conceivable that image recognition, such as pose estimation of the target object 11, may become difficult to acquire through machine learning due to the changes in appearance of the non-stationary parts 12 as described above. Therefore, in this embodiment, a machine learning system 1 is provided that can easily acquire image recognition of the target object 11 through machine learning, even if the target object 11 has non-stationary parts 12 that cause changes in appearance other than changes in posture, as described above.

[0022] Figure 2 illustrates the operational state of the robot 20 in this system 1. The robot 20 in this system 1 is, for example, an industrial robot and comprises a camera 2, a robot arm 21, and a robot control device 4. As shown in Figure 2, for example, the robot 20 in this embodiment is operated in a site 15 such as a factory or logistics warehouse to perform image recognition of a target object 11 from an image captured by the camera 2 for grasping by the end effector 6 of the robot arm 21.

[0023] The robot 20 of this system 1 can operate by, for example, gripping a steady portion 13 of a target object 11. Furthermore, in the actual site 15 where such a robot 20 is operated, it is assumed that various backgrounds other than the target object 11 may be captured in the image captured by the camera 2. Such backgrounds include, for example, a conveyor belt, a container box, a transport container, etc. In this embodiment, a machine learning system 1 is provided that makes it easier to perform machine learning that can recognize the target object 11 even when the background in the captured image is complex. The details of the configuration of this system 1 will be described below.

[0024] 1.2. Details of Configuration FIG. 3 is a block diagram illustrating the configuration of the machine learning system 1 according to the present embodiment. In the present system 1, for example, the display device 10 and the robot 20 are communicatively connected to the machine learning device 5.

[0025] The display device 10 of the present system 1 is constituted by various display devices such as an organic light emitting diode display (OLED), a micro LED display, a mini LED display, or a liquid crystal display (LCD), and has a display surface for displaying various images. The display device 10 includes a plurality of pixels arranged in a two-dimensional array to form the display surface, a communication circuit that communicatively connects to an external device to receive image data, and a control circuit that causes the received image data to be displayed on the display surface. The display device 10 is not necessarily limited to a display device having a display surface, and may be, for example, a projector. Alternatively, the display device 10 may be a device that switches display by arranging a substance such as paper or cloth on the display surface.

[0026] 1.2.1. Regarding the Robot In the robot 20 of the present system 1, as shown in FIG. 3, for example, the robot arm 21 and the camera 2 are each communicatively connected to the robot control device 4.

[0027] The robot arm 21 is, for example, an articulated manipulator having a plurality of movable axes. For example, the robot arm 21 includes a drive unit such as a motor or an actuator for each joint, a drive mechanism that constitutes the joint, an encoder that detects the drive state of the drive mechanism, and an end effector 6 for operations such as grasping. The robot arm 21 also includes a communication circuit that receives control commands from the robot control device 4, and a control circuit that controls the drive unit of each joint and drives the end effector 6 based on the control commands.

[0028] The camera 2 is attached to the robot arm 21 of the robot 20, for example. For example, the camera 2 is arranged near the end effector 6 of the robot arm 21 so that the target object 11 to be grasped is included within the angle of view range.

[0029] The camera 2 includes, for example, an optical system including a lens that forms a subject image, an image sensor that captures the subject image formed via the optical system to generate image data, and a communication circuit that transmits the image data to an external device such as the robot control device 4. The image sensor is, for example, a CMOS image sensor or a CCD image sensor, and captures a color image or a monochrome image. The camera 2 may be a depth camera and may include a distance image sensor.

[0030] The camera 2 is an example of an imaging device in the machine learning system 1. The imaging device in the present system 1 is not limited to one camera 2, and may include a plurality of cameras 2. The camera 2 of the present system 1 does not particularly need to be provided in the robot 20, and may be an external component of the robot 20.

[0031] 1.2.2. Robot Control Device The robot control device 4 is an arithmetic device that controls each part of the robot 20 such as the robot arm 21. As shown in, for example, FIG. 3, the robot control device 4 includes a control unit 40, a storage unit 41, and a communication interface 42. Hereinafter, the interface is abbreviated as "I / F". The robot control device 4 is configured by an information processing device such as various computers.

[0032] The control unit 40 controls, for example, the overall operation of the robot control device 4. The control unit 40 includes, for example, a CPU, and realizes predetermined functions in the robot control device 4 by executing a program (software) with the CPU. For example, the control unit 40 has a self-localization estimation function that estimates the posture of the camera 2 and the like in the robot 20 based on information on various driving states received from the robot arm 21. The self-localization estimation function may be realized by arithmetic processing for solving forward kinematics using encoder information representing angles of each joint of the robot arm 21, for example, or Visual-SLAM or the like using information on imaging results obtained by the camera 2 may be used.

[0033] The control unit 40 may include a processor consisting of dedicated electronic circuits designed to perform predetermined functions, either in place of or in addition to the CPU. That is, the control unit 40 can be implemented with various processors such as a CPU, MPU, GPU, DSP, FPGA, ASIC, etc. The control unit 40 may consist of one or more processors.

[0034] The memory unit 41 is a storage medium that stores programs and data necessary to realize the functions of the robot control device 4, and is composed of, for example, an HDD or SSD. The above programs may be provided from a communication network such as the Internet, or they may be stored on a portable recording medium. The memory unit 41 stores, for example, a trained posture estimation model 3 (details will be described later) when the robot 20 is in operation.

[0035] Furthermore, the storage unit 41 may include RAM such as DRAM or SRAM, and may function as a buffer memory for temporarily storing (i.e., holding) data. The storage unit 41 may also function as a working area of ​​the control unit 40, and may include a storage area in the internal memory of the control unit 40.

[0036] The communication interface 42 is a circuit for connecting the robot control device 4 to external devices such as a camera 2, a robot arm 21, and a machine learning device 5, and for performing data communication according to a predetermined communication standard. For example, predetermined communication standards include USB, IEEE 1395, IEEE 802.3, IEEE 802.11a / 11b / 11g / 11ac, Wi-Fi, and Bluetooth®. The communication interface 42 may be configured to include connection terminals for connecting to external devices, and may communicate with external devices via a communication network or directly.

[0037] 1.2.3. About the Machine Learning Device The machine learning device 5 is an information processing device that performs information processing for machine learning in this system 1, and generates, for example, at least one of a training dataset and a trained model. The machine learning device 5 comprises, for example, a control unit 50, a storage unit 51, a communication interface 52, an operation unit 53, and a display unit 54, as shown in Figure 3. The machine learning device 5 is composed of, for example, an information processing device such as a personal computer (PC).

[0038] The control unit 50 includes, for example, a CPU that works in cooperation with software to realize predetermined functions. The control unit 50 controls the overall operation of the machine learning device 5, for example. The control unit 50 reads data and programs stored in the memory unit 51 and performs various calculations to realize various functions. The control unit 50 includes, for example, a learning unit 50a that performs calculations for machine learning of the posture estimation model 3 as a functional configuration.

[0039] The control unit 50 executes a program that includes a set of instructions for realizing each of the above functions. This program may be provided via a communication network or stored on a portable recording medium. The control unit 50 may also be a processor composed of dedicated electronic circuits or reconfigurable electronic circuits designed to realize each of the above functions. The control unit 50 may be implemented using various processors such as a CPU, MPU, GPU, GPGPU, NPU, TPU, microcontroller, DSP, FPGA, and ASIC. The control unit 50 may consist of one or more processors.

[0040] The memory unit 51 is a storage medium that stores programs and data necessary to realize the functions of the machine learning device 5. The memory unit 51 includes, for example, an SSD or HDD, and stores parameters, data, and control programs for realizing predetermined functions such as machine learning. For example, the memory unit 51 stores the background dataset DS1, which will be described later.

[0041] Furthermore, the storage unit 51 may include RAM such as DRAM or SRAM, and may function as a buffer memory for holding data. The storage unit 51 may also function as a working area for the control unit 50, and may include a storage area in the internal memory of the control unit 50.

[0042] The communication interface 52 is a circuit for connecting the machine learning device 5 to external devices such as the display device 10 and the robot control device 4, and for performing data communication according to a predetermined communication standard. For example, predetermined communication standards include USB, HDMI®, IEEE 1395, IEEE 802.3, IEEE 802.11a / 11b / 11g / 11ac, Wi-Fi, and Bluetooth. The communication interface 52 may be configured to include connection terminals for connecting to external devices, and may communicate with external devices via a communication network or directly.

[0043] The operation unit 53 is a general term for the operating components that the user operates. The operation unit 53 is composed of, for example, a keyboard, mouse, trackpad, touchpad, buttons, and switches, or a combination thereof. The operation unit 53 may also constitute a touch panel together with the display unit 54. The operation unit 53 acquires various information input by the user's operation. In the machine learning device 5, the operation unit 53 may be omitted. For example, the machine learning device 5 may include connection terminals for connecting external operating components.

[0044] The display unit 54 is composed of, for example, an OLED or LCD, and displays various information. The display unit 54 may display various icons for operating the operation unit 53, and various information such as information input from the operation unit 53. In the machine learning device 5, the display unit 54 may be omitted.

[0045] The configuration of System 1 described above is merely an example, and the configuration of System 1 is not limited to this. For example, the machine learning device 5 may consist of one or more server devices connected via a communication network, or it may consist of a collaboration between a server device and an information terminal. The machine learning device 5 may also be configured as a network-type system. Furthermore, some or all of the functions of the machine learning device 5 may be implemented in the robot control device 4, and the machine learning device 5 and the robot control device 4 may be configured as an integrated unit.

[0046] 2. Operation The operation of the machine learning system 1 of this embodiment, which is configured as described above, will be explained below.

[0047] The machine learning system 1 of this embodiment performs machine learning on a posture estimation model 3, which is an image recognition model that estimates the posture of an object 11 based on images captured by the camera 2 in the robot 20. The posture of the object 11 includes, for example, the position and orientation (angle) of the object 11 in three-dimensional space.

[0048] In machine learning for estimating the pose of such object 11, a typical conventional technique involves manually annotating the regions in captured images where the object 11 is present to prepare training data. For example, Non-Patent Document 1 uses manual prompts such as placing points or boxes on the target of segmentation in the image, or uses a segmentation dataset containing more than one billion masks to build the model.

[0049] However, with these conventional techniques, the annotation burden arises from sequentially identifying the target object 11 in a large number of captured images and distinguishing the corresponding region from the background, resulting in high annotation costs in machine learning for image recognition, such as pose estimation of the target object 11.

[0050] Therefore, the machine learning system 1 of this embodiment provides a machine learning method that enables pose estimation of the target object 11 even when the background of the site 15 where the robot 20 is operated is complex, by using machine learning that does not use annotation to sequentially identify the target object 11 in the captured image by a human.

[0051] 2.1. Overall Operation The overall operation of System 1 will be explained using Figures 4 to 6B. Figure 4 is a flowchart illustrating the operation of System 1. Figure 5 is a diagram illustrating the operation of System 1. Figure 6A is a functional block diagram illustrating the machine learning process (S2) in System 1. Figure 6B is a flowchart illustrating the operation of the trained robot control device 4 by System 1.

[0052] First, in this system 1, for example, the control unit 50 of the machine learning device 5 controls image capture using the camera 2 and the display device 10 to collect data to be used for machine learning of the posture estimation model 3, i.e., training data (S1). The training data capture process (S1) is performed in this system 1 by placing the target object 11 on the display surface of the display device 10, as shown in Figure 1, for example.

[0053] In the learning data capture process (S1) of this embodiment, the background image, which is an image showing the background to be displayed on the display device 10, and the orientation of the camera 2 are changed, for example, randomly, while images of the target object 11 including the background image are captured multiple times. The orientation of the camera 2 during such image capture is an example of the imaging orientation in this embodiment. Multiple captured images are illustrated in Figures 5(A) and (B).

[0054] Figure 5(A) shows an example of a first captured image A1 obtained by camera 2 in a first imaging posture C1 of the target object 11 in this system 1. Figure 5(B) illustrates a second captured image A2 obtained by camera 2 in a second imaging posture C2, which is different from the first imaging posture C1 in Figure 5(A), of the target object 11. In step S1, each captured image A1 to A2 is captured, for example, with the target object 11 stationary on the display surface, and the camera 2's posture is set to the respective imaging postures C1 to C2.

[0055] The first captured image A1 includes, for example, as shown in FIG. 5(A), the target object 11 viewed from a first capturing posture C1 and a background image B1. The second captured image A2 includes, for example, as shown in FIG. 5(B), the target object 11 viewed from a second capturing posture C2 different from the first capturing posture C1 and a background image B2 different from the background image B1 of the first captured image A1. In the second captured image A2 illustrated in FIG. 5(B), the appearance of the unsteady portion 12 is made different from that in the first captured image A1 due to the deformation of the unsteady portion 12 in addition to the posture change from the first captured image A1.

[0056] FIG. 5(C) is a diagram for explaining the correspondence between the posture of the target object 11 in each of the captured images A1 and A2 of FIGS. 5(A) and (B) in the present system 1 and each of the capturing postures C1 and C2. According to the learning data photographing process (S1) of the present embodiment, as shown in FIG. 5(C), a consistent relationship is established between the change in the posture of the steady portion 13 of the target object 11 between the first and second captured images A1 and A2 and the change between the first and second capturing postures C1 and C2. Such a relationship is expressed as the following relational expression (1) using, for example, a homogeneous transformation matrix. [Math 1] c1 T c2 = c1 T o ・( c2 T o ) -1 …(1)

[0057] In the above formula (1), the homogeneous transformation matrix c1 T c2 represents a change from the first capturing posture C1 to the second capturing posture C2. The homogeneous transformation matrix c1 T o represents the posture of (the steady portion 13 of) the target object 11 viewed from the first capturing posture C1, that is, the first object posture. The homogeneous transformation matrix c2 T o represents the posture of (the steady portion 13 of) the target object 11 viewed from the second capturing posture C2, that is, the second object posture. The various homogeneous transformation matrices described above are described as a 4-row 4-column matrix using a 3-row 3-column rotation matrix and a 3-dimensional translation vector for a three-dimensional space, for example.

[0058] According to the training data acquisition process (S1) of this embodiment, by changing background images B1 and B2, etc., training data can be generated such that the above relational expression (1) does not hold in the transient portion 12 of the target object 11 other than the steady portion 13 and in the background portion in each captured image A1 and A2. Details of the process in step S1 will be described later.

[0059] Next, for example, the control unit 50 of the machine learning device 5 performs machine learning on the posture estimation model 3 based on the training data generated by the training data acquisition process (S1) (S2). An overview of the machine learning process (S2) in this system 1 will be explained using Figure 6A.

[0060] In the machine learning process (S2) of this embodiment, first, as shown in Figure 6A, the first and second captured images A1 and A2 are input to the pose estimation model 3 being trained. Based on the first and second object poses Tc1 and Tc2 output from the pose estimation model 3 as a result, and the first and second imaging poses C1 and C2 at the time of capturing each captured image A1 and A2 (S1), the learning unit 50a of the control unit 50, for example, performs machine learning to satisfy the above equation (1) (S2). Details of the process in step S2 will be described later.

[0061] Next, for example, the machine learning device 5 applies the learning results to the robot control device 4 (S3) so that the robot 20 is operated using the posture estimation model 3 generated as a result of the machine learning process (S2). For example, in step S3, the control unit 50 of the machine learning device 5 transmits the data of the learned posture estimation model 3 to the robot control device 4 via the communication I / F 52.

[0062] For example, in step S3, the control unit 40 of the robot control device 4 receives the posture estimation model 3 via the communication I / F 42, stores it in the memory unit 41, and performs image recognition processing using the posture estimation model 3 when the robot 20 is in operation. Figure 6B illustrates the operation of the robot control device 4 after it has been trained by the machine learning system 1.

[0063] The process shown in the flow chart of Figure 6B is repeatedly executed by the control unit 40 of the robot control device 4 at a predetermined period, such as the frame period of the camera 2, when the trained posture estimation model 3 is stored in the memory unit 41 during the operation of the robot 20. First, the control unit 40 sequentially receives captured images from the camera 2 attached to the robot arm 21, for example (S31). The control unit 40 inputs the received captured images into the trained posture estimation model 3 and obtains the object posture estimated by the posture estimation model 3 (S32). The control unit 40 operates the robot arm 21 by controlling the various drive units of the robot arm 21 to grasp the steady portion 13 of the target object 11 based on the estimated object posture (S33).

[0064] System 1 completes the process illustrated in the flow chart of Figure 4, for example, by applying a trained posture estimation model 3 to the robot control device 4 (S3).

[0065] Based on the above operation, even if there are transient parts 12 in the target object 11 that cause changes in appearance other than changes in posture, System 1 can construct a posture estimation model 3 using machine learning to estimate the posture of the target object 11 by focusing on the steady parts 13 that are less likely to cause such changes in appearance. Furthermore, System 1 can obtain a posture estimation model 3 that accurately estimates the posture of the target object 11, robustly against complex backgrounds, using machine learning, even without annotation.

[0066] 2.2. Training Data Acquisition Processing The details of the training data acquisition processing in step S1 of Figure 4 will be explained using Figures 7 to 9.

[0067] Figure 7 is a flowchart illustrating the training data acquisition process (S1) in this system 1. Each process shown in the flowchart of Figure 7 is executed, for example, by the control unit 50 of the machine learning device 5.

[0068] In this system 1, first, for example, the control unit 50 of the machine learning device 5 prepares a background dataset DS1 (S10). The background dataset DS1 is a dataset that manages the image data of each background image Bk (where k is an integer from 1 to K).

[0069] Figure 8 illustrates the data structure of the background dataset DS1 in this system 1. The background dataset DS1 illustrated in Figure 8 stores image data of multiple background images B1 to BK and background IDs indicating identification information for each background image B1 to BK, in relation to each other. For example, in step S10, the control unit 50 reads the background dataset DS1 that has been previously stored in the storage unit 51. Alternatively, the control unit 50 may acquire part or all of the background dataset DS1 from an external device or communication network of the machine learning device 5, for example, via the communication interface 42 (S10).

[0070] The background dataset DS1 may consist of a dataset of images provided as a preset, or it may consist of images prepared by the user of this system 1. For example, the control unit 50 of the machine learning device 5 may register background images taken by the user of various backgrounds at the site 15 into the background dataset DS1. During such registration, the background ID may be automatically assigned by the control unit 50, or it may be assigned by the user.

[0071] The background images B1 to BK in the background dataset DS1 include images of backgrounds with complexity, such as complex structures or multiple objects overlapping in a complex manner (e.g., images of forests). The background dataset DS1 may include various background images B1 to BK, and may also include a background image Bk that is different from the actual background of the site 15. Furthermore, the background dataset DS1 may or may not include a background image Bk of the actual background of the site 15. In addition, the background images B1 to BK are not limited to real-world images, but may also include images generated appropriately through various image processing methods. The number K of background images B1 to BK in the background dataset DS1 is set appropriately, for example, from the perspective of background diversity in machine learning (e.g., a few images to tens of thousands of images).

[0072] In the background dataset DS1, background images B1 to BK with different background IDs are set to be dissimilar images such that, for example, the relationship in equation (1) above does not hold for two images. For example, equation (1) above may hold between multiple similar images obtained by different projection transformations for a 3D object with the same background. By using such dissimilar images in the background dataset DS1, the accuracy of machine learning using the training data D2 obtained in the training data acquisition process (S1) can be improved.

[0073] Furthermore, if objects other than the target object 11 and background images B1 and B2, such as a part of the housing of the display device 10, are captured in the first and second captured images A1 and A2, the above equation (1) will also hold true for the captured portion, raising concerns about its impact on the machine learning process. Therefore, the system 1 may, for example, set the imaging postures C1 and C2 such that in at least one of the two images, the area in which the background images B1 and B2 are displayed is wider than the imaging range.

[0074] Returning to Figure 7, the control unit 50 selects background image B1 from the background dataset DS1 and displays the selected background image B1 on the display device 10 (S11). The selection of background image B1 (S11) may be in a predetermined order, such as the order of background IDs in the background dataset DS1, or it may be in a random order.

[0075] For example, in step S11, the control unit 50 transmits the image data of the selected background image B1 to the display device 10 via the communication I / F 52. When the display device 10 receives the image data from, for example, the machine learning device 5, it displays the background image B1 indicated by the received image data on its display surface. For example, the control unit 50 manages the background ID of the background image B1 displayed in step S11. The selection of background image B1 (S11) is performed over multiple (N) steps S11 to S14 so that background images B1 to BK with various background IDs can be selected, and the background IDs can overlap.

[0076] Furthermore, for example, in step S11, the appearance of the transient portion 12 of the target object 11 may be changed. Such changes in the appearance of the transient portion 12 in this system 1 are performed in such a way that equation (1) above does not hold true for the transient portion 12 before and after the change. For example, if the transient portion 12 of the target object 11 is flexible, it may be deformed by the user of this system 1. Alternatively, if the transient portion 12 is transparent or mirror-like, the appearance of the transient portion 12 may be changed by changing the optical image transmitted or reflected by the transient portion 12 due to a change in the background image B1. In this case, changes in the appearance of the steady portion 13 other than changes in posture are suppressed so that equation (1) above holds true for the steady portion 13.

[0077] Furthermore, the control unit 50, for example, performs drive control of the robot 20 to dynamically set the imaging posture of the camera 2 (S12). The process in step S12 is performed to capture an image An (where n is an integer from 1 to N) for various background images B1 to BK from various imaging postures. The drive control in multiple steps S12 may change the imaging posture of the camera 2 continuously or discretely.

[0078] For example, in step S12, the control unit 50 transmits an instruction to the robot control device 4 via the communication interface 52 to drive the robot arm 21. In the robot control device 4, for example, the control unit 40 receives the drive instruction via the communication interface 42 and generates a control command for the robot arm 21 according to the received drive instruction. The control unit 40 transmits the control command to the robot arm 21 via the communication interface 42 to drive the robot arm 21. The control command may include the drive amount for each part of the robot arm 21, or the drive amount may be calculated by the control circuit in the robot arm 21. The control unit 40 may receive feedback of various drive amounts from the robot arm 21.

[0079] In step S12, the imaging posture of camera 2 can be managed using the robot 20's self-position estimation function. For example, the control unit 40 of the robot control device 4 sequentially calculates the imaging posture of camera 2 based on the amount of drive of the robot arm 21 and records it in the storage unit 41. The setting of the imaging posture of camera 2 by the drive control of robot 20 (S12) is performed, for example, within a predetermined range in which the target object 11 is expected to be captured in the image of camera 2, based on the field of view range of camera 2. The processing order of steps S11 and S12 is not particularly limited and may be in the same order as shown, in reverse order, or simultaneously.

[0080] Next, the control unit 50 acquires image data (S13) showing the captured image A1, which is the result of the shooting in the state where the background image B1 is displayed on the display device 10 in steps S11 and S12 and the imaging posture of the camera 2 is set by the robot 20. For example, the camera 2 on the robot 20 periodically captures images at a predetermined frame rate to generate image data, and sequentially transmits the generated image data to the robot control device 4.

[0081] For example, in step S13, the control unit 50 of the machine learning device 5 requests the robot control device 4 via the communication interface 52 for captured image A1 of the background image B1 currently being displayed. The control unit 40 of the robot control device 4, for example, in response to the request from the machine learning device 5, identifies the image data captured at the time of the request from the image data received from the camera 2 via the communication interface 42. The control unit 40 transmits the identified image data and the corresponding imaging posture to the machine learning device 5 via the communication interface 42. In this way, the control unit 50 of the machine learning device 5 acquires the image data and imaging posture of the requested captured image A1 as the shooting result via the communication interface 52 (S13).

[0082] Furthermore, the control unit 50 determines, for example, whether the number of captured images A1 obtained by image capture in each step S11 to S13 as described above has reached a predetermined number N (S14). The number N is set appropriately from the viewpoint of collecting sufficient training data for machine learning of the posture estimation model 3 in this system 1 (for example, a few images to tens of thousands of images). The number N may be greater than or equal to the number K of background images B1 to BK in the background dataset DS1. For example, even if the respective numbers N and K are only a few images, machine learning processing (S2) can be performed using Bayesian optimization.

[0083] If, for example, the number of captured images A1 has not reached N (NO in S14), the control unit 50 repeats the processing from step S11 onwards. By repeating the processing in steps S11 to S13, N captured images A1 to AN are taken for various background images B1 to BK in various imaging orientations of the camera 2 (S11 to S14).

[0084] For example, when the number of captured images A1 to AN reaches N (YES in S14), the control unit 50 generates a training dataset containing multiple training data based on the image data and imaging pose of the acquired captured images A1 to AN (S15).

[0085] Figure 9 illustrates the data structure of the training dataset DS2 in this system 1. The training dataset DS2 illustrated in Figure 9 manages each training data D2 with identification numbers 1 to N by associating it with the captured images A1 to AN, the capturing poses C1 to CN, and the background ID. This data structure of the training dataset DS2 defines the processing for each pair of training data D2 (i.e., training data pair DP2a) sampled in the machine learning process (S2), as will be described later. In this way, the training dataset DS2 constitutes a data structure that conforms to the program used for machine learning.

[0086] For example, for the first image capture (S11 to S13), the control unit 50 associates the background ID (BG001) of the background image B1 selected in step S11 with the image data of the captured image A1 and the capturing orientation C1 acquired in step S13 to construct the corresponding learning data D2. For each image capture (S11 to S13), the control unit 50 constructs the learning data D2 in the same manner as above and stores them sequentially in, for example, the storage unit 51. Based on the learning data D2 thus accumulated, the control unit 50 generates a learning dataset DS2 containing each piece of learning data D2 (S14).

[0087] The control unit 50 terminates the training data capture process (S1) illustrated in Figure 7 by generating the training dataset DS2 (S15), and proceeds to the machine learning process (step S2 in Figure 4).

[0088] According to the above learning data acquisition process (S1), the system 1 can acquire images of the target object 11 from various imaging poses in various background images B1 to BK while changing the appearance of the non-stationary portion 12 of the target object 11 (S11 to S13), and acquire learning data D2 of each acquired image A1 to AN as a result of these acquisitions. Furthermore, according to the system 1, the learning data D2 can be acquired by utilizing the functions of the robot 20, and the processing load for generating the learning dataset DS2 can be reduced.

[0089] The selection of background images B1 to BK in step S11 described above may be restricted so that background IDs do not overlap between multiple steps S11. For example, the number of captured images A1 to AN N in the training dataset DS2 may be set to be less than or equal to the number of background images B1 to BK in the background dataset DS1, K. In addition, in the training dataset DS2 of this system 1, the management of background IDs may be omitted as appropriate if they are not used in machine learning processing (S2), etc.

[0090] Furthermore, in the above description, an example was described in which the display device 10 is connected to the machine learning device 5 via communication and the display of background images B1 to BK (S11) is controlled by the control unit 50 in this system 1. However, the system 1 is not limited to this, and for example, the display device 10 may be connected to the robot control device 4 via communication, and the display control in step S11 may be performed by the control unit 40 of the robot control device 4. For example, in the robot control device 4, the communication I / F 52 may support video communication standards such as HDMI, and the background data set DS1 may be pre-stored in the storage unit 41.

[0091] Furthermore, each process in the learning data capture process (S1) is not limited to the control unit 50 of the machine learning device 5, but may also be executed by the control unit 40 of the robot control device 4. In this embodiment, the robot control device 4 may be an example of a machine learning device that generates the learning dataset DS2. In this system 1, the robot control device 4 may transmit the generated learning dataset DS2 to the machine learning device 5 via the communication I / F 42.

[0092] Furthermore, the above description described an example in which the imaging posture as seen from the camera 2 of the target object 11 is set (S12) by changing the posture of the camera 2 attached to the robot arm 21 while the target object 11 is stationary on the display device 10. The System 1 is not limited to this, however, for example, the imaging posture as seen from the camera 2 of the target object 11 may be set (S12) by moving the target object 11 without changing the posture of the camera 2, and an image A1 of the background image B1 and the target object 11 may be captured from the camera 2 (S13). In this case, the posture of the target object 11 at each imaging time may be appropriately managed as the imaging posture, and the training dataset DS2 may be generated (S15) in the same manner as in the above embodiment. That is, the imaging posture managed by the System 1 is not limited to the posture of the camera 2, but may also be the posture of the target object 11.

[0093] Furthermore, the system 1 may automatically perform changes in appearance (S11), such as deformation of the non-steady portion 12. For example, the various control units 50 and 40 of the system 1 may deform the non-steady portion 12 using actuators such as the robot arm 21. Also, the change in imaging posture in step S12 of the system 1 may be changed by various means, not limited to the robot arm 21. For example, by changing the imaging posture of the target object 11 using an electromagnet or a thin wire that is difficult to see with the camera 2, it is possible to change the imaging posture of the target object 11 without actuators such as the robot arm 21 appearing in the captured images A1 and A2. Along with this, changes in appearance (S11), such as deformation of the non-steady portion 12 of the target object 11, may be performed.

[0094] Furthermore, in step S13, the imaging operation by camera 2 does not have to be performed periodically; for example, it may be performed in response to a request from the machine learning device 5 to acquire an image A1. The control unit 40 of the robot control device 4 may send a control signal to camera 2 via the communication I / F 42 to execute the imaging operation of camera 2 in response to such a request. Also, the image data of these images A1 to AN does not have to be sent to the machine learning device 5 one by one after each image capture (S13); the image data of multiple images may be sent together.

[0095] Furthermore, although the above embodiment describes an example in which the imaging posture (S12) is set by changing the posture of the camera 2 mounted on the robot arm 21, the System 1 is not limited to this. For example, the System 1 may have multiple cameras 2 installed in advance and imaging may be performed from each camera 2, thereby setting different imaging postures. For example, the imaging posture of each camera 2 may be measured in advance using an AR marker or the like, or the mounting position of the camera may be mechanically restricted by a jig to achieve the desired imaging posture. In addition, the camera 2 used in such a machine learning system 1 does not necessarily have to be attached to the robot arm 21, and may be appropriately positioned separately from the robot arm 21, etc., so as to be able to image the target object 11.

[0096] 2.3. The details of the training data acquisition process in step S2 of Figure 4 of the machine learning process will be explained using Figures 10 to 12.

[0097] Figure 10 is a flowchart illustrating the machine learning process (S2) in this system 1. Each process shown in the flowchart of Figure 10 is executed by the control unit 50 of the machine learning device 5, for example, after the training dataset DS2 has been prepared in step S1.

[0098] First, the control unit 50 of the machine learning device 5 selects a training data pair DP2a from the training dataset DS2 stored in the memory unit 51 (S21). Below, an example of processing will be described when a training data pair DP2a consisting of training data D2 of the first captured image A1 and training data D2 of the second captured image A2 is selected in step S21. The first and second captured images A1 and A2 in the training data pair DP2a in step S21 are selected as inputs for calculation processing (S22 to S24) to find the error from the above-mentioned relational expression (1) in the pose estimation model 3 (Figure 11) under training.

[0099] For example, in step S21, the control unit 50 selects a pair of training data DP2a by referencing the background ID in the training dataset DS2 and sampling two training data D2 with different background IDs. This makes it easy to exclude combinations of captured images with the same background image in the machine learning process (S2), even if the training data acquisition process (S1) allows for overlapping background images B1 to BK, and makes it easier to have the pose estimation model 3 in equation (1) ignore the background during training.

[0100] Next, the control unit 50 inputs the image data of the first and second captured images A1 and A2 from the selected training data pair DP2a to the training pose estimation model 3, and obtains the first and second object poses Tc1 and Tc2 as training outputs from the pose estimation model 3, respectively (S22, S23). The processing in steps S22 and S23 will be explained with reference to Figure 11.

[0101] Figure 11 is a functional block diagram illustrating the attitude estimation model 3 in the machine learning system 1 of this embodiment. The attitude estimation model 3 of this embodiment includes, for example, a segmentation unit 31 and an attitude estimation unit 32, as shown in Figure 11.

[0102] The segmentation unit 31 is provided to perform image recognition processing (i.e., image segmentation) that generates a mask image in the input image that shows a region in the input image where a specific object, such as the steady portion 13 of the target object 11, is visible, in response to the input image, such as an captured image. The segmentation unit 31 can be configured by fine-tuning an existing trained model such as SAM or YOLO. Alternatively, the segmentation unit 31 may be configured by training an image recognition model, such as a segmentation model implemented with a neural network, from an initial state.

[0103] The pose estimation unit 32 includes, for example, a cropping unit 32a for a pre-prepared 3D model 30 and a cropping unit 32b for a mask image from the segmentation unit 31 and an input captured image. The position estimation unit 32 is configured to perform image recognition processing to estimate the object pose of the target object 11 shown by the 3D model 30 using, for example, each cropping unit 32a, 32b. The pose estimation unit 32 may be configured using existing methods such as FoundationPose, SAM-6D, DenseFusion, or Iterative Closest Point (ICP), or it may be configured by learning an image recognition model implemented with a neural network from an initial state. B. Wen, et. al. "FoundationPose: Unified 6D pose estimation and tracking of novel objects," in CVPR, 2024. J. Lin, et. al. "SAM-6D: Segment anything model meets zero-shot 6d object pose estimation," in CVPR, 2024. W. Chen, et. al. "DenseFusion: 6D object pose estimation by iterative dense fusion." in CVPR, 2019.

[0104] The 3D model 30 is model data representing the three-dimensional shape of the target object 11, and can be created, for example, using a known 3D scanner. Alternatively, the 3D model 30 may be generated from image data of the target object 11, and 3D model generation technologies such as NeRF, BundleSDF, and 3D Gaussian Splatting (see below) may be used. W. Bowen, et. al. "BundleSDF: Neural 6-DOF tracking and 3D reconstruction of unknown objects." in CVPR, 2023. K. Bernhard, et. al. "3D gaussian splatting for real-time radiance field rendering." ACM Trans. Graph. 42.4 (2023): 139-1.

[0105] The 3D model 30 is an example of shape data in this embodiment. The shape data in this embodiment is not necessarily limited to the 3D model 30, but may be, for example, a collection of two-dimensional images of the target object 11.

[0106] Figure 12 is a diagram illustrating the processing of the attitude estimation model 3 in this system 1. Figure 12(A) shows an example of an image captured and input to the attitude estimation model 3. Figure 12(B) shows an example of a cropped image from the image captured in Figure 12(A). Figure 12(C) shows an example of a 3D model 30 of the target object 11 in this system 1. Figure 12(D) shows an example of a rendered image from the 3D model 30 in Figure 12(C). Figure 12(E) shows an example of the estimated object attitude from Figures 12(B) and (D) in this system 1.

[0107] The attitude estimation model 3 of this system 1 first acquires a mask 13a on the captured image that indicates the region of the steady portion 13 of the target object 11, for example, using the segmentation unit 31, as illustrated in Figure 12(A). The cropping unit 32b of the position estimation unit 32 performs cropping on the captured image (Figure 12(A)) based on the information of the mask 13a, for example, to generate a cropped image as illustrated in Figure 12(B).

[0108] The cropped image of the crop unit 32b (Figure 12(B)) is, for example, an image extracted from the captured image (Figure 12(A)) that contains a rectangular or other region of interest that includes the steady portion 13. The crop unit 32b calculates the centroid of the region of interest as the reference position of the region of interest by, for example, extracting the contour of the mask 13a, and determines the size of the region of interest based on the camera parameters and depth of the camera 2.

[0109] Furthermore, in this embodiment, the crop portion 32a for the 3D model 30 is set with a parameter indicating a region of interest in the 3D model 30, i.e., an ROI parameter Pm, as shown in Figure 12(C). The ROI parameter Pm includes, for example, a three-dimensional coordinate position corresponding to the steady portion 13 in the 3D model 30, and a radius from this coordinate position. The ROI parameter Pm constitutes, for example, a learning parameter targeted for machine learning in this system 1.

[0110] In the posture estimation model 3 of this system 1, the cropping unit 32a performs rendering and cropping processing to image the 3D model 30 (Figure 12(C)) based on, for example, the ROI parameter Pm, and generates a rendered image (Figure 12(D)) in which the region of interest of the 3D model 30 has been cut out. The position estimation unit 32, for example, sets various postures and causes the cropping unit 32 to generate a predetermined number of rendered images (for example, 252 images) of the region of interest of the 3D model 30 viewed from each posture.

[0111] The position estimation unit 32 of this system 1 compares each rendering image (Figure 12(D)) based on these posture hypotheses with the cropped image (Figure 12(B)) to calculate the object posture of the target object 11 (the steady portion 13) as shown in Figure 12(E). The posture hypotheses of the position estimation unit 32 are repeatedly adjusted, for example, by iteratively performing the generation of a predetermined number of rendering images a predetermined number of times (for example, 5 times). The rendering and cropping processes by the cropping unit 32a for the 3D model 30 may be performed sequentially or integrally.

[0112] Returning to Figure 10, for example in step S22, the pose estimation model 3 of this embodiment first inputs the first captured image A1 to the segmentation unit 31, as illustrated in Figure 11, and outputs a first mask image. Next, the pose estimation unit 32 calculates the first object pose Tc1 as an estimation result in the current learning state based on the first mask image from the segmentation unit 31, the first captured image A1, and the 3D model 30.

[0113] Furthermore, in step S23, similar to step S22, for example in the posture estimation model 3 shown in Figure 11, the segmentation unit 31 generates a second mask image based on the second captured image A2. Next, the posture estimation unit 32 calculates the second object posture Tc2 based on the second mask image, the second captured image A2, and the 3D model 30. The processing in steps S22 and S23 is performed with the same learning parameters, etc., in the posture estimation model 3, as shown in Figure 11, for example.

[0114] Next, the control unit 50, for example as the learning unit 50a (Figure 6A), calculates an error function F(T) based on the calculated first and second object poses Tc1 and Tc2 and the first and second imaging poses C1 and C2 corresponding to the first and second imaging images A1 and A2 in the selected learning data pair DP2a (S24). The error function F(T) is a loss function that shows the error from the above-mentioned relational equation (1) in the estimation result during learning, and can be expressed as, for example, the following equation (2). [Equation 2] F(T) = |Vt(T)| + α × |Vr(T)| ... (2) T = ( c1 T o ) -1 ・ c1 T c2 ・ c2 T o

[0115] In equation (2) above, the homogeneous transformation matrix T has translational and rotational transformation components corresponding to the error between the left and right sides in relation (1). The translation vector Vt(T) is a three-dimensional vector representing the translational component of the homogeneous transformation matrix T for the error. The rotational vector Vr(T) is a three-dimensional vector representing the rotational component of the homogeneous transformation matrix T, and can be obtained from the rotation matrix, for example, using Rodrigues' rotation formula. The coefficient α is a parameter that adjusts the weighting of the error in the translational component and the error in the rotational component. For example, if the unit of the translational error is m and the unit of the rotational error is rad, α may be set to 0.1.

[0116] In the right-hand side of equation (2) above, the first term is the norm of the translation vector Vt(T), and represents the magnitude of the translation component in the homogeneous transformation matrix T (i.e., the translation amount of the error). The second term is the norm of the rotation vector Vr(T), and represents the magnitude of the rotation component in the homogeneous transformation matrix T (i.e., the rotation angle of the error). According to equation (2) above, the attitude estimation model 3 can be trained to estimate the object attitudes Tc1 and Tc2 in such a way that relation (1) is satisfied by minimizing the error function F(T).

[0117] Furthermore, the control unit 50, acting as the learning unit 50a, updates the learning parameters in the pose estimation model 3 based on the calculation result of the error function F(T) (S24), for example, as shown in Figure 6A (S25). The learning parameters include, for example, weighting parameters of the neural network and other components that constitute the pose estimation model 3 to be learned, such as the segmentation unit 31 and the pose estimation unit 32, as well as the ROI parameter Pm.

[0118] For updating the learning parameters (S25), optimization methods such as gradient descent can be applied, for example, if the model is differentiable. For example, in step S25, the control unit 50 updates the learning parameters to minimize the error function F(T) using backpropagation. Alternatively, the control unit 50 may perform reinforcement learning or Bayesian optimization to minimize the error function F(T). Existing learning algorithms such as PPO or TRPO can be used for reinforcement learning. For Bayesian optimization, parameter search methods such as Gaussian process regression or TPE may be used.

[0119] Furthermore, in step S25, the learning parameters of both the segmentation unit 31 and the attitude estimation unit 32 may be updated simultaneously, or they may be updated alternately. The system 1 may fix the learning parameters of one of the segmentation unit 31 and the attitude estimation unit 32 and update the learning parameters of the other (S25), or it may exclude one from the learning target and adopt an existing model, etc. Also, the updating of the learning parameters (S25) is not limited to each step S23, but may be performed, for example, at predetermined intervals.

[0120] Subsequently, the control unit 50 determines, for example, whether or not predetermined completion conditions have been met (S26). The completion conditions are the conditions for completing the machine learning of the posture estimation model 3, and are set appropriately according to the learning method adopted.

[0121] If the control unit 50 does not meet the conditions for completion of machine learning (NO in S26), it repeats the processing from step S21 onwards. In this way, by repeating steps S21 to S25, a training data pair DP2a containing two of various captured images A1 to AN is selected from the training dataset DS2 (S21), and the pose estimation model 3 is trained to minimize the error function F(T) calculated based on the corresponding captured poses C1 to CN.

[0122] If the control unit 50 satisfies the conditions for completion of machine learning (YES in S26), it generates the pose estimation model 3 of the learning result as a trained model by, for example, storing the final learning parameters in the storage unit 51 (S27).

[0123] The control unit 50 terminates the machine learning process (S2) exemplified in Figure 10 by generating the posture estimation model 3 from the learning results (S27), and proceeds to applying the learned posture estimation model 3 to the robot control device 4 (step S3 in Figure 4).

[0124] According to the above machine learning process (S2), in this system 1, the machine learning device 5 can perform machine learning based on the training data D2 using relational equation (1) (S21 to S25) to generate a pose estimation model 3 as a trained model (S27).

[0125] Generally, learning the segmentation unit requires mask data, which incurs high annotation costs. In contrast, this system 1 makes it possible to perform machine learning on the pose estimation model 3, including the segmentation unit 31, without annotation.

[0126] In the above machine learning process (S2), the processing by the learning unit 50a may be performed on an external server of the machine learning device 5. The pose estimation model 3 may also be executed on an external server of the machine learning device 5. The control unit 50 of the machine learning device 5 may transmit various information used for processing to the external server and receive the learned pose estimation model 3 from the external server.

[0127] The above explanation describes an example of using the error function F(T) of equation (2) in the machine learning process (S2), but System 1 is not limited to this. For example, System 1 may perform the machine learning process (S2) using an error function F(T1, T2) expressed as shown in the following equation (3) instead of the error function F(T) of equation (2). [Equation 3] F(T1, T2) = |Vt(T1) - Vt(T2)| + 2 × α × Φ(T1, T2) ... (3) T1 = w T c1 ・ c1 T o T2 = w T c2 ・ c2 T o Φ(T1, T2) = arccos(|q(T1)・q(T2)|)

[0128] In equation (3) above, the first term on the right-hand side is the Euclidean distance between the respective translation vectors Vt(T1) and Vt(T2) of the homogeneous transformation matrices T1 and T2. The function Φ(T1, T2) in the second term on the right-hand side is the inverse cosine of the norm of the inner product of the quaternion representations q(T1) and q(T2) of the respective rotation components of the homogeneous transformation matrices T1 and T2. The coefficient α is a parameter that adjusts the weighting of the errors in the translation and rotation components, similar to equation (2). Homogeneous transformation matrix w T c1、 w T c2These represent, for example, the first imaging posture C1 and the second imaging posture C2 as seen from the robot's base coordinate system w. System 1 can also train the posture estimation model 3 to estimate object postures Tc1 and Tc2 by machine learning that minimizes the error function F(T1, T2) in equation (3) above, thereby satisfying relation (1).

[0129] In the above explanation, an example was described in which two training data D2s are sampled from the training dataset DS2 in the selection of training data pairs DP2a (S21). The system 1 is not limited to this, and for example, in the training data acquisition process (S1), pairs of background images B1 and B2 may be sampled from the background dataset DS1, and each training data pair DP2a may be determined by sampling multiple times. For example, the control unit 50 may perform image acquisition (S11 to S13) for each background image B1 and B2 and generate the training dataset DS2 so as to sequentially store the training data pair DP2a corresponding to each pair of background images B1 and B2 (S15). In this case, the machine learning process (S2) may be performed to sequentially select training data pairs DP2a from the training dataset DS2 (S21).

[0130] In the above embodiment, an example was described in which a change in the appearance of the unsteady portion 12 of the target object 11 occurs due to deformation or the like in a pair of captured images A1 and A2 of the training data pair DP2a. The System 1 is not limited to this, and it is not necessary for a change in the appearance of the unsteady portion 12 due to deformation or the like to occur in a pair of captured images A1 and A2 of the training data pair DP2a. Even in such a case, if a change in the appearance of the unsteady portion 12 due to deformation or the like occurs in a pair of captured images A3 and A4 of training data pair DP2b, which is different from training data pair DP2a, then, similar to the above embodiment, it is possible to perform machine learning for image recognition even if the target object 11 has an unsteady portion 12. In this case, the captured images A1 and A2 of one training data pair DP2a are an example of the first captured images, and the captured images A3 and A4 of the other training data pair DP2b are an example of the second captured images.

[0131] 3. Summary As described above, in the machine learning system 1 of this embodiment, the machine learning device 5 comprises a communication I / F 52, which is an example of a communication unit, and a control unit 50. The communication I / F 52 is communicably connected to a robot 20, which is an example of an external device including a camera 2, which is an example of an imaging device. The control unit 50 controls the robot 20 via the communication I / F 52. The target object 11 includes a transient portion 12, which is an example of a first part, having physical properties that cause changes in appearance separate from changes in posture, and a steady portion 13, which is an example of a second part, that is less likely to cause changes in appearance due to the above-mentioned physical properties than the transient portion 12. The control unit 50, via the communication I / F 52, causes the camera 2 to capture a first image A1 including the target object 11 in a first imaging posture C1 as viewed from the camera 2, and via the communication I / F 52, causes the camera 2 to capture a second image A2 in a second imaging posture C2 different from the first imaging posture C1, in which the unsteady portion 12 of the target object 11 has changed in appearance from the first image A1 due to the above physical characteristics, and generates at least one of a training dataset DS2 in machine learning in which the unsteady portion 12 of the target object 11 is image recognized, and an example of a trained model, an orientation estimation model 3.

[0132] According to the machine learning device 5 described above, even if the target object 11 has non-stationary parts 12, machine learning focusing on the stationary parts 13 becomes possible, making it easier to perform machine learning for image recognition of the target object 11. For example, this system 1 can reduce the processing load of such image recognition machine learning.

[0133] In the machine learning apparatus 5 of this embodiment, the physical properties of the transient portion 12 include at least one of flexibility, light transmission, and specular reflectivity. The machine learning apparatus 5 of this embodiment makes it easier to perform machine learning for image recognition even for an object 11 in which the transient portion 12 undergoes a change in appearance separate from a change in posture due to these physical properties. The transient portion 12 of the object 11 has the above physical properties to the extent that its appearance can be changed such that the relationship in equation (1) above does not hold between the first and second captured images A1 and A2. Furthermore, the steady portion 13 of the object 11 does not have to have the above physical properties, or it may have them within a lower range than the transient portion 12.

[0134] In the machine learning device 5 of this embodiment, the first captured image A1 is captured including the first background image B1, and the second captured image A2 is captured including the second background image B2, which is different from the first background image B1. By capturing the first and second captured images A1 and A2, which include these different background images B1 and B2, the machine learning device 5 of this embodiment can facilitate machine learning for image recognition of the target object 11. For example, this system 1 can realize annotation-less machine learning for image recognition such as pose estimation of the target object 11.

[0135] In this embodiment, the control unit 50 selects, for example, the first and second captured images A1 and A2 from the background dataset DS1 (S21), thereby setting the first and second background images B1 and B2 such that the correspondence between the change in the posture of the target object 11 in the first and second captured images A1 and A2 and the change in the first and second captured postures C1 and C2 does not hold true for the first and second background images B1 and B2. This makes it easier to perform machine learning image recognition of the target object 11 by preventing the above equation (1) from holding true in the background portion of each captured image A1 and A2.

[0136] In setting background images B1 and B2 as described above, it is also possible to avoid settings that would satisfy the correspondence in equation (1) above, such as simply shifting the position of a background with a repeating pattern like a checkerboard. In other words, background images B1 and B2 may be set in a way that breaks the correspondence that would appear to be a physical change in their orientation.

[0137] In this embodiment, the imaging range of the camera 2 in at least one of the first and second captured images A1 and A2 may be narrower than the area in which the corresponding background image among the background images B1 and B2 is displayed on the display device 10. This avoids the influence of reflections from the edges of the display device 10 on machine learning, and makes it easier to perform machine learning for image recognition of the target object 11.

[0138] In this embodiment, the imaging range of the camera 2 in at least one of the first and second captured images A1 and A2 may be set so that the background images B1 and B2 are captured in all parts other than the target object 11. This avoids the influence of parts other than the target object 11 on machine learning and makes it easier to perform machine learning for image recognition of the target object 11.

[0139] In this embodiment, the first and second imaging postures are, for example, the posture of an imaging device such as the camera 2 when the first and second imaging images A1 and A2 are captured, respectively. The communication I / F 52 is connected to a robot 20 equipped with the camera 2, which is an example of an external device. The control unit 50 controls the robot via the communication I / F 52 to make the posture of the camera 2 different in the first imaging posture C1 and the second imaging posture C2 (S12), thereby causing the robot to capture the first and second imaging images A1 and A2 (S13). This makes it easier to perform machine learning for image recognition of the target object 11 using the robot 20 in this system 1. The robot 20 in this system 1 may be operated based on a trained posture estimation model 3. The posture estimation model 3 may perform image recognition to estimate the posture of the target object 11 based on the images captured by the camera 2.

[0140] In this embodiment, the first and second imaging postures may be the postures of the target object 11 when the first and second imaging images A1 and A2 are captured, respectively. Even in this case, the system 1 can easily perform machine learning for image recognition of the target object 11 in the same manner as described above, using training data D2 that associates the posture of the target object 11 with the corresponding imaging images.

[0141] In this embodiment, the communication interface 52 is connected to a display device 10 that displays an image on the background in which the target object 11 is placed. The control unit 50 causes the display device 10 to display a first background image B1 via the communication interface 52, causing the camera 2 to capture a first captured image A1, and then causes the display device 10 to display a second background image B2 via the communication interface 52, causing the camera 2 to capture a second captured image A2 (S11 to S13). As a result, the system 1 can easily perform machine learning for image recognition of the target object 11 by changing the background images B1 and B2 in each captured image A1 and A2 through a simple process of controlling the display device 10.

[0142] In this embodiment, for example, the control unit 40 of the robot control device 4 may perform image recognition of the target object 11 using a posture estimation model 3 obtained as a result of machine learning based on the first and second captured images A1, A2 and the first and second captured postures C1, C2. The robot control device 4 and the machine learning device 5 may be provided as an integrated unit. For example, the robot control device 4 may be provided with a program in which it executes various functions of the machine learning device 5 stored in a storage unit 41.

[0143] In the machine learning apparatus 5 of this embodiment, the control unit 50 may generate a training dataset DS2 based on the first, second, ..., nth captured images A1, A2, ..., AN and the first, second, ..., nth captured poses C1, C2, ..., CN (S15). The training dataset DS2 includes the first, second, ..., nth captured images A1, A2, ..., AN and the first, second, ..., nth captured poses C1, C2, ..., CN in relation to each other, as shown in Figure 9, for example. The machine learning apparatus 5 of this embodiment can facilitate machine learning for image recognition of the target object 11 by generating such a training dataset DS2. In this system 1, the subsequent machine learning processing (S2) may be performed outside the machine learning apparatus 5.

[0144] In the machine learning device 5 of this embodiment, the control unit 50 may perform machine learning to establish a correspondence between the change in the posture of the target object 11 in the first and second captured images A1 and A2 and the change in the first and second captured postures C1 and C2, thereby generating a trained posture estimation model 3 (S27). The machine learning device 5 of this embodiment makes it easier to perform machine learning for image recognition of the target object 11. The machine learning device 5 of this embodiment does not need to generate a training dataset DS2; for example, it may acquire a training dataset DS2 generated by the robot control device 4 and execute the machine learning process (S2).

[0145] In this embodiment, the posture estimation model 3 includes a posture estimation unit 32 that performs image recognition to estimate the posture of the target object 11 in each captured image. The machine learning device 5 of this embodiment makes it easier to perform machine learning of image recognition for posture estimation of the target object 11 by the posture estimation model 3.

[0146] In the posture estimation model 3 of this embodiment, the posture estimation unit 32 refers to a 3D model 30 which is an example of the shape model of the target object 11, and estimates the posture of the target object 11 in each captured image by comparing the rendered image of the 3D model 30 with each captured image (see Figures 12(A) to (E)). As a result, the posture estimation model 3 of this embodiment can estimate the object posture by matching it with the rendered image of the 3D model 30.

[0147] In the attitude estimation model 3 of this embodiment, the attitude estimation unit 32 generates a rendering image based on the ROI parameter Pm, which is an example of a parameter representing the region of the steady-state portion 13 in the 3D model 30 (Figure 12(C)). As a result, the attitude estimation model 3 of this embodiment can accurately estimate the attitude of the target object 11 by focusing on the steady-state portion 13 in the 3D model 30.

[0148] In the posture estimation model 3 of this embodiment, the posture estimation unit 32 compares a cropped image (Figure 12(B)) obtained by cutting out a region including the steady portion 13 from each captured image with a rendered image (Figure 12(D)) to estimate the posture of the target object 11 (Figure 12(E)). As a result, the posture estimation model 3 of this embodiment can accurately estimate the posture of the target object 11 by focusing on the steady portion 13 in the image of the target object 11 captured by the camera 2.

[0149] In this embodiment, the posture estimation model 3 further includes a segmentation unit 31 that extracts the region of the target object 11 in each captured image. The posture estimation unit 32 estimates the posture of the target object 11 in each captured image based on a mask image showing the region extracted by the segmentation unit 31. This makes it easier to estimate the posture of the target object 11 in the posture estimation model 3 of this embodiment.

[0150] In this embodiment, a machine learning method is provided that is performed, for example, by a control unit 50 of a machine learning device 5. This method includes the steps of: causing the camera 2 to capture a first image A1 including the target object 11 and a first background image in a first imaging posture C1 from which the camera 2 views the target object 11; causing the camera 2 to capture a second image A2 including the target object 11 and a second background image different from the first background image in a second imaging posture C2 different from the first imaging posture C1; and generating at least one of training data D2 for machine learning in which the target object 11 is image recognized and a posture estimation model 3 based on the first and second image images A1, A2 and the first and second imaging postures C1, C2.

[0151] In this embodiment, a program is provided to cause the control unit 50 to execute the above-described machine learning method. According to the machine learning method of this embodiment, it is possible to easily perform machine learning for image recognition of the target object 11.

[0152] In this embodiment, the robot control device 4 includes a communication I / F 42, which is an example of a communication unit, that is communicably connected to a robot arm 21, which is an example of a manipulator on which a camera 2 is positioned, and a control unit 40 that controls the robot arm 21 via the communication I / F 42. The control unit 40 controls the robot arm 21 via the communication I / F 42 to cause the camera 2 to capture a first image A1 including the target object 11 and a first background image in a first imaging posture C1 as viewed from the camera 2 (S12-S13), and controls the robot arm 21 via the communication I / F 42 to cause the camera 2 to capture a second image A2 including the target object 11 and a second background image different from the first background image in a second imaging posture C2 different from the first imaging posture C1 (S12-S13). The control unit 40 operates the robot arm 21 by applying a posture estimation model 3, which has been machine-learned based on the first and second captured images A1 and A2 and the first and second captured postures C1 and C2 (S3).

[0153] According to the robot control device 4 of this embodiment, by controlling the robot 20, captured images A1 and A2 with different imaging poses C1 and C2 can be obtained, making it easier to perform machine learning for image recognition of the target object 11. Furthermore, in the robot control device 4 of this embodiment, the control unit 40 can operate the robot arm 21 by performing image recognition of the target object 11 in the captured image by the camera 2 during operation, for example, based on the trained pose estimation model 3 applied in step S3.

[0154] (Embodiment 2) Embodiment 2 will be described below with reference to Figures 13 to 14. Embodiment 1 described a machine learning system 1 in which a segmentation unit 31 is included in the pose estimation model 3. Embodiment 2 describes a machine learning system 1 in which the pose estimation model 3A automatically generates prompts for the segmentation unit 31.

[0155] Hereinafter, descriptions of the machine learning system 1 according to this embodiment will be omitted as appropriate, and descriptions of the configuration and operation similar to those of the machine learning system 1 according to Embodiment 1 will be given.

[0156] 1. Regarding the posture estimation model, Figure 13 illustrates the posture estimation model 3A in the machine learning system 1 of Embodiment 2. In addition to having the same configuration as the posture estimation model 3 of Embodiment 1 (Figure 11), the posture estimation model 3A of this embodiment further includes a prompt generation unit 33, as shown in Figure 13. The prompt generation unit 33 is composed of an image recognition model such as a neural network, and automatically generates prompts for the segmentation unit 31. Template matching or the like may be used in the prompt generation unit 33.

[0157] Furthermore, the segmentation unit 31 may generate multiple mask images for a single input image. The machine learning system 1 of this embodiment may use the weights of these multiple mask images, i.e., mask weights, generated by the segmentation unit 31, as parameters to be learned.

[0158] The prompts generated by the prompt generation unit 33 are information used by the segmentation unit 31 to specify the steady portion 13 of the target object 11 for segmentation when generating a mask image. Examples of prompts include one or more key points indicating the position of the steady portion 13 of the target object 11 in the image, a bounding box surrounding the target object 11, or both. Such key points are examples of indices in this embodiment.

[0159] The prompt generation unit 33 receives an image, such as a first or second captured image, and reference information as input. The reference information may include a template image of the target object 11, a template mask, template keypoints, or language input, or a combination of these.

[0160] The machine learning system 1 of this embodiment performs steps S1 to S3 in Figure 4, similar to Embodiment 1, and for example, the machine learning device 5 performs machine learning of the above-mentioned pose estimation model 3A in the machine learning process (S2). In the pose estimation model 3A of this embodiment, the prompt generation unit 33 can learn using backpropagation, reinforcement learning, Bayesian optimization, etc., similar to Embodiment 1.

[0161] In the machine learning process (S2) of this embodiment, machine learning may be performed with some of the learning parameters of the prompt generation unit 33, segmentation unit 31, and attitude estimation unit 32 fixed, or all learning parameters may be updated simultaneously. For example, the attitude estimation model 3A of this embodiment may be generated as a trained model by training the prompt generation unit 33 using a base model for the segmentation unit 31 and the attitude estimation unit 32.

[0162] 2. Regarding the prompt generation unit, Figure 14 shows an example of the configuration of the prompt generation unit 33 in the machine learning system 1 of this embodiment. The prompt generation unit 33 in this example configuration comprises, for example, a feature extraction unit 33a, a template feature calculation unit 33b, and a similarity calculation unit 33c, as shown in Figure 14.

[0163] For example, in steps S22 and S23 of the machine learning processing (S2), the feature extraction unit 33a generates a feature map when it receives images such as the first or second captured images A1 and A2 as input. The feature map is a two-dimensional map having a feature vector for each pixel of the input image. The feature extraction unit 33a can use the output of the intermediate layer of various models such as SAM, Stable Diffusion, DINO, and DINOv2 (see below) as features. Alternatively, the feature extraction unit 33a may extract features from multiple models and use the fusion result, such as the average, as features. C. Mathilde, et. al. "Emerging properties in self-supervised vision transformers." in ICCV, 2021. O. Maxime, et. al. "DINOv2: Learning robust visual features without supervision." arXiv preprint arXiv:2304.07193 (2023).

[0164] In the machine learning process (S2) of this embodiment, for example, the control unit 50 of the machine learning device 5 inputs a template image of the target object 11 to the feature extraction unit 33a and outputs a feature map. Next, for example, the template feature calculation unit 33b extracts feature vectors of the pixel regions corresponding to the target object 11 on the feature map based on the feature map and the template mask of the target object 11, and calculates template features by aggregation such as averaging.

[0165] Next, the similarity calculation unit 33c calculates a similarity, such as the cosine similarity with a template feature, for each pixel of the feature map, such as the first captured image A1, and generates a similarity map. Furthermore, the similarity calculation unit 33c detects the peak position of the pixel value in the similarity map and generates a prompt using the detected position as a keypoint.

[0166] In this configuration example, when the pose estimation model 3A is used after training, the template feature calculation unit 33b may be omitted in the prompt generation unit 33, and for example, the template features may be set in advance in the similarity calculation unit 33c.

[0167] In the prompt generation unit 33 of the above configuration example, the aggregation of template features is not limited to the mean; for example, the maximum value may be used, or the sum of the mean and the maximum value may be used. Alternatively, the template feature calculation unit 33b may cluster the features using a method such as kmean and extract multiple representative template features. In this case, the similarity calculation unit 33c may calculate a similarity map for each template feature.

[0168] Figure 15 is a diagram illustrating keypoint transfer in the attitude estimation model 3A of this system 1. Figure 15(A) shows an example of a template image in this system 1. Figure 15(B) shows an example of an image acquired after keypoint transfer from Figure 15(A).

[0169] In this system 1, for example, a reference key point P0 indicating a steady portion 13 of the target object 11 may be pre-assigned to the template image as reference information, as illustrated in Figure 15(A). The prompt generation unit 33 of this system 1 may, for example, apply the above processing to such reference information and transfer the reference key point P0 of the template image to images such as the first or second captured images A1 and A2, as illustrated in Figures 15(A) and (B). For example, the template feature calculation unit 33b generates feature quantities for the key point.

[0170] As illustrated in Figure 15(B), even if the captured image differs from the template image (Figure 15(A)) in appearance of the transient portion 12 of the target object 11, it is possible to obtain a keypoint P1 that indicates the position of the steady portion 13 of the target object. The prompt generation unit 33 of this embodiment may generate the keypoint P1 thus transferred as a prompt. Models such as DINO-ViT, SD+DINO, and GeoAware-SC (see below) can be used as the prompt generation unit 33. A. Shir, et. al. "Deep ViT features as dense visual descriptors." arXiv preprint arXiv:2112.05814 2.3 (2021): 4. Z. Junyi, et. al. "A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence." in NeurIPS, 2023. Z. Junyi, et. al. "Telling left from right: Identifying geometry-aware semantic correspondence." in CVPR, 2024.

[0171] Note that the template image (Figure 15(A)) does not need to be an image of the same object as the target object 11; it may be an image of an object similar to the target object 11. Even in such cases, if index information such as keypoints is added to the portion of the template image corresponding to the steady portion 13 of the target object 11, it is possible to identify the steady portion 13 of the captured image using techniques such as the Semantic Correspondence described above.

[0172] Furthermore, when using keypoints for prompts generated by the prompt generation unit 33, each keypoint may be assigned a positive or negative label. For example, a positive keypoint indicates that the area is within the region of the target object 11 in the image, and a negative keypoint indicates that the area is outside the region of the target object 11 in the image. For example, if SAM or the like is used in the segmentation unit 31, the target object 11 can be segmented with greater accuracy by inputting positive and negative prompts to the segmentation unit 31, respectively.

[0173] In the prompt generation unit 33 employing the above configuration, the machine learning process (S2) of the system 1 may be performed to learn the positive and negative labels of keypoints. The positive and negative labels of keypoints may be learned using methods such as backpropagation, reinforcement learning, or Bayesian optimization, similar to the various learning parameters described above.

[0174] Furthermore, the prompt generation unit 33 is not limited to the configuration example using keypoints as described above. For example, it may be configured to estimate a bounding box surrounding the target object 11 from an image. For instance, the prompt generation unit 33 may accept language input specifying the target object 11, estimate the bounding box, and generate the estimation result as a prompt. The language input can be natural language specifying the target object 11, such as "a photo of a plastic bottle." Such a prompt generation unit 33 can use visual language models (VLM) such as OWL-ViT or bounding box estimation models such as R-CNN or YOLO (see below). M. Matthias, et. al. "Simple open-vocabulary object detection." in ECCV, 2022. G. Ross, et. al. "Rich feature hierarchies for accurate object detection and semantic segmentation." in CVPR, 2014.

[0175] 3. Summary As described above, in the machine learning system 1 of this embodiment, the posture estimation model 3A further includes a prompt generation unit 33 that automatically generates prompts to be input to the segmentation unit 31 based on each captured image. This makes it easier to perform machine learning image recognition of the target object 11 using the segmentation unit 31 in this system 1.

[0176] In the pose estimation model 3A of this embodiment, the prompt generation unit 33 refers to a template image of an example of a reference image of the target object 11 (Figure 15(A)) and transfers the reference keypoint P0 of an example index indicating the transient portion 12 in the template image to the transient portion 12 of the target object 11 in each captured image (Figure 15(B)). The prompt generation unit 33 generates a prompt based on the keypoint P1 of the transferred example index. As a result, the pose estimation model 3A of this embodiment can identify the region of the transient portion 12 from each captured image using the index transfer, making it easier to perform machine learning for image recognition of the target object 11.

[0177] (Embodiment 3) Embodiment 3 will be described below with reference to Figure 16. Embodiment 3 describes a machine learning system 1 in which the posture estimation model 3B does not use the segmentation unit 31 of Embodiments 1 and 2.

[0178] Hereinafter, descriptions of the configuration and operation similar to those of the machine learning system 1 in Embodiments 1 and 2 will be omitted as appropriate, and the machine learning system 1 according to this embodiment will be described.

[0179] Figure 16 illustrates the attitude estimation model 3B in the machine learning system 1 of Embodiment 3. The attitude estimation model 3B of this embodiment has the same configuration as the attitude estimation model 3 of Embodiment 1, but instead of the segmentation unit 31, it includes a slot attention unit 35 as shown in Figure 16. In this embodiment, the target object 11 is not limited to one type of object, but may include multiple types of objects.

[0180] In the pose estimation model 3B of this embodiment, the slot attention unit 35 is a module that uses slots set on the image for each type of object, such as the target object 11, for image segmentation attention. The slot attention unit 35 includes, for example, an initial slot generation unit 35a and a slot update unit 35b, as shown in Figure 16. The slots represent image features related to objects such as the target object 11 as latent representations acquired by machine learning.

[0181] The initial slot generation unit 35a generates a number of initial slots, i.e., initial slots, that is, a number of slots with a predetermined value. The slot update unit 35b automatically clusters regions in the image with similar feature quantities by repeatedly updating the initial slots and assigns them to each slot. Such a slot attention unit 35 can be configured, for example, in the same manner as in Non-Patent Document 2.

[0182] For example, in steps S23 and S24, the slot attention unit 35 generates slots for the first and second captured images A1 and A2, respectively, based on the first and second captured images A1 and A2 input, as shown in Figure 16, and outputs them to the attitude estimation unit 32. Based on these slots for the first and second captured images A1 and A2, the attitude estimation unit 32 calculates the first and second object attitudes Tc1 and Tc2 from the slot corresponding to the target object 11, for example.

[0183] In this slot attention unit 35, whether a region in the image is assigned to a slot depends on the initial slot. Similar to the embodiments described above, this system 1 performs machine learning on the pose estimation model 3B, which includes the slot attention unit 35, to satisfy the relational equation (1) described above. With this system 1, the initial slot generation unit 35a generates initial slots with high accuracy, and the slots can appropriately cluster the target object 11.

[0184] As described above, in the machine learning system 1 of this embodiment, the target object 11 includes one or more objects, and the posture estimation model 3 further includes a slot attention unit 35 that generates slots set for each object in each captured image. This system 1 makes it easier to perform machine learning using slot attention technology.

[0185] (Other Embodiments) As described above, Embodiments 1 to 3 have been explained as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited thereto and can be applied to embodiments that have been modified, substituted, added, or omitted as appropriate. Furthermore, it is possible to create new embodiments by combining the components described in each of the above embodiments. Therefore, other embodiments are described below as examples.

[0186] In the above embodiments 1 to 3, a machine learning system 1 was described that uses changes in the appearance of the unsteady portion 12 of the target object 11 due to deformation of the unsteady portion 12 (see Figure 5) in the machine learning of the posture estimation model 3. However, the machine learning system 1 of this embodiment is not limited to this. Such modifications will be explained with reference to Figures 17 to 18.

[0187] Figure 17 illustrates a modified configuration of the machine learning system 1A. In this embodiment, the machine learning system 1A may cause changes in the appearance of the transient portion 12 by, for example, changing the lighting conditions for the target object 11. In addition to the same configuration as in Embodiment 1, the system 1A may further include a lighting device 14, as shown in Figure 17. The lighting device 14 includes, for example, various light sources and a control circuit for controlling the emission of light from the light sources, and is arranged to irradiate the target object 11 with illumination light. The lighting device 14 may be used in the field 15 during operation, or it may be set up specifically for learning. The lighting device 14 of the system 1A is configured to change any of the lighting conditions, such as the amount of light, color, direction of irradiation, and position of irradiation light, as appropriate.

[0188] Figure 18 is a flowchart illustrating the training data acquisition process (S1) in a modified version of Figure 17, machine learning system 1A. In this modified version, machine learning system 1A, for example, operates in the same way as in Embodiment 1 (Figure 4), and as illustrated in Figure 18, further controls the change in lighting conditions in the training data acquisition process (S1) (S11A). For example, the control unit 50 of the machine learning device 5 in this system 1A controls the lighting device 14 via the communication I / F 52 to change the lighting conditions each time it acquires each captured image A1 to A2 (S13) (S11A). For example, if the non-stationary portion 12 of the target object 11 has metallic luster, such a change in lighting conditions can cause a change in the appearance of the non-stationary portion 12 in the corresponding captured image.

[0189] As described above, in the machine learning device 5 of the machine learning system 1A in this embodiment, the first and second captured images A1 and A2 are captured with the target object 11 illuminated under different lighting conditions. This makes it easier for the system 1 to perform machine learning on the target object 11 by changing the appearance of the transient portion 12 of the target object 11 by changing the lighting conditions.

[0190] In each of the embodiments described above, a background dataset DS1 is prepared (S10), and a machine learning system 1 used for the training data capture process (S1), which is the collection of training data D2, is described. In the machine learning system 1 of this embodiment, the background dataset DS1 may be prepared so as not to include background images that unintentionally have a negative impact on training (S10). For example, if the background dataset DS1 contains multiple identical or similar image data, there is concern about the impact on machine learning in the relational expression (1) described above. Therefore, for example, in step S10, the control unit 50 may calculate the pixel error between background images and exclude images that fall below a predetermined threshold from the background dataset DS1.

[0191] Furthermore, in the training data acquisition process (S1), if the differences in background images among the multiple captured images unintentionally resemble changes in posture, there is concern about the impact on the machine learning described above. Therefore, the machine learning system 1 of this embodiment may, for example, create a background check dataset by capturing images multiple times while changing the background image before step S1, when the target object 11 is not yet set up. In this case, for example, the control unit 50 may sample a pair of captured images and their respective imaging postures, calculate the pixel error between the image obtained by deforming one of the captured images in accordance with the change from one imaging posture to the other, and the other captured image, and check in advance whether there are any pairs of captured images that fall below a predetermined threshold. In the subsequent training data acquisition process (S1), by capturing images with the same imaging posture and displaying the same background image as during the pre-check while the target object 11 is set up, it is possible to guarantee that no background data that would adversely affect learning is included.

[0192] Furthermore, in each of the above embodiments, a machine learning system 1 using a background dataset DS1 containing multiple background images B1 to BK has been described. In this embodiment, the machine learning system 1 may use video data that sequentially displays multiple background images B1 to BK, i.e., background video data, instead of or in addition to the background dataset DS1. For example, in the learning data capture process (S1) of this embodiment, the control unit 50 may manage the background image in the captured image by synchronizing the timing of playback display of the background video data with the timing of the camera 2 capturing the captured image. Also, in this embodiment, the background video data may be generated in such a way as to ensure that it does not contain background data that would adversely affect learning, for example, as described above.

[0193] In each of the embodiments described above, a posture estimation model 3 in which a 3D model 30 is input to a posture estimation unit 32 has been described. In this embodiment, the posture estimation model 3 may perform preprocessing before inputting the 3D model 30 to the posture estimation unit 32. For example, this preprocessing may involve shifting the origin of the 3D model 30 by an offset amount defined by position coordinates such as three dimensions, and setting an offset origin. For example, the posture estimation unit 32 may crop a range of interest, such as a predetermined diameter range, from the 3D model 30, centered on this offset origin, and use it for posture estimation. This allows posture estimation to be performed by focusing only on a narrow area around the offset origin, thereby improving the accuracy of posture estimation. In the machine learning process (S2) of this embodiment, the offset amount and range of interest of this preprocessing may also be learned by machine learning in the same way as the various learning parameters in each of the embodiments described above. Furthermore, when the posture estimation model 3 performs model-based posture estimation by rendering the 3D model 30 and comparing it with captured images A1 and A2 to estimate the posture of the target object 11, the lighting conditions or camera settings (e.g., internal parameters or distortion coefficients) during rendering may also be machine-learned in the same way as the various learning parameters in each of the embodiments described above.

[0194] Furthermore, while the above embodiments describe an example in which a 3D model 30 is input to the pose estimation unit 32, the pose estimation model 3 of this embodiment is not particularly limited to this, and does not necessarily require the use of a 3D model 30. For example, the pose estimation model 3 of this embodiment may be configured using a model-free pose estimation method such as SAMPose or Gen6D (see below). S. Wubin, et. al. "SAMPose: Generalizable model-free 6D object pose estimation via single-view prompt." RA-L, 2025. L. Yuan, et. al. "Gen6D: Generalizable model-free 6-DOF object pose estimation from RGB images." in ECCV, 2022.

[0195] Furthermore, in each of the above embodiments, background images B1 to BK were displayed on a display device 10 located behind the target object 11 as seen from the camera 2. However, background images B1 to BK are not limited to the background; they may also be in the foreground of the target object 11. For example, in this system 1, a see-through display may be placed as the display device 10 between the camera 2 and the target object 11, and the background image may be displayed on this display device 10. This makes it possible to perform machine learning to improve the accuracy of image recognition of the target object 11, even if there are objects in the foreground of the target object 11 in the captured image.

[0196] Furthermore, in each of the above embodiments, a communication I / F 52 was exemplified as an example of the communication unit of the machine learning device 5. In this embodiment, the communication unit may be a circuit configuration such as an internal bus for internal communication of the embedded device, and for example, it may perform signal communication between the machine learning device 5 and an external device in an embedded device in which the two are configured as an integrated unit.

[0197] Furthermore, in each of the above embodiments, examples of pose estimation models 3, 3A, and 3B were described as the learning targets for the machine learning system 1. In this embodiment, the learning targets for the machine learning system 1 are not limited to pose estimation models 3, 3A, and 3B, but may be various image recognition models. For example, in this system 1, the segmentation unit 31, pose estimation unit 32, prompt generation unit 33, and slot attention unit 35 described above may each be used as the image recognition models to be learned.

[0198] Such machine learning can be performed, for example, in the machine learning process (S2), by fixing the learning parameters of parts other than the target of learning and updating the learning parameters of the target part in the pose estimation models 3, 3A, and 3B of each embodiment described above. The image recognition to be learned in this system 1 may be image segmentation, keypoint estimation, keypoint transfer, bounding box estimation, or slot attention, or detection of the target object 11 by these various estimations. As a result of such machine learning process (S2), this system 1 can generate trained models of various image recognition models.

[0199] Furthermore, in each of the above embodiments, a machine learning system 1 used in machine learning processing (S2) was described in which a training dataset DS2 was generated in the training data capture process (S1). However, this system 1 does not necessarily have to generate a training dataset DS2. For example, the machine learning device 5 of this embodiment may perform processing for each training data pair DP2a (S22 to S25) in machine learning processing (S2) while controlling image capture in the same way as steps S11 to S13 of the training data capture process (S1). In this case, for example, the training data pair DP2a is temporarily held in the storage unit 51 of the machine learning device 5. The machine learning system 1 of this embodiment can also generate a trained model in the same way as in each of the above embodiments.

[0200] Furthermore, although the above embodiments describe a machine learning system 1 that changes the background image of the target object 11 for machine learning, the machine learning system 1 of this embodiment is not limited to this, and does not need to change the background image. In the learning data acquisition process (S1) of this embodiment, the control unit 50 of the machine learning device 5 may, as appropriate, change the appearance of the non-stationary portion 12 and change the imaging posture (S12) without specifically performing background image selection and display (S11) to acquire the image captured by the camera 2 (S13).

[0201] This also makes it easier for the system 1 to perform machine learning for image recognition even when the target object 11 has a non-stationary portion 12, similar to the embodiments described above. The system 1 does not necessarily have to be applied to a site 15 with a particularly complex background. Furthermore, the system 1 does not need to be equipped with a display device 10, nor does it need to store the background dataset DS1 in the storage unit 51. The learning dataset DS2 of this embodiment can omit information about the background image, such as the background ID.

[0202] Furthermore, in each of the above embodiments, a robot control device 4 that grips the steady portion 13 of the target object 11 has been described. The robot control device 4 of this embodiment does not necessarily have to grip the steady portion 13, and may perform various controls such as manipulator control using the posture estimation of the steady portion 13.

[0203] Furthermore, in each of the above embodiments, a machine learning system 1 was described in which a trained posture estimation model 3 is applied to the robot control device 4. This system 1 is not limited to controlling the robot 20, but may be applied to various applications that use various image recognition methods for various target objects 11.

[0204] (Examples of embodiments) The following are examples of embodiments related to this disclosure.

[0205] A first aspect of the present disclosure is a machine learning apparatus comprising a communication unit that is communicably connected to an external device including an imaging device that captures an image of a target object, and a control unit that controls the external device via the communication unit. The target object includes a first part having physical properties that cause a change in appearance separate from a change in posture, and a second part that is less susceptible to changes in appearance due to physical properties than the first part. The control unit, via the communication unit, causes the imaging device to capture a first image including the target object and a first background image in a first imaging posture as seen from the imaging device, and via the communication unit, causes the imaging device to capture a second image in which the first part of the target object has undergone a change in appearance due to physical properties from the first image in a second imaging posture different from the first imaging posture, and generates at least one of a training dataset and a trained model in machine learning in which the second part of the target object is image recognized based on the first and second image images and the first and second imaging postures.

[0206] A second embodiment is a machine learning apparatus according to the first embodiment, wherein the physical properties of the first part include at least one of flexibility, light transmission, and specular reflectivity.

[0207] The third embodiment is a machine learning apparatus according to the first or second embodiment, wherein the first captured image is captured including a first background image, and the second captured image is captured including a second background image different from the first background image.

[0208] The fourth embodiment is a machine learning apparatus as described in the third embodiment, wherein the control unit sets the first and second background images such that the correspondence between the change in the pose of the target object in the first and second captured images and the change in the first and second captured pose does not hold true for the first and second background images.

[0209] The fifth aspect is a machine learning apparatus according to the third or fourth aspect, wherein the imaging range of the imaging apparatus in at least one of the first and second captured images is narrower than the area in which the corresponding background image among the first and second background images is displayed on the display device.

[0210] The sixth embodiment is a machine learning apparatus described in any of the third to fifth embodiments, wherein the communication unit is connected to a display device that displays an image on the background on which the target object is placed. The control unit causes the display device to display a first background image via the communication unit, causing the imaging device to capture a first image, and causes the display device to display a second background image via the communication unit, causing the imaging device to capture a second image.

[0211] The seventh aspect is a machine learning apparatus described in any of the first to sixth aspects, wherein the first and second imaging postures are the postures of the imaging apparatus when the first and second imaging images are captured, respectively.

[0212] The eighth aspect is the machine learning apparatus described in the seventh aspect, wherein the external equipment is a robot equipped with an imaging device. The control unit controls the robot via the communication unit so that the orientation of the imaging device differs between the first imaging orientation and the second imaging orientation, thereby causing the robot to capture first and second images.

[0213] The ninth aspect is a machine learning apparatus described in any of the first to sixth aspects, wherein the first and second imaging postures are the postures of the target object when the first and second imaging images are captured, respectively.

[0214] The tenth embodiment is a machine learning apparatus according to any of the first to ninth embodiments, wherein the control unit generates a training dataset based on first and second captured images and first and second imaging poses. The training dataset includes the first captured image and the first imaging pose associated with each other, and includes the second captured image and the second imaging pose associated with each other.

[0215] The eleventh embodiment is a machine learning apparatus according to any of the first to tenth embodiments, wherein the control unit performs machine learning to establish a correspondence between the change in the pose of the target object in the first and second captured images and the change in the first and second captured poses, thereby generating a trained model.

[0216] The twelfth embodiment is a machine learning apparatus according to any of the first to eleventh embodiments, wherein the trained model includes an pose estimation unit that performs image recognition to estimate the pose of the target object in each captured image.

[0217] The thirteenth aspect is a machine learning apparatus described in the twelfth aspect, in which the attitude estimation unit refers to a shape model of the target object and estimates the attitude of the target object in each captured image by comparing a rendered image of the shape model with each captured image.

[0218] In the fourteenth aspect, in the machine learning apparatus described in the thirteenth aspect, the posture estimation unit generates a rendering image based on parameters representing the region of the second part of the shape model.

[0219] The fifteenth embodiment is a machine learning apparatus according to the thirteenth or fourteenth embodiment, in which the posture estimation unit compares a cropped image obtained by cutting out a region including the second portion from each captured image with a rendered image to estimate the posture of the target object.

[0220] The sixteenth embodiment is a machine learning apparatus according to any of the twelfth to fifteenth embodiments, wherein the trained model further comprises a segmentation unit that extracts the region of the target object in each captured image. The pose estimation unit estimates the pose of the target object in each captured image based on the region extracted by the segmentation unit.

[0221] The 17th embodiment is a machine learning apparatus according to the 16th embodiment, wherein the trained model further comprises a prompt generation unit that automatically generates prompts to be input to the segmentation unit based on each captured image.

[0222] The 18th aspect is a machine learning apparatus as described in the 17th aspect, wherein the prompt generation unit refers to a reference image of the target object, transfers an index indicating a second portion in the reference image to the second portion of the target object in each captured image, and generates a prompt based on the transferred index.

[0223] The 19th aspect is a machine learning apparatus according to the 12th aspect, wherein the target object includes one or more objects, and the trained model further comprises a slot attention unit that generates slots set for each object in each captured image.

[0224] The 20th embodiment is a machine learning apparatus described in any of the first to 19 embodiments, wherein the first and second captured images are captured with the target object illuminated under different lighting conditions.

[0225] The 21st aspect is a machine learning method performed by a control unit using an imaging device that captures an image of a target object. The target object includes a first part having physical properties that cause a change in appearance separate from a change in posture, and a second part that is less susceptible to changes in appearance due to physical properties than the first part. The method includes the steps of: causing the imaging device to capture a first image including the target object in a first imaging posture as viewed from the imaging device; causing the imaging device to capture a second image in a second imaging posture different from the first imaging posture, in which the first part of the target object has undergone a change in appearance due to physical properties from the first image; and generating at least one of a training dataset and a trained model for machine learning in which the second part of the target object is image-recognized, based on the first and second image images and the first and second imaging postures.

[0226] The 22nd aspect is a program for causing a control unit to execute the machine learning method described in the 21st aspect.

[0227] The 23rd embodiment is a robot control device comprising a communication unit that is communicably connected to a manipulator on which an imaging device for capturing images of a target object is arranged, and a control unit that controls the manipulator via the communication unit. The target object includes a first part having physical properties that cause changes in appearance separate from changes in posture, and a second part that is less likely to cause changes in appearance due to physical properties than the first part. The control unit operates the manipulator via the communication unit based on an image of the target object captured by the imaging device and a trained model in which machine learning is performed to image recognize the second part of the target object. For example, the control unit performs image recognition of the target object in the image captured by the imaging device based on the trained model and operates the manipulator to grasp the second part of the target object.

[0228] As described above, embodiments have been explained as examples of the technology in this disclosure. For this purpose, accompanying drawings and a detailed description have been provided.

[0229] Therefore, the components described in the attached drawings and detailed descriptions may include not only components essential for solving the problem, but also components that are not essential for solving the problem, provided that they illustrate the technology described above. For this reason, the mere presence of these non-essential components in the attached drawings and detailed descriptions should not be immediately assumed to mean that they are essential.

[0230] Furthermore, since the embodiments described above are for illustrative purposes of the technology described herein, various modifications, substitutions, additions, omissions, etc., can be made within the claims or their equivalents.

[0231] This disclosure is applicable to various technologies that use machine learning for image recognition of target objects, and can be applied to, for example, robot control.

Claims

1. A machine learning device comprising: a communication unit that is communicably connected to an external device including an imaging device that captures an image of a target object; and a control unit that controls the external device via the communication unit, wherein the target object includes a first part having physical properties that cause a change in appearance separate from a change in posture, and a second part that is less susceptible to changes in appearance due to the physical properties than the first part; the control unit, via the communication unit, causes the imaging device to capture a first image including the target object in a first imaging posture as viewed from the imaging device; the control unit, via the communication unit, causes the imaging device to capture a second image in which the first part of the target object has undergone a change in appearance due to the physical properties from the first image, in a second imaging posture different from the first imaging posture; and generates at least one of a training dataset and a trained model in machine learning in which the second part of the target object is image recognized based on the first and second image and the first and second imaging postures.

2. The machine learning apparatus according to claim 1, wherein the physical properties in the first part include at least one of flexibility, light transmission, and specular reflectivity.

3. The machine learning apparatus according to claim 1, wherein the first captured image is captured including a first background image, and the second captured image is captured including a second background image different from the first background image.

4. The machine learning apparatus according to claim 3, wherein the control unit sets the first and second background images such that a correspondence between the change in the posture of the target object in the first and second captured images and the change in the first and second captured posture does not exist in the first and second background images.

5. The machine learning apparatus according to claim 3, wherein the imaging range of the imaging device in at least one of the first and second captured images is narrower than the area in which the corresponding background image of the first and second background images is displayed on the display device.

6. The machine learning apparatus according to claim 3, wherein the communication unit is connected to a display device that displays an image on the background on which the target object is placed, the control unit causes the display device to display the first background image via the communication unit, causing the imaging device to capture the first image, and causes the display device to display the second background image via the communication unit, causing the imaging device to capture the second image.

7. The machine learning apparatus according to claim 1, wherein the first and second imaging postures are the postures of the imaging apparatus when the first and second imaging images are captured, respectively.

8. The machine learning apparatus according to claim 7, wherein the external device is a robot equipped with the imaging device, and the control unit controls the robot via the communication unit to cause the posture of the imaging device to be different in the first imaging posture and the second imaging posture, thereby causing the robot to capture first and second images.

9. The machine learning apparatus according to claim 1, wherein the first and second imaging postures are the postures of the target object when the first and second imaging images are captured, respectively.

10. The machine learning apparatus according to claim 1, wherein the control unit generates the learning dataset based on the first and second captured images and the first and second imaging postures, and the learning dataset includes the first captured image and the first imaging posture in relation to each other, and includes the second captured image and the second imaging posture in relation to each other.

11. The machine learning apparatus according to claim 1, wherein the control unit performs machine learning to establish a correspondence between the change in the posture of the target object in the first and second captured images and the change in the first and second captured postures, thereby generating the trained model.

12. The machine learning apparatus according to claim 1, wherein the trained model comprises a posture estimation unit that performs image recognition to estimate the posture of the target object in each captured image.

13. The machine learning apparatus according to claim 12, wherein the posture estimation unit estimates the posture of the target object in each captured image by comparing a rendered image of the shape model with each captured image, with respect to the shape model of the target object.

14. The machine learning apparatus according to claim 13, wherein the posture estimation unit generates the rendering image based on parameters representing the region of the second part in the shape model.

15. The machine learning apparatus according to claim 13, wherein the posture estimation unit estimates the posture of the target object by comparing a cropped image obtained by cutting out a region including the second portion from each captured image with the rendered image.

16. The machine learning apparatus according to claim 12, wherein the trained model further comprises a segmentation unit that extracts the region of the target object in each of the captured images, and the pose estimation unit estimates the pose of the target object in each of the captured images based on the region extracted by the segmentation unit.

17. The machine learning apparatus according to claim 16, further comprising a prompt generation unit that automatically generates prompts to be input to the segmentation unit based on each of the captured images, wherein the trained model is further a machine learning apparatus according to claim 16.

18. The machine learning apparatus according to claim 17, wherein the prompt generation unit refers to a reference image of the target object, transfers an index indicating the second portion in the reference image to the second portion of the target object in each captured image, and generates the prompt based on the transferred index.

19. The machine learning apparatus according to claim 12, wherein the target object includes one or more objects, and the trained model further comprises a slot attention unit that generates slots set for each object in each captured image.

20. The machine learning apparatus according to claim 1, wherein the first and second captured images are captured with the target object illuminated under different lighting conditions.

21. A machine learning method performed by a control unit using an imaging device that captures an image of a target object, wherein the target object includes a first part having physical properties that cause a change in appearance separate from a change in posture, and a second part that is less susceptible to changes in appearance due to the physical properties than the first part, and the method includes the steps of: causing the imaging device to capture a first image including the target object in a first imaging posture as viewed from the imaging device; causing the imaging device to capture a second image in a second imaging posture different from the first imaging posture, in which the first part of the target object has undergone a change in appearance due to the physical properties from the first image; and generating at least one of a training dataset and a trained model in machine learning in which the second part of the target object is image recognized, based on the first and second image and the first and second imaging postures.

22. A program for causing the control unit to execute the machine learning method described in claim 21.

23. A robot control device comprising: a communication unit that is communicably connected to a manipulator on which an imaging device for capturing an image of a target object is arranged; and a control unit that controls the manipulator via the communication unit, wherein the target object includes a first part having physical properties that cause a change in appearance separate from a change in posture, and a second part that is less likely to cause a change in appearance due to the physical properties than the first part, and the control unit operates the manipulator via the communication unit based on an image of the target object captured by the imaging device and a trained model in which machine learning is performed to image recognize the second part of the target object.