MACHINE LEARNING OF OBJECT RECOGNITION USING A ROBOT-GUIDED CAMERA
Patent Information
- Application Number
- DE502020012632
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-06
- Filing Date
- 2020-05-05
- Publication Date
- 2026-02-12
- Estimated Expiration
- 2040-05-05
AI Technical Summary
Existing machine learning methods for object recognition in robotics require manual marking of rectangular areas in numerous training images, especially when handling different objects, which is time-consuming and inefficient.
A method involving positioning a robot-guided camera in various poses to capture localization images, determining a virtual model of the learning object, and using machine learning to automate the determination of object poses, particularly through deep artificial neural networks, allowing for flexible and reliable object recognition.
Enables efficient and automated object recognition, enabling robots to interact flexibly with their environment by grasping or manipulating objects without prior knowledge of their positioning, improving robot operation and reducing manual intervention.
Description
[0001] The present invention relates to a method and system for machine learning object recognition using at least one robot-guided camera and a learning object, or for operating a robot using the learned object recognition, and to a computer program product for carrying out the method.
[0002] Object recognition enables robots to interact more flexibly with their environment, for example, by grasping, manipulating, or otherwise manipulating objects that are positioned in ways that are not known beforehand. Reference is made to "Vision-Guided Robot Control for 3D Object Recognition and Manipulation"; SQ et al.; 2008; InTech," which concerns image processing methods with a focus on improving the robustness and accuracy of robot systems. Furthermore, reference is made to "Machine Vision Approach for Robotic Assembly"; Peña-Cabrera, M et al.; Assembly Automation AA25-3, which describes a method for the online detection and classification of parts in robot assembly tasks and its application in an intelligent manufacturing cell.Furthermore, reference is made to DE 20 2017 001 227 U1, which relates to an object recognition system comprising a control device, a base whose position and orientation in space is known, such that base coordinate values which characterize the position and orientation of the base in space are stored in the control device.
[0003] Object recognition can be advantageously learned using machine learning. In particular, artificial neural networks can be trained to identify bounding boxes or rectangles, masks, or similar shapes in captured images.
[0004] Up to now, this has required manually marking the corresponding rectangular areas in a large number of training images, especially when different objects or object types are to be recognized or handled robotically.
[0005] One objective of an embodiment of the present invention is to improve machine learning for object recognition. Another objective of the present invention is to improve the operation of a robot.
[0006] These problems are solved by a method with the features of claim 1 and 6, respectively. Claims 8 and 9 protect a system and a computer program product for carrying out a method described herein. The dependent claims relate to advantageous embodiments.
[0007] According to one embodiment of the present invention, a method for machine learning object recognition using at least one robot-guided camera and at least one learning object comprises the following step: Positioning the camera in different poses relative to the learning object using a robot, wherein at least one localization image, in particular a two-dimensional and / or a three-dimensional localization image, is recorded and stored in each pose, which depicts the learning object.
[0008] This allows various localization images to be advantageously captured at least partially automatically, whereby the shooting perspectives or poses of the images relative to each other are known due to the known positions of the camera-guiding robot or the correspondingly known poses of the robot-guided camera, and are (specifically) predetermined in one version.
[0009] In one version, a pose describes a one-, two- or three-dimensional position and / or a one-, two- or three-dimensional orientation.
[0010] According to one embodiment of the present invention, the method comprises the following steps: Determining a virtual model of the learning object based on the poses and at least some of these localization images; and determining a pose of a reference of the learning object in one or more training images captured with the camera, in a version in one or more of the localization images and / or one or more images with one or more interfering objects not depicted in at least one of these localization images, based on this virtual model.
[0011] By determining a virtual model of the learning object and using this model to determine a pose of a reference of the learning object, this pose can be advantageously at least partially automated and thus easily and / or reliably determined and then used for machine learning, especially in the case of learning objects that are not known beforehand.
[0012] Accordingly, the method according to one embodiment of the present invention comprises the following step: Machine learning of object recognition of the reference based on the determined pose(s) in the training image(s).
[0013] The reference can be a simplified representation of the learning object in one version, a body, in particular a three-dimensional body, in particular an envelope body in one version, an (enclosing) polyhedron, in particular a cuboid, or the like, a curve, in particular a two-dimensional curve, in particular an envelope curve, an (enclosing) polygon, in particular a rectangle, or the like, a mask of the learning object, or the like.
[0014] In one embodiment, the robot has at least three, in particular at least six, and in another embodiment at least seven, axes, in particular rotary joints.
[0015] In one embodiment, advantageous camera positions can be reached by means of at least three axes, advantageous camera poses by means of at least six axes, and advantageously redundantly by means of at least seven axes, so that, for example, obstacles can be avoided or the like.
[0016] In one embodiment, the machine learning comprises training a deep artificial neural network, in particular a deep convolutional neural network or deep learning. This is a machine learning method particularly suitable for the present invention. Similarly, the object recognition (to be learned or learned by machine) comprises an artificial neural network (to be trained or trained), and can be implemented in this way.
[0017] In one implementation, determining the virtual model involves reconstructing a three-dimensional scene from localization images; in another implementation, it involves a method for simultaneous visual position determination and map creation ("visual SLAM").
[0018] This allows the virtual model to be determined advantageously, in particular simply and / or reliably, in a single execution, even when the pose and / or shape of the learning object is unknown.
[0019] Additionally or alternatively, determining the virtual model in one implementation includes at least partially eliminating an environment depicted in localization images.
[0020] This allows the virtual model to be determined advantageously, in particular simply and / or reliably, in a single execution, even when the pose and / or shape of the learning object is unknown.
[0021] Additionally or alternatively, the learning object is arranged in a known environment, which is empty in one version, while the localization images are taken, in particular on an (empty) surface, in a version of known color and / or position, for example a table or the like.
[0022] This allows the elimination of the environment to be improved in one execution, in particular made simpler and / or more reliable.
[0023] Additionally or alternatively, determining the virtual model in one execution includes filtering, in one execution before and / or after reconstructing a three-dimensional scene and / or before and / or after eliminating the environment.
[0024] This allows the virtual model to be determined more advantageously, and in particular more reliably, in a single execution.
[0025] In one implementation, determining the virtual model includes determining a point cloud model. Accordingly, the virtual model can be a point cloud model, in particular.
[0026] This allows the virtual model to be determined in a particularly advantageous, especially simple, flexible and / or reliable way in a single execution, even when the pose and / or shape of the learning object is unknown.
[0027] Additionally or alternatively, determining the virtual model in one implementation includes determining a mesh model, in particular a polygon (mesh) model, in one implementation based on the point cloud model. Accordingly, the virtual model can have a (polygon) mesh model, in particular be one.
[0028] This can improve the (further) handling or use of the virtual model in one version.
[0029] In one implementation, determining the pose of the reference involves transforming a three-dimensional reference into one or more two-dimensional references. Specifically, the pose of a three-dimensional mask or three-dimensional (enveloping) body in the reconstructed three-dimensional scene and the corresponding individual localization images can first be determined and then transformed, in particular mapped, into the corresponding pose of a two-dimensional mask or a two-dimensional (enveloping) curve.
[0030] This makes it advantageous, and especially simple, to determine the pose of the two-dimensional reference.
[0031] In one implementation, determining the pose of the reference involves transforming a three-dimensional virtual model into one or more two-dimensional virtual models. Specifically, the pose of the virtual model can first be determined in the reconstructed three-dimensional scene and the corresponding individual localization images, and then the corresponding pose of a two-dimensional mask or a two-dimensional (enveloping) curve can be determined from these.
[0032] This allows the pose of the two-dimensional reference to be determined advantageously, and in particular reliably.
[0033] According to one embodiment of the present invention, a method for operating a, in particular the, robot comprises the following steps: Determining the pose of one or more references of an operating object using object recognition learned with a method or system described herein; and operating, in particular controlling and / or monitoring, this robot based on this pose.
[0034] An object recognition system learned according to the invention is particularly advantageously used to operate a robot, wherein, in one embodiment, this robot is also used to position the camera in different poses. In another embodiment, the camera-guiding robot and the robot that is operated based on the pose determined by means of object recognition are different robots. Controlling a robot, in one embodiment, comprises path planning and / or online control, in particular regulation. Operating the robot, in one embodiment, comprises contacting, in particular grasping, and / or manipulating the object being operated.
[0035] In one embodiment, at least one camera, guided by the robot being operated or by another robot, captures one or more recognition images, each depicting the operating object. In another embodiment, at least one recognition image is captured in different poses (relative to the operating object). The pose of the reference(s) of the operating object is / are determined in one embodiment based on this recognition image(s).
[0036] This allows the robot to interact advantageously, and in particular flexibly, with its environment, for example by contacting objects that are positioned in a way that is not known in advance, in particular by grasping, processing or the like.
[0037] In one implementation, the (used) object recognition is selected based on the operating object from several existing object recognitions, which have been learned for a specific object type in a single implementation using a method described here. In this implementation, after the respective training, the coefficients of the trained artificial neural network are stored, and a neural network is parameterized with the stored coefficients for each operating object on which the robot is to be based, or for its type, for example, for an object to be contacted, especially to be grasped or processed.
[0038] This allows an artificial neural network, particularly one with a uniform structure, to be parameterized for each operating object (type) in a single implementation, thereby selecting an operating object (type)-specific object recognition from several machine-learned object (type)-specific object recognitions. This allows for an improvement in the object recognition used to operate the robot and, consequently, in the robot's operation.
[0039] In one implementation, a one- or multi-dimensional environment and / or camera parameter is specified for object recognition, based on an environment and / or camera parameter from the machine learning process, specifically identical to the environment or camera parameter used in the machine learning process. The parameter can include, in particular, exposure, (camera) focus, or the like. In another implementation, the environment and / or camera parameter from the machine learning process is stored together with the learned object recognition.
[0040] This allows for object recognition used to operate the robot in one version, thereby improving the robot's operation.
[0041] In one implementation, determining the pose of a reference of the operating object involves transforming one or more two-dimensional references into a three-dimensional reference. In another implementation, the poses of two-dimensional references in various recognition images are determined using object recognition, and the pose of a corresponding three-dimensional reference is derived from these. For example, if the pose of a two-dimensional rectangular prism is determined using object recognition in each of three mutually perpendicular recognition images, the pose of a three-dimensional cuboid can be derived from these.
[0042] In one implementation, the position of the operating object is determined based on the position of the reference object; in another implementation, it is determined based on a virtual model of the operating object. For example, if the position of a three-dimensional envelope is determined, the virtual model of the operating object can then be aligned within this envelope, and the position of the operating object can be determined in this way. Similarly, in another implementation, the position of the operating object within the envelope can be determined using a matching process, particularly a three-dimensional one.
[0043] This can improve the operation of the robot, in particular contacting, especially gripping, and / or processing the operating object, for example by determining the position of suitable contact, especially gripping, or processing surfaces based on the virtual model or the pose of the operating object, or the like.
[0044] In one embodiment, one or more working positions of the robot, in particular one or more working poses of an end effector of the robot, are determined and planned based on the pose of the reference of the operating object, in particular the (derived) pose of the operating object. In another embodiment, this is done based on operating data specified for the operating object, in particular specified contact surfaces, especially gripping surfaces, or the like. In a further embodiment, a movement, in particular a path, of the robot for contacting, in particular gripping, or processing the operating object is planned and / or executed based on the determined pose of the reference or the operating object, and in a further development, based on the specified operating data, in particular contact surfaces, especially gripping surfaces, or processing surfaces.
[0045] According to one embodiment of the present invention, a system, in particular hardware and / or software, especially programming technology, is set up to carry out a method described herein.
[0046] According to one embodiment of the present invention, a system comprises: Means for positioning the camera in different poses relative to the learning object using a robot, wherein at least one localization image, in particular a two-dimensional and / or a three-dimensional localization image, depicting the learning object, is recorded and stored in each pose; means for determining a virtual model of the learning object based on the poses and at least some of the localization images; means for determining a pose of a reference of the learning object in at least one training image recorded with the camera, in particular at least one of the localization images and / or at least one image with at least one interfering object not depicted in at least one of the localization images, based on the virtual model; and means for machine learning object recognition of the reference based on this determined pose in the at least one training image.
[0047] Additionally or alternatively, the system features: Means for determining a pose of at least one reference of an operating object using object recognition learned with one described herein; and means for operating the robot based on this pose.
[0048] In one version, the system or its means has: Means for training an artificial neural network; and / or means for reconstructing a three-dimensional scene from localization images; and / or means for at least partially eliminating an environment depicted in localization images; and / or means for filtering; and / or means for determining a point cloud model; and / or means for determining a network model; and / or means for transforming a three-dimensional reference to at least one two-dimensional reference and / or a three-dimensional virtual model to at least one two-dimensional virtual model; and / or means for capturing at least one recognition image using at least one camera, in particular a robot-guided camera, depicting the operating object, and means for determining the pose based on this recognition image;and / or means for selecting object recognition based on the operating object from several existing object recognitions learned using a method described herein; and / or means for specifying an environment and / or camera parameter for object recognition based on an environment and / or camera parameter in machine learning; and / or means for transforming at least one two-dimensional reference into a three-dimensional reference; and / or means for determining a pose of the operating object based on the pose of the operating object's reference, in particular based on a virtual model of the operating object; and / or means for determining at least one working position of the robot, in particular a working pose of an end effector of the robot, based on the pose of the operating object's reference, in particular the pose of the operating object, in particular based on operating data specified for the operating object.
[0049] A means according to the present invention can be configured as hardware and / or software, in particular comprising a processing unit, preferably a microprocessor unit (CPU), graphics processing unit (GPU), or the like, preferably connected to a storage and / or bus system via data or signals, and / or comprising one or more programs or program modules. The processing unit can be configured to execute instructions implemented as a program stored in a storage system, to acquire input signals from a data bus, and / or to output signals to a data bus. A storage system can comprise one or more, in particular different, storage media, in particular optical, magnetic, solid-state, and / or other non-volatile media. The program can be configured to embody the methods described herein.is capable of executing such procedures, enabling the processing unit to perform the steps of such procedures and thus, in particular, to learn object recognition or operate the robot. A computer program product may, in one version, include a storage medium, particularly non-volatile, for storing a program or containing a program, wherein the execution of this program causes a system or controller, particularly a computer, to execute a procedure described herein or one or more of its steps.
[0050] In one implementation, one or more, in particular all, steps of the procedure are carried out fully or partially automatically, in particular by the system or its means.
[0051] In one version, the system features the robot.
[0052] Further advantages and features will become apparent from the dependent claims and the exemplary embodiments. These are shown, in part schematically: Fig. 1: according to an embodiment of the present invention; and Fig. 2: according to an embodiment of the present invention.
[0053] Fig. 1 Figure 1 shows a system according to an embodiment of the present invention with a robot 10, to whose gripper 11 a (robot-guided) camera 12 is attached.
[0054] First, using a robot controller 20, object recognition of a reference of a learning object 30, which is arranged on a table 40, is learned by machine.
[0055] To do this, the robot 10 positions S10 in one step (see Fig. 2 ) the camera 12 in different poses relative to the learning object 30, whereby in each pose a two-dimensional and a three-dimensional localization image is recorded and stored, which depicts the learning object 30. Fig. 1 The robot-guided camera 12 shows an example of such a pose.
[0056] Using a method for visual simultaneous position determination and map creation, a three-dimensional scene with the learning object 30 and the table 40 is reconstructed from the three-dimensional localization images ( Fig. 2 : step S20), from which an environment depicted in these localization images in the form of table 40 is eliminated, in particular segmented out ( Fig. 2 : Step S30).
[0057] After filtering out interference signals ( Fig. 2 (Step S40, where steps S30 and S40 can also be swapped) is used to determine a virtual point cloud model ( Fig. 2 : step S50), from which, for example using a Poisson method, a virtual network model of polygons is determined ( Fig. 2 : Step S60), which represents the learning object 30.
[0058] In step S70, a three-dimensional reference of the learning object 30 is determined in the form of a bounding box or other mask, and in step S80, this is transformed into the two-dimensional localization images from step S10. The pose of the three-dimensional reference is determined in each three-dimensional localization image, and from this, the corresponding pose of the two-dimensional reference is determined in the associated two-dimensional localization image, which the camera 12 captured in the same pose as this three-dimensional localization image.
[0059] Subsequently, in step S90, interfering objects 35 are placed on the table 40, and then further two-dimensional training images are captured. These images depict both the learning object 30 and the interfering objects 35 that were not shown in the localization images from step S10. The camera 12 is preferably positioned again in the same poses in which it captured the localization images. The three-dimensional reference of the learning object 30 is also transformed into these further training images in the manner described above.
[0060] Then, in step S100, an artificial neural network Al is trained to determine the two-dimensional reference of the learning object 30 in the two-dimensional training images, which now each contain the learning object 30, its two-dimensional reference and partly also additional interfering objects 35.
[0061] The machine-learned object recognition or the neural network Al trained in this way can now determine the corresponding two-dimensional reference, in particular an enclosing rectangle or another mask, in two-dimensional images in which the learning object 30 or a (sufficiently) similar, in particular identical, object 30' is depicted.
[0062] In order to grasp such operating objects 30' with the gripper 11 of the robot 10 (or another robot (gripper)), in a step S110 the corresponding object recognition, in particular the corresponding (trained) artificial neural network, is first selected, in an execution by parameterizing an artificial neural network Al with the corresponding parameters stored for the object recognition of these operating objects.
[0063] Then, in step S120, recognition images are taken with camera 12 in different poses relative to the operating object, which depict the operating object, and the two-dimensional reference is determined in these recognition images using the selected or parameterized neural network Al ( Fig. 2 : Step S130) and from this a three-dimensional reference of the operating object 30' in the form of an enclosing cuboid or another mask is determined by means of transformation ( Fig. 2 : Step S140).
[0064] Based on the pose of this three-dimensional reference and object-type-specific operating data specified for the operating object 30', for example a virtual model, specified gripping points or the like, a suitable gripping pose of the gripper 11 is then determined in step S150, which the robot approaches in step S160 and grips the operating object 30' ( Fig. 2 : Step S170).
[0065] Although exemplary embodiments were explained in the preceding description, it should be noted that a multitude of modifications are possible. Furthermore, it should be emphasized that the exemplary embodiments are merely examples and are not intended to restrict the scope of protection, applications, or structure in any way. Rather, the preceding description provides the skilled person with a guideline for implementing at least one exemplary embodiment, whereby various modifications, particularly with regard to the function and arrangement of the described components, can be made without departing from the scope of protection as defined by the claims and these equivalent combinations of features. Reference symbol list
[0066] 10 Robot 11 Gripper 12 Camera 20 Robot Controller 30 Learning Object 30 Operating Object 35 Disturbance Object 40 Table (Environment) Artificial Neural Network
Claims
1. A method of machine learning of an object recognition with the aid of at least one camera (12) guided by a robot and at least one learning object (30), the method comprising the steps of: - positioning (S10) the camera in different camera poses relative to the learning object with the aid of a robot (10), wherein at least one localisation image, in particular a two-dimensional and / or a three-dimensional localisation image, is captured in each of the camera poses and stored, wherein the localisation image depicts the learning object; - determining (S20-S60) a virtual model of the learning object on the basis of the camera poses and at least some of the localisation images, wherein the virtual model comprises a point cloud model; - determining (S70-S90), on the basis of the virtual model, a pose of a reference of the learning object in at least one training image captured with the camera, in particular in at least one of the localisation images and / or in at least one image with at least one interfering object not depicted in at least one of the localisation images, wherein the reference comprises a bounding volume; and - machine learning (S100) of an object recognition of the reference on the basis of this determined pose in the at least one training image.
2. The method according to claim 1, characterised in that the robot has at least three axes, in particular rotary joints.
3. The method according to any one of the preceding claims, characterised in that the machine learning comprises training an artificial neural network (AI).
4. The method according to any one of the preceding claims, characterised in that the determining of the virtual model comprises: - reconstructing (S20) a three-dimensional scene from localisation images; - at least partially eliminating (S30) an environment (40) depicted in localisation images; - filtering (S40); - determining (S50) a point cloud model.
5. The method according to any one of the preceding claims, characterised in that the determining of the pose of the reference comprises a transformation (S80) of a three-dimensional reference into at least one two-dimensional reference and / or a transformation (S80) of a three-dimensional virtual model into at least one two-dimensional virtual model.
6. A method of operating a robot (10), the method comprising the steps of: - determining (S130, S140) a pose of at least one reference of an operating object (30') with the aid of an object recognition that has been learned using a method according to any one of the preceding claims, wherein the reference comprises a bounding volume; and - operating (S150, S160) the robot on the basis of this pose.
7. The method according to the preceding claim, characterised in that - at least one camera, in particular a camera guided by a robot, captures (S120) at least one recognition image which depicts the operating object, and the pose is determined on the basis of this recognition image; - the object recognition is selected (S110) on the basis of the operating object from several existing object recognitions, each of which has been learned for an object type and which have been learned using a method according to any one of the preceding claims 1 to 5; - for the object recognition, an environment parameter and / or a camera parameter is specified on the basis of an environment parameter and / or a camera parameter which has been used in the machine learning; - the determining of the pose of the reference of the operating object comprises a transformation (S140) of at least one two-dimensional reference into a three-dimensional reference; - a pose of the operating object is determined on the basis of the pose of the reference of the operating object, in particular on the basis of a virtual model of the operating object; and / or - at least one working position of the robot, in particular a working pose of an end effector of the robot, is determined (S150) on the basis of the pose of the reference of the operating object, in particular on the basis of the pose of the operating object determined from this, in particular on the basis of operating data specified for the operating object.
8. A system for machine learning of an object recognition with the aid of at least one camera (12) guided by a robot and at least one learning object (30), and / or for operating a robot (10), which system is set up to carry out a method according to any one of the preceding claims and / or which comprises: - means for positioning the camera in different camera poses relative to the learning object with the aid of a robot (10), wherein at least one localisation image, in particular a two-dimensional and / or a three-dimensional localisation image, is captured in each of the camera poses and stored, wherein the localisation image depicts the learning object; - means for determining a virtual model of the learning object on the basis of the camera poses and at least some of the localisation images, wherein the virtual model comprises a point cloud model; - means for determining, on the basis of the virtual model, a pose of a reference of the learning object in at least one training image captured with the camera, in particular in at least one of the localisation images and / or in at least one image with at least one interfering object not depicted in at least one of the localisation images, wherein the reference comprises a bounding volume; and - means for machine learning of an object recognition of the reference on the basis of this determined pose in the at least one training image and / or comprises: - means for determining a pose of at least one reference of an operating object (30') with the aid of an object recognition that has been learned using a method according to any one of the preceding claims 1 to 5; and - means for operating the robot on the basis of this pose.
9. A computer program product with a program code which is stored on a computer-readable medium, for carrying out a method according to any one of the preceding claims 1 to 7.