DEVICE AND METHOD FOR MEASURING, INSPECTING OR PROCESSING OBJECTS

DE502021008102D1Active Publication Date: 2025-08-07CARL ZEISS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE502021008102
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-11
Filing Date
2021-11-09
Publication Date
2025-08-07
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

Existing coordinate measuring machines and robotic systems face challenges in achieving fast and precise positioning of measuring or machining tools relative to objects, particularly in industrial settings, due to their stationary nature and the need for the object to be brought to the machine, which limits throughput and accuracy.

Method used

A mobile platform with a kinematic system and an instrument head, equipped with sensors and a controller, uses a combination of initial and refined pose estimation methods, including machine learning and sensor fusion, to accurately position the instrument head on the target object, enabling precise alignment with an accuracy of up to ±0.1 mm.

Benefits of technology

The solution allows for rapid and precise positioning of the instrument head on the object, enhancing throughput and accuracy in industrial applications, supporting both measurement and processing tasks with high precision.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present application relates to devices and methods for measuring, inspecting or processing objects, which can be used in particular in the industrial production of objects such as motor vehicles, aircraft, components thereof or for measuring industrial facilities, but are not limited thereto.

[0002] Various devices and methods for measuring and processing objects are known in industry. Such devices and methods can be used, for example, for final inspection of a manufactured product, for inspection during production, or for processing a product during production.

[0003] One example of such devices is coordinate measuring machines. Such coordinate measuring machines typically comprise a measuring head system with a sensor for measuring the object, or a positioning system with which the measuring head can be moved along the object to be measured. Conventional coordinate measuring machines are stationary and require the coordinate measuring machine to be larger than the target object, which in this case is also referred to as the measurement object. Furthermore, the target object must be brought to the coordinate measuring machine.

[0004] In this regard, DE 10 2016 109 919 A1 discloses a device for measuring objects, in which a measuring head is attached to a measuring robot positioned on a mobile platform. In this way, the measuring robot can be brought to the object using the mobile platform and then measure the object. Various sensors are used to position the platform and the measuring robot.

[0005] Precise positioning of the measuring head is particularly necessary when specific parts of the object, such as individual points, need to be measured. Fast and precise positioning is also important to achieve sufficient throughput in industrial production.

[0006] US 2019 / 0 291 275 A1 discloses a method for a robot with a mobile platform, in which, based on a CAD model of the object, a plurality of waypoints are calculated based on the range of a robot arm, and a robot is moved accordingly. The robot arm is mounted on a mobile platform and has various sensors. Therefore, corresponding waypoints must be calculated.

[0007] DE 10 2018 008 209 A1 discloses another method for positioning a mobile robot in conjunction with an assembly line that moves an object to be processed. For this purpose, so-called "virtual gap points" of the object's contour are determined.

[0008] Further devices and methods for positioning robots with measuring heads are known from EP 2 869 157 A1 or US 2019 / 0 321 977 A1.

[0009] Similar problems also arise when machining objects with pinpoint accuracy. In this case, a corresponding machining tool is attached to the robot arm in addition to or as an alternative to the measuring head. Otherwise, the same applies as stated above regarding measurement.

[0010] Based on the known methods, it is an object of the present invention to provide improved devices and methods with which a faster and / or more accurate positioning of a robot arm relative to an object is possible.

[0011] According to the invention, a device according to claim 1 and a method according to claim 10 are provided. The subclaims define further embodiments.

[0012] A device for measuring, inspecting, and / or processing objects is provided, comprising a mobile platform for moving the device through a spatial area, kinematics attached to the mobile platform, and an instrument head attached to the kinematics. The kinematics are configured to move the instrument head relative to the mobile platform.

[0013] Furthermore, the device has at least one sensor. Finally, the device has a controller which is configured to determine a first estimate of a pose of a target object based on signals from the at least one sensor, to control the mobile platform based on the first estimate to move towards the object, to determine a second, more precise (i.e., more precise than the first estimate) estimate of the pose of the target object based on signals from the at least one sensor, wherein the determination of the first estimate and / or the determination of the second estimate is additionally carried out based on a digital representation of the target object, and to control the kinematics to position the instrument head on the target object based on the second estimate.

[0014] In this way, precise positioning of the instrument head on the object can be achieved. For example, in some embodiments, the positioning accuracy can be in the range of 0.1 mm.

[0015] A mobile platform is understood to be a device by means of which the device can be moved. For this purpose, the mobile platform can, for example, have wheels with a drive. In another embodiment, the mobile platform can be a rail-mounted mobile platform that moves on rails through the spatial area.

[0016] In the context of this application, the term target object refers to an object on which measurements are to be carried out and / or which is to be processed.

[0017] Kinematics refers to a movable device that can move the instrument head. An example of such a kinematic is a robot arm.

[0018] A pose is the combination of position and orientation, as defined, for example, in DIN ISO 8373, 2nd edition of December 3, 2012, under 4.5. The pose can be specified, for example, in three translation coordinates, which specify the position of an object, and three angular coordinates, which specify the orientation.

[0019] A digital representation is data that specifies the shape of the object. Such a digital representation can, for example, be a 3D model of the object created based on measurements of a sample object, and / or it can be CAD ("computer-aided design") data of the object. It should be noted that this digital representation does not have to exactly match the object. In particular, the digital representation can reflect a desired shape of the object, and in some embodiments, the device can then be used to detect deviations from this desired shape during a measurement (e.g., missing holes or deformations). The digital representation can optionally include a texture of the object, which can then additionally be used to determine the pose estimate.

[0020] In general, it should be noted that terms such as "one image" do not exclude the use of additional images or a video, i.e. a sequence of images, for the assessment.

[0021] An instrument head generally comprises one or more sensors for measuring or inspecting the target object and / or one or more tools for processing the target object.

[0022] In this context, measurement is understood to mean, in particular, a quantitative measurement (e.g., measuring a length), whereas inspection can be qualitative, for example, checking for the presence of certain features, such as drill holes, or detecting defects such as cracks without quantitatively measuring them, or performing a completeness analysis. This type of inspection can be used, in particular, for quality assurance.

[0023] The at least one sensor whose signals are used for the initial estimation can, for example, comprise a camera such as a wide-field camera. In this case, the signals from the sensor correspond to a captured image. The at least one sensor can additionally or alternatively comprise a LIDAR sensor, whereby a LIDAR sensor can also be used in combination with a camera.

[0024] The controller can further be configured to identify the target object in the image before determining the first estimate and then to determine the first estimate based on the identified object. LIDAR measurements can also be performed based on the identified object. Conventional image recognition methods can be used for this purpose. The digital representation can also be used for identification. Such identification can facilitate pose estimation.

[0025] An overview of different pose estimation methods, which can be used here, for example, for the initial estimation, can be found in T. Hodan et al., "BOP: Benchmarkň of 6d object pose estimation", European Conference on Computer Vision (ECCV) 2018.

[0026] The controller can further be configured to determine a distance of the target object from the device based on a measurement from the at least one sensor before determining the first estimate. For this purpose, the at least one sensor can comprise, for example, a LIDAR sensor, a radar sensor, or another sensor suitable for distance measurement, such as a time-of-flight (TOF) sensor. The distance determination can then be incorporated into the pose estimation, since, for example, the dimensions of the object from the digital representation reveal how large the object appears in the image, thus facilitating detection. Furthermore, the movement of the mobile platform can be determined based on the distance.

[0027] The controller can further be configured to track the pose of the target object, e.g., according to the first estimate, while moving the mobile platform and / or while moving the kinematics; however, this feature can also be omitted in other embodiments. In other words, the pose can be continuously adjusted during the movement in order to be able to correct the movement. This can be particularly helpful if the target object itself moves during the movement. Region-based tracking, in particular, can be used. Such region-based tracking is described, for example, in H. Tjaden et al., "A Region-based Gauss-Newton Approach to Real-Time Monocular Multiple Object Tracking," IEEE transactions on pattern analysis and machine intelligence 41.8 (2018), pages 1797-1812.

[0028] The controller may also include trained machine learning logic for determining the first estimate and / or the second estimate. Machine learning logic is understood to mean a device, in particular a computing device, that has been trained using machine learning methods. Examples include trained neural networks such as CNNs ("convolutional neural networks"), neural networks with a plurality of layers, including hidden layers, or supported vector machines. To train such machine learning logic, images and digital representations of example objects can be fed to the machine learning logic, with the pose of the object then being provided using coordinates (three translation coordinates and three angles) or in another way, for example, using coordinates of vertices of a cuboid or other body that circumscribes the object.These coordinates are then the result of the trained state when corresponding images and digital representations are fed into the trained machine learning logic. In particular, views of objects from different directions can be considered for training, and a large number of images, for example, approximately 10,000 images, can be used. This makes it possible to recognize the pose of an object from different directions, in the presence of occlusions, and / or under different lighting conditions.

[0029] The second estimate can be determined based on the image or another image, for example, an image of the target object taken after the platform has been moved, the digital representation, and the first estimate. The first estimate is thus used as the basis for the second estimate, allowing for a gradual refinement of the pose determination. Tracking can also take place here. Optimization methods such as Direct Directional Chamfer Optimization (D 2< CO) can be used for this purpose. In this way, an accurate pose determination is possible even for comparatively complex objects. The D 2< CO method is also suitable for shiny surfaces, for example. A description of the method can be found, for example, in M. Imperoli and A.Pretto, "D2CO: Fast and Robust Registration of 3D Textureless Objects using the Directional Chamfer Distance," Proceedings of the 10th international conference on computer vision systems (ICVS), 2015, pp. 316ff. However, other algorithms can also be used. Such algorithms can, for example, be implemented in parallel on graphics processors.

[0030] Corresponding methods for controlling a device for measuring and / or processing objects, which device comprises a mobile platform for moving the device through a spatial region and kinematics attached to the mobile platform with an instrument head attached to the kinematics and a sensor, are also provided. Such a method comprises a first estimation of a pose of a target object based on signals from a sensor, controlling the mobile platform to move towards the object, a second, more precise estimation of the pose of the target object based on signals from the at least one sensor, wherein the determination of the first estimate and / or the determination of the second estimate is also carried out based on a digital representation of the target object, and controlling the kinematics based on the second estimate in order to position the instrument head on the target object.

[0031] The invention is explained in more detail below with reference to the accompanying drawings. They show: Fig. 1 an application example of a device according to the invention, Fig. 2 a block diagram of a device according to an embodiment, Fig. 3 a flowchart of a method according to an embodiment, Fig. 4 a flowchart of a method according to a further embodiment, Fig. 5 a flowchart of a method according to a further embodiment together with explanatory figures, Fig. 6 a neural network, as may be used in some embodiments, Fig. 7A und 7B Views to explain the functioning of the neural network of the Fig. 6 , Fig. 8 a diagram illustrating pose tracking, Fig. 9 a diagram illustrating a method for more accurate pose estimation, Fig. 10 a more detailed implementation example of the procedure of Fig. 9 , and Fig. 11 a flowchart of a method for more accurately determining a pose according to an embodiment.

[0032] Various embodiments are explained in detail below. These are for illustrative purposes only and are not to be construed as limiting.

[0033] While embodiments are described with a variety of features, not all of these features (e.g., elements, components, method steps, operations, algorithms, etc.) are necessary for implementation, and other embodiments may include alternative features or additional features, or some features may be omitted. For example, specific algorithms and methods for pose determination are presented, but other algorithms and methods may be used in other embodiments.

[0034] Features of different embodiments may be combined with one another unless otherwise stated. Variations and modifications described for one embodiment are also applicable to other embodiments.

[0035] In Fig. 1 An embodiment of a device 22 for measuring objects, for example, motor vehicles 20, 21, is shown. The motor vehicles 20, 21 are merely an example of objects, and other types of objects, for example, parts of motor vehicles, parts of other devices, or the like, can also serve as target objects for measuring and / or processing. Even though a survey is used as an example here, processing or inspection of objects can also take place instead of or in addition to the survey.

[0036] The device 22 comprises a mobile platform 24 with wheels, crawler tracks, or other means of locomotion and a drive, so that the device 22 can travel through one or more rooms or even outdoors to objects 22, 21 to be measured. The mobile platform 24 represents an example of a mobile platform as can be used in exemplary embodiments. This travel can, as will be explained later, be carried out autonomously by the device 22 by means of a controller.

[0037] Furthermore, the device 22 has a measuring head 25 which is attached to a kinematics 23, for example a robot arm. By means of the kinematics 23, the measuring head 25 can be positioned precisely on the object to be measured (i.e. the respective target object), for example the motor vehicles 20 or 21. As will be explained below, in order to control the device 22, a pose of the object, i.e. its position and orientation in space, is determined and the measuring head 25 is aligned accordingly, i.e. the pose of the measuring head relative to the object is also determined. Ultimately, only the relative position of the measuring head and the object is important. Therefore, the following explanations of how the device 22 is controlled in order to position the measuring head 25 on the object 20 or 21 should always be understood as positioning relative to the object.Such a movement and positioning is also called differential movement, since it does not have to take place in an absolute, fixed coordinate system, but only positions two objects, measuring head 25 and object 20 or 21, relative to each other.

[0038] The actual measurement is then performed using the measuring head 25, also referred to as the sensor head. For this purpose, the measuring head 25 can comprise, for example, a confocal chromatic multispot sensor (CCMS), another type of optical sensor, a tactile sensor, or any other suitable sensor to perform a desired measurement on the object to be measured. If, instead of measuring, the object is to be processed, appropriate processing tools can be provided. In general, the measuring head 25 can be referred to as an instrument head, which has one or more sensors and / or one or more tools, for example for screwing, drilling, riveting, gluing, soldering, or welding.

[0039] Instead of a robot, other kinematics can also be used, such as an autonomous column measuring machine. While the mobile platform is shown here with wheels, other solutions are also possible, such as a permanently installed mobile platform that runs on rails. The latter is particularly suitable if the area where the measurements are to be taken is well-defined, for example, within a factory hall, so that movement on rails is also possible.

[0040] To explain this further, Fig. 2 a block diagram, which explains in more detail an example of the structure of a device such as device 22. The operation of the device and methods for operating the device are then explained with reference to the Fig. 3 bis 11 In the example of the Fig. 2 A device for measuring objects comprises an instrument head 30, which, for example, as explained for the measuring head 25, comprises one or more sensors for measuring the target objects, and / or one or more tools for processing target objects. The measurement and / or processing is controlled by a controller 31. For this purpose, the controller 31 can, for example, have one or more microcontrollers, microprocessors, and the like, which are programmed by a corresponding computer program to control the device and execute functions explained in more detail below. Implementation entirely or partially by application-specific hardware is also possible. It should be noted that the controller 31 of the device does not have to be fully implemented on the mobile platform 24.Rather, some of the control tasks, such as calculations, can be performed in an external computing device such as a computer and transferred to a control component on the platform via a suitable interface, such as a radio interface.

[0041] Furthermore, the device comprises the Fig. 2 a drive 35, which is used, for example, to drive the mobile platform 24 of the Fig. 1 The drive 35 is controlled by the controller 31. For this control of the drive and the control of a kinematics such as the kinematics 22 of the Fig. 1 The device can have various sensors. As an example, a LIDAR sensor 32 ("light detection and ranging") and a wide-field camera 33, for example a fisheye camera or a wide-angle camera, are shown. In addition, further sensors 34, for example acceleration sensors, angle sensors, combinations thereof, e.g., so-called inertial measurement units (IMUs), magnetometers, a thermometer for temperature compensation, further cameras, for example, a high-resolution camera with a smaller field of view than the wide-field camera 33, and the like, can be provided in the device, e.g., on the kinematics 23. A pattern-based projection camera, such as that used in gesture recognition systems, can also be used as a further sensor.In other embodiments, the sensors may include a navigation system, such as a differential GPS ("global positioning system") or the like. Odometry data, e.g., from the wheels of the mobile platform, may also be measured. Sensors 32-34 may have different update rates. Measurement data from different sensors may be used in combination, for which conventional sensor fusion methods may be applied.

[0042] The controller 31 additionally receives a digital representation of the target object, for example the objects 20, 21 of the Fig. 1 or parts thereof to be machined. Based on the digital representation 36 of the target object and data from the sensors 32 to 34, the controller 31 controls the drive 35 and the kinematics 23 in order to position the measuring head 25 or the instrument head 30 on the target object, for example the objects 20, 21 of the Fig. 1 , to position.

[0043] Various methods and techniques for this purpose, which can be implemented in the controller 31, will now be described with reference to the Fig. 3 bis 11 explained.

[0044] The Fig. 3 shows roughly a sequence of methods according to various embodiments.

[0045] In step 10, the platform is moved to the target object, ie, the object to be processed and / or measured. As explained later, this can be done based on a first, rougher, estimate of the object's pose. In step 11, an instrument head such as the instrument head 30 of the Fig. 2 aligned with the target object. This can be done, as will be explained in more detail later, based on a more precise estimate of the object's pose. The target object will be referred to simply as the object below.

[0046] Next, with reference to the Fig. 4 a more detailed procedure is explained.

[0047] In step 40, the object is detected and the distance to the object is determined. For this purpose, an image can be recorded, for example by the wide-field camera 33 of the Fig. 3 , and in particular for distance measurement of the LIDAR sensor 32 of the Fig. 2 Machine learning techniques such as YOLO ("you only look once") methods or R-CNN, Fast R-CNN, and Faster R-CNN methods can be used for object detection. With such approaches, the presence of the target object can be determined and its position can be determined with an accuracy of, for example, 100 mm, depending on the object type and size. These and other numerical values given here are for illustrative purposes only and should not be interpreted as limiting.

[0048] In step 41, the pose of the object is then estimated, which will be referred to as a rough estimate below to distinguish it from the fine estimate that occurs later in step 43. The term "estimate" or "estimate" expresses that the determination is subject to a certain degree of uncertainty, which is greater in the rough estimate of step 41 than in the fine estimate of step 43. A camera image can be used for this purpose, whereby the camera can be attached to both the mobile platform and the instrument head. The camera can be the same camera as in step 40, or a separate camera can be used for this purpose, which, for example, has a smaller field of view and is aimed at the object detected in step 40. In addition, the digital representation of the object, from which the object's shape can be seen, is used here.In some embodiments, the object is additionally segmented, for example, by determining the contour of the object in the image. Pose estimation and segmentation, if applicable, can be performed using machine learning techniques, such as a so-called "single shot approach," PVNet, or dense fusion method. Such methods are described, for example, in S. Peng et al., "PVNet: Pixel-wise Voting Network for 6DoF Pose Estimation," arXiv 1812.11788, December 2018; C. Wang et al., "DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion," arXiv 1901.04780, January 2019; or K. Kleeberger and M. Huber, "Single Shot 6D Object Pose Estimation," arXiv 2004.12729, April 2020.

[0049] Using such methods, the object's pose can be determined with an accuracy in the range of ±10 mm to ±50 mm for object diameters of approximately 1 m. The achievable accuracy may depend on the object size.

[0050] The pose determined in this way can be tracked, in particular in real time, in order to take changes in the object's pose into account. Here, area-based object tracking can be used, which, for example, uses local color histograms of image recordings that are evaluated over time and optimized for pose tracking using a Gauss-Newton method. The tracking update rate can be adapted to the movement speed of the mobile platform relative to the target object. The aforementioned method allows up to approximately 20 recalculations per second and thus enables the tracking of fast movements. The update rate can be selected in such a way that the object's position does not change too greatly between successive images, i.e. the position change is only large enough for the method to ensure reliable tracking.

[0051] In step 42, the platform is then moved to the object. In addition to the rough pose estimate from step 41, LIDAR data from the LIDAR sensor 32 of the Fig. 2 SLAM ("simultaneous localization and mapping") techniques can be used to process the LIDAR data to enable differential movement toward the object while avoiding obstacles. Robot Operating System (ROS) programs such as ROS Navigation Stack or ROS Differential_Drive can be used for this purpose. In some embodiments, such techniques allow the platform to be positioned with a positioning accuracy of ±10 mm relative to the object in a platform's motion plane.

[0052] In step 43, the aforementioned fine estimate of the object's pose is then performed. This can start from the rough estimate from step 41 and further refine it. Sensor data, such as camera data, can be used for this purpose. Preferably, a high-resolution camera is used, for example, in the instrument head 30 of the Fig. 3 can be provided. In addition, the digital representation of the object can be incorporated into the calculation. In step 43, methods such as edge detection-based methods can be used. One example is the D 2< CO method ("direct directional chamber optimization"), which can deliver robust results even with texture-free and glossy object surfaces.

[0053] In step 44, based on the object pose estimated in step 43, the instrument head is aligned with the object by controlling a kinematic system such as kinematic system 23. This allows a pose determination accuracy of up to ±0.1 mm in some embodiments. Positioning accuracy can then also be within this range if the kinematic system used also allows at least such an accuracy.

[0054] In step 45, the measurement and / or processing of the object is then performed using the positioned instrument head and the sensors and / or tools attached to it. For example, gaps in an object can be measured and / or it can be detected whether various parts of the object are flush.

[0055] An example of the rough estimation of step 41 and a possible follow-up of the pose will now be described with reference to the Fig. 5 , 6, 7A und 7B explained.

[0056] In step 50 of the Fig. 5 As symbolized by images 56, a plurality of image recordings of an object are generated from different directions. This plurality of image data from different directions can be viewed as a digital representation of the object. In other embodiments, such images can be generated virtually from an existing digital representation, e.g., from 3D CAD data. One such approach is described, for example, in S. Thalhammer et al., "SyDPose: Object Detection and Pose Estimation in Cluttered Real-World Depth Images using only Synthetic Data," 2019 International Conference on 3D Vision. The recorded object in images 56 is an example of an object to be measured later in production, for example, a prototype or an object from series production.The images are provided with the object's pose relative to a viewer position (the camera position in the case of real images, a virtual observer position in the case of virtual generation). This information can also be generated during virtual generation.

[0057] In step 51, auxiliary points are automatically generated in the images generated in step 50, which characterize the pose of the object, in particular for the method described here. In an image 57, these can be, for example, a plurality of 2D points 58 on the respective image of the object, which characterize a pose of the object. These auxiliary points can be, for example, corner points and the center point of a bounding box surrounding the object in the image (shown as a rectangle in image 57) or points on the surface of the object, for example on edges that can be recognized in the image by image recognition. These auxiliary points characterize the pose of the object for the method.

[0058] Based on this data, a neural network can then be trained in step 52 and, after training, used to determine the pose. For example, the poses of the images generated in step 50 are correlated with the auxiliary points from step 51, so that after training, the corresponding auxiliary points and thus the pose can be determined for essentially any image of the object. Such a neural network is described as an example in Fig. 6 shown. The neural network receives an image 60 as input in an input layer 61. From an output layer 65, for example, the auxiliary points that describe the object are then output (see the explanations for Figure 57 and step 51). Processing takes place via a plurality of intermediate layers 62A to 62H, whereby branches can also occur (see layers 63, 64). These layers 62A to 62H, 63, 64 are referred to as hidden layers of the neural network. The dimensionality of the problem can be changed in the layers. The neural network shown can be designed as a CNN ("convolutional neural network") and operate, for example, according to the YOLO method. The result is obtained, as shown in the Fig. 7A und 7B shown, a 3D bounding box represented around the object 71 as a cuboid 73 in the image 70 with the estimated vertices of the cuboid 73, which represents the pose of the object. For illustration, the 3D position is backprojected onto the 2D image.

[0059] Fig. 7A und B show an example of an output of the trained network. Fig. 7A the 2D points in the image are visualized, which according to Fig. 7B have the corresponding 3D representation - here according to an example where the auxiliary points are defined to reflect the vertices of the 3D bounding box corresponding to cuboid 73.

[0060] From a captured image such as image 59 of the Fig. 5 In step 53, the trained neural network, to which the image is fed, then produces a rough estimate of the pose in the six coordinates (6D), i.e., three translational coordinates and three angular coordinates. As shown schematically as an example, the estimated pose, as symbolized by a dashed line 512, already approximates the actual pose of the object relatively well, but is still subject to errors.

[0061] As already explained, a region-based tracking of the pose can then follow in step 41. For this purpose, areas around the contour of the object's pose can be worked out, as indicated in an image 510 by circles 511, and foreground and background can be distinguished here. This is shown in Fig. 8 explained in more detail.

[0062] A block 80 of the Fig. 8 symbolizes the rough estimation of the pose, where, for example, the digital representation is adapted to the image by a transformation T, and thus the transformation T (translation and rotation) represents the estimated pose. In a block 81, areas are then evaluated. The image 82 corresponds to the image 510 of the Fig. 5 , where the individual circles represent individual areas. As indicated in Figure 83, histograms can then be created for the individual circles in the transition area between the foreground and background, thus allowing the contour of the object to be better determined. This can be done with an optimization as expressed in block 84. The cost function to be optimized is a function E rbphm, which consists of three components. Here, E rb represents a cost function of area-based pose detection as shown in block 81, with a weighting factor λ rb , E ph represents a cost function of photometric pose detection with a weighting factor λ ph , and E m represents a cost function of a motion component, with a weighting factor λ m . The weighting factors are set in order to weight the cost functions relative to one another.

[0063] Photometric pose recognition relies on the assumption that the colors of any surface point of the object are similar in every image, regardless of perspective (this is called photometric consistency). The cost function represents the pixel-wise photometric error as a function of the pose parameters (translation and orientation), which are optimized.

[0064] The motion component represents information about the movement of a camera used for image acquisition, for example based on sensors such as inertial measurement units (IMUs) or other motion sensors of the device, for example the mobile platform or kinematics.

[0065] This results in an updated 6D pose in block 85. In addition, the digital representation, referred to here as the model, can also be updated.

[0066] Again on Fig. 5 Referring to this, in step 55, a tracked 6D pose results, as shown in an image 513. Even though the object has moved relative to the image acquisition compared to image 512, the pose, represented here by a contour 514, matches the object even better than in image 59.

[0067] Next, approaches for the fine-tuning estimation are presented in step 43 of the Fig. 4 explained in more detail. Fig. 9 shows an approach for fine estimation according to an embodiment.

[0068] The inputs used here are a high-resolution input image 90, in which a target object 91, in this case a car door, can be seen. This input image 90 is recorded with a camera of the device. Furthermore, the rough estimate of the pose, represented here by a contour 93, which was obtained in step 41, is used as input information, as shown in an image 92. Finally, a digital representation 94 is used. This input information is fed to a pose matching algorithm 95, examples of which will be explained in more detail later. This then results in the fine estimate of the pose, represented here by a contour 97, as shown in an image 96.

[0069] The Fig. 10 shows a possible implementation of the approach of Fig. 9 using the aforementioned D 2< CO method. A high-resolution input image is processed by edge analysis in different directions to a directional edge distance tensor (in English "directional chamber distance tensor"), represented by image 1001, which is a three-dimensional tensor with dimensions width times height of the image times the number of evaluated directions. This tensor, the initial pose specified by a translation vector T with components tx, ty and tz and an orientation vector Ω with orientation coordinates rx, ry, rz and a model of the object 1003, which can be represented, for example, as a cloud of n 3D points P i are fed to an optimization process 1004, which optimizes the vector components of the pose, ie rx , ry , rz , tx , ty , tz. The optimization process is the one described in Fig. 10 In the case presented, the so-called "Powell's dog leg" method is used.

[0070] Alternatively, another method, such as the Levenberg-Marquardt method, can be used. The result is a refined pose 1005 with correspondingly refined vectors T and Ω, each of which is marked with an apostrophe (') to distinguish it from the initial pose 1002.

[0071] The Fig. 11 shows a flow diagram of the application of the D 2< CO process, as it is also in Fig. 10 shown, to the fine estimation of the object's pose. In step 1101, a high-resolution input image, an initial pose from the coarse estimation, and a digital representation, for example a 3D model, are provided as input values, as already described in the Fig. 9 and 10 Then the tensor 1001 of the Fig. 10 In step 1102, edges are detected in the input image. Conventional edge detection methods can be used for this purpose.

[0072] In step 1103, direction extraction is then performed, ie different directions in which the model extends are recognized.

[0073] Then, a loop 1104 is executed, which runs over all directions, where N dir is the number of directions extracted in step 1103. For each pass, one of the directions is selected in step 1105 using a Harris transform. A direction transformation then follows in step 1106. This creates a 2D Euclidean direction transformation for each combination of direction and detected edges, also called a direction-edge map. The maps are then stacked to form a Directional Chamfer Direction (DCD) tensor, as represented in Figure 1001.

[0074] This is followed by a forward / backward propagation in step 1107. In this step, the DCD tensor is calculated based on the 2D transformations from step 1106 and the directions of the detected edges. Finally, in step 1108, the tensor is smoothed according to its orientation, for which a simple Gaussian filter can be used. This tensor is then fed to a nonlinear optimization method 1109, corresponding to the optimization 1004 of the Fig. 10 As already mentioned, this is done according to the "Powell's dog leg" method or another method such as the Levenberg-Marquardt method.

[0075] The optimization involves an optimization loop 1110 as long as the pose does not converge sufficiently (for example, as long as the components of the vectors T and Ω still change by more than a specified amount or a specified relative size per iteration). In each iteration of the loop, the digital representation, for example a 3D model, is rasterized in step 1111. In the process, n sample points are extracted.

[0076] In step 112, the points of the 3D model are reprojected. For this purpose, the example points selected in step 1111 are projected onto edges in the 2D image space using the current pose T, Ω, and calibration data of a camera used, and their directions are calculated. In step 1113, Jacobian matrices with respect to T and Ω are calculated. In step 1114, a corresponding cost function is then evaluated, based on which the pose is updated in step 1115. As a result, the fine estimate of the pose is obtained in step 1116.

[0077] Since the D 2< CO method is a well-known method in itself, the above explanation is kept relatively brief. As already mentioned, other methods can also be used, or the D 2< CO method can be combined with other methods, such as edge weighting or so-called annealing (randomized delta transformations) or systematic offsets, which can increase robustness. Other methods include oriented chamfer matching as described in H. Kai et al., "Fast Detection of Multiple Textureless 3-D Objects," Proceedings of the 9th International Conference on Computer Vision Systems (ICVS), 2013; fast directional chamfer matching as described in M.-Y. Liu et al., "Fast Object Localization and Pose Estimation in Heavy Clutter for Robotic BinPicking," Mitsubishi Electric Research Laboratories TR2012-00, 2012; or point-pair feature matching.The latter approach involves extracting 3D feature points from depth images using a specific strategy or randomly. These observed feature points are then efficiently mapped to the digital representation of the object (the visible surface of the CAD model). Hash tables can be used to map point pairs to accelerate computation at runtime. Additionally, this approach utilizes various ICP (Iterative Closest Point) algorithms for pose registration, such as "Picky-ICP" or "BC-ICP" (ICP with biunique correspondences). One such method is described in B. Drost et al., "Model Globally, Match Locally: Efficient and Robust 3D Object Recognition," published at http: / / campar.in.tum.de / pub / drost2010CVPR / drost2010CVPR.pdf.

[0078] It is again pointed out that the above embodiments, in particular preliminary details of the methods used, are for illustrative purposes only and that other methods can also be used.

Claims

1. Apparatus (22) for measuring, inspecting and / or processing objects, comprising: a mobile platform (24) for moving the apparatus through a spatial region, a kinematic system (23) attached to the mobile platform (24), and an instrument head (25; 30) attached to the kinematic system (23), wherein the kinematic system (23) is configured to move the instrument head (25; 30) relative to the mobile platform (24), at least one sensor (32-34) attached to the mobile platform (24) or the kinematic system (23) and a high-resolution camera attached to the mobile platform (24) or the kinematic system (23), and a controller (31) configured: to determine a first estimation of a pose of a target object (20, 21; 91) on the basis of signals from the at least one sensor (32-34), to control the mobile platform (24) on the basis of the first estimation to move to the target object, to determine a second, more accurate estimation of the pose of the target object on the basis of signals from the high-resolution camera and the first estimation and to control the kinematic system (23), wherein determining the first estimation and determining the second estimation are additionally effected on the basis of a digital representation (36; 94; 1003) of the target object (20, 21; 91), and to position the instrument head at the target object (20, 21; 91) on the basis of the second estimation by way of the control of the kinematic system, and wherein the controller (31) is further configured to track the pose of the target object (20, 21; 91) during the movement of the mobile platform and / or during a movement of the kinematic system (23).

2. Apparatus according to Claim 1, wherein the controller (31) is configured to identify the target object (20, 21; 91) in an image before determining the first estimation and to determine the first estimation on the basis of the identified target object.

3. Apparatus according to Claim 1 or 2, wherein the controller (31) is configured to determine a distance between the target object (20, 21; 91) and the apparatus (22) on the basis of a measurement of the at least one sensor (32-34) before determining the first estimation.

4. Apparatus according to any of Claims 1 to 3, wherein the at least one sensor comprises a LIDAR sensor and / or a camera.

5. Apparatus according to any of Claims 1 to 4, wherein the controller (31) is configured to track the pose by means of region-based tracking.

6. Apparatus according to any of Claims 1 to 5, wherein the controller (31) comprises a trained machine learning logic for determining the first estimation and / or the second estimation.

7. Apparatus according to any of Claims 1 to 6, wherein determining the second estimation is effected on the basis of the image or a further image of the target object (20, 21; 91), the digital representation (36; 94; 1003) and the first estimation.

8. Apparatus according to Claim 7, wherein determining the second estimation is effected on the basis of a directional chamfer optimization.

9. Method for controlling an apparatus (22) for measuring, inspecting and / or processing objects, wherein the apparatus comprises a mobile platform (24) for moving the apparatus through a spatial region, a kinematic system (23) attached to the mobile platform (24), and an instrument head (25; 30) attached to the kinematic system (23), wherein the kinematic system (23) is configured to move the instrument head (25; 30) relative to the mobile platform (24), and wherein the apparatus comprises at least one sensor (32-34) attached to the mobile platform (23) or the kinematic system (23) and a high-resolution camera attached to the mobile platform (23) or the kinematic system (23), wherein the method comprises: determining a first estimation of a pose of a target object (20, 21; 91) on the basis of signals from the at least one sensor (32-34), controlling the mobile platform (24) on the basis of the first estimation to move to the target object, determining a second, more accurate estimation of the pose of the target object on the basis of signals from the high-resolution camera and the first estimation, wherein determining the first estimation and determining the second estimation are additionally effected on the basis of a digital representation (36; 94; 1003) of the target object (20, 21; 91), and controlling the kinematic system (23) in order to position the instrument head at the target object (20, 21; 91) on the basis of the second estimation, and tracking the pose of the target object (20, 21; 91) during the movement of the mobile platform and / or during a movement of the kinematic system (23).

10. Method according to Claim 9, further comprising, before determining the first estimation, identifying the target object (20, 21; 91) in an image, wherein the first estimation is determined on the basis of the identified target object.

11. Method according to Claim 9 or 10, further comprising, before determining the first estimation, determining a distance between the target object (20, 21; 91) and the apparatus (22) on the basis of a measurement of the at least one sensor (32-34).

12. Method according to any of Claims 9 to 11, wherein the tracking comprises region-based tracking.

13. Method according to any of Claims 9 to 12, wherein determining the first estimation comprises using a machine learning logic.

14. Method according to any of Claims 9 to 13, wherein determining the second estimation is effected on the basis of the image or a further image of the target object (20, 21; 91), the digital representation and the first estimation.

15. Method according to Claim 14, wherein determining the second estimation is effected on the basis of a directional chamfer optimization.