Devices and methods for training machine learning models that recognize object topology from object images

By generating training data image pairs and using the Laplace-Beltrami operator to determine edge weights, the problem of sensor calibration dependence in robot object recognition topology technology is solved, achieving higher accuracy and robust object recognition.

CN114494312BActive Publication Date: 2026-03-10ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing robot object recognition topology technologies rely on sensor calibration accuracy and self-supervised learning, resulting in poor interpretability and sensitivity to calibration errors, making it difficult to accurately identify objects in different postures.

Method used

By generating training data image pairs, the descriptor component values ​​of the mesh vertices of the 3D model are used to adapt or add descriptor component values ​​to enhance the robustness of the machine learning model. The model is trained using supervised learning, and edge weights are determined by combining the Laplace-Beltrami operator to generate target images to reduce the impact of calibration errors.

Benefits of technology

This improved the robot's object recognition accuracy and model robustness under different postures, reduced the dependence on sensor calibration, and enhanced the interpretability of the model and the effectiveness of the training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494312B_ABST
    Figure CN114494312B_ABST
Patent Text Reader

Abstract

A method for training a machine learning model for recognizing objects from an object image, according to various embodiments, includes: obtaining a 3D model of the object; determining descriptor component values ​​for each vertex of a mesh; and generating training data image pairs having a training input image and a target image, respectively. The target image is generated by determining vertex positions in the training input image; assigning descriptor component values ​​determined for vertices at vertex positions to positions in the target image; and adapting at least some of the descriptor component values ​​assigned to positions in the target image or adding descriptor component values ​​to positions in the target image such that descriptor component values ​​within a predetermined spacing within the object are closer to descriptor component values ​​assigned to positions outside the object than descriptor component values ​​assigned to positions inside the object, wherein the positions inside the object are farther from object edges than the predetermined spacing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to a device and a method for training a machine learning model for recognizing a topology of an object from an image of the object. BACKGROUND

[0002] In order to realize the flexible manufacturing or processing of objects by means of robots, it is desirable that the robot is able to manipulate the object irrespective of the pose in which the object is arranged in the workspace of the robot. Thus, the robot should be able to recognize which part of the object is in which position, so that it can, for example, grasp the object at the correct site in order to fix the object, for example, at another object or to weld the object at the current site. This means that the robot should be able to recognize the topology (surface) of the object from one or more images recorded by a camera fixed at the robot. A solution for realizing this consists in determining for a part of the object, i.e. a pixel representing the object in the image plane, a descriptor, i.e. a point (vector), in a predefined descriptor space, wherein the robot is trained to assign the same descriptor to the same part of the object irrespective of the current pose of the object.

[0003] In the publication "Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation" by Peter Florence et al., hereinafter referred to as reference 1, a dense object net is described, which is a model of self-supervised dense descriptor learning.

[0004] However, the effectiveness of the solution of reference 1 depends to a large extent on the quality of the training data collected and the sensors involved, and the descriptors determined from the self-supervised learning often lack interpretability. In response thereto, it is desirable to improve the solution for training for recognizing the topology of an object in this respect. SUMMARY

[0005] According to various embodiments, a method for training a machine learning model for recognizing an object topology of an object from an object image is provided, the method comprising: obtaining a 3D model of the object, wherein the 3D model comprises a mesh of vertices; determining descriptor component values for each vertex of the mesh; generating training data image pairs, wherein each training data image pair comprises a training input image and a target image showing the object, and wherein generating the target image comprises: determining vertex positions of vertices of an object model of the object, the vertices having the vertex positions in the training input image; assigning, for each determined vertex position in the training input image, a descriptor component value determined for the vertex at the vertex position to a position in the target image; and adapting at least some of the descriptor component values assigned to positions in the target image or adding descriptor component values to positions of the target image, such that the adapting or adding assigns descriptor component values to positions within the object and within a preset distance from an object edge in the target image that are closer to descriptor component values assigned to positions outside the object than to descriptor component values assigned to positions in the interior of the object, wherein positions in the interior of the object are further from the object edge than the preset distance. The method further comprises training the machine learning model by supervised learning using the training data image pairs as training data.

[0006] By adapting the descriptors in the interior of the object near the edge, i.e. within the preset distance, towards the direction of the descriptors as they appear in the background, i.e. outside the object, it is achieved that the object edge in the target image is less sharp, in turn making the training robust with respect to errors in the training data, especially calibration errors of the camera used to record the training input image. Thereby, it is avoided that the training is negatively affected in case the object position in the target image does not exactly coincide with the object position in the recorded camera image that is used as the pertaining training input image.

[0007] Various examples are explained below.

[0008] Example 1 is the method for training a machine learning model for recognizing an object topology of an object from an object image as described above.

[0009] Example 2 is the method of example 1, having the descriptor for each vertex of the mesh, wherein the descriptor is a vector having a vector component for each of a plurality of channels, and

[0010] adapting the descriptor component values of the channel such that the adapting assigns descriptor component values of the channel to positions within the object and within a preset distance from an object edge in the target image that are closer to descriptor component values of the channel assigned to positions outside the object than to descriptor component values of the channel assigned to positions in the interior of the object, wherein positions in the interior of the object are further from the object edge than the preset distance; and / or

[0011] adding a channel with descriptor component values such that the descriptor component values of the added channel assigned to positions within the object and within a preset distance from the object edge in the target image are closer to the descriptor component values of the added channel assigned to positions outside the object than to the descriptor component values of the added channel assigned to positions within the object and further away from the object edge than the preset distance.

[0012] Thus, the target image can be an image with a plurality of channels comprising descriptor component values and one or more channels can be adapted or a channel can be added in order to achieve that the training of the machine learning model is more robust with respect to calibration errors.

[0013] Example 3 is the method of example 2, wherein one channel of the plurality of channels is replaced by the added channel.

[0014] By replacing one channel by adding one channel, it is achieved that the number of channels remains unchanged. Thereby, by adding no training data volume is increased and the dimension of the output of the machine learning model can remain unchanged.

[0015] Example 4 is the method of example 2 or 3, wherein the descriptor component values of the added channel are chosen such that the descriptor component values monotonically vary when traversing the positions from positions outside the object to positions in the object center.

[0016] In said example, intuitively a mask is created, wherein the descriptor component values are more strongly masked the closer they are to the background, since the risk of the descriptor component values negatively influencing the training of the machine learning model due to calibration errors increases the closer to the object edge (and thus the closer to the background). Such a mask can be simply created using the 3D model (and the known positions of the object in the target image) and thus achieved with low additional effort to improve the training.

[0017] Example 5 is the method of any one of examples 1 to 4, wherein the descriptor component values are adapted such that the descriptor component values assigned to positions within the object and within a preset distance and the descriptor component values assigned to positions outside the object are balanced.

[0018] Intuitively, in this example, the descriptor component values at the object edge blend with the descriptor component values of the background in order to avoid that calibration errors negatively influence the training of the machine learning model. This can be done by simply automatically post-processing the target image and thus achieved with low additional effort to improve the training.

[0019] Example 6 is the method of any of examples 1 to 5, wherein generating the training data image pair comprises obtaining a plurality of images of the object having different poses, and generating the training data image pair from each of the obtained images by generating a target image for the obtained image.

[0020] This enables training of a machine learning model (e.g., of a robot having a robot control device implementing the machine learning model) so as to recognize a topology of an object irrespective of a pose of the object, e.g., in a workspace of the robot.

[0021] Example 7 is the method of any of examples 1 or 6, comprising determining, from respective poses of the object in the training input image (e.g., in a camera coordinate system), vertex positions of vertices of an object model of the object, the vertices having the vertex positions in the training input image.

[0022] This enables accurate determination of the vertex positions, which in turn enables accurate target images for supervised training.

[0023] Example 8 is the method of any of examples 1 to 7, wherein the vertices of the 3D model are connected by edges, wherein each edge has a weight specifying a proximity of two vertices of the object connected by the edge, and wherein the descriptor of each vertex of the mesh is determined by finding a vertex of the pair of descriptors that minimizes a distance between the descriptors of the pair of descriptors with respect to the connected vertices in a manner weighted by the weights of the edges between the pairs of vertices.

[0024] This approach enables training of a machine learning model (such as a neural network) so as to perform more accurate predictions (i.e., descriptor determinations) compared to using self-supervised learning (i.e., enabling a more widespread use of the network). Furthermore, it provides greater flexibility for adapting the machine learning model so that it can be used in different problems, reduces training data requirements (e.g., the amount of training data required) and results in an interpretable machine learning aid.

[0025] Example 9 is the method of example 8, wherein finding the descriptor comprises determining an eigenvector of a Laplacian matrix of a graph formed by the vertices and edges of the 3D model and adopting components of the eigenvector as the descriptor component values.

[0026] This enables efficient determination of (nearly) optimal descriptors. For example, each descriptor is a vector having a dimension d and in turn having d components (e.g., (3, 4) is a vector having a dimension 2 and having components 3 and 4).

[0027] Example 10 is the method of example 9, comprising associating each vertex with a component position in the eigenvector, and for each vertex adopting a component in the eigenvector at the component position associated with the vertex as the descriptor component value for the vertex.

[0028] In particular, the descriptor space dimension can be flexibly chosen by selecting multiple feature vector components for the descriptor, depending on the desired descriptor space dimension.

[0029] Example 11 is the method of example 10, wherein determining the descriptor further comprises compressing components of feature vectors whose feature values differ by less than a pre-set threshold into a single component.

[0030] This achieves that the descriptor space dimension is reduced when the object is symmetric. In other words, an unnecessary distinction between symmetric parts of the object can be avoided.

[0031] Example 12 is the method of any of examples 1 to 11, wherein obtaining the 3D model of the object comprises obtaining a 3D mesh of vertices and edges modeling the object, and assigning edge weights of a Laplace-Beltrami operator as weights of the edges, the Laplace-Beltrami operator being used at the mesh.

[0032] Using this approach achieves that the geometry of the model is taken into account when determining the proximity of two vertices by employing geodesic distances on the object instead of Euclidean metric in the ambient space, which can be inaccurate when the object is curved.

[0033] Example 13 is a method for controlling a robot, comprising training a machine learning model according to any of examples 1 to 12, obtaining an image showing an object, inputting the image into the machine learning model, determining a pose of the object from an output of the machine learning model, and controlling the robot according to the determined pose of the object.

[0034] Example 14 is the method of example 13, wherein determining the pose of the object comprises determining a position of a certain part of the object, and wherein controlling the robot according to the determined pose of the object comprises controlling an end effector of the robot to move to the position of the part of the object and to interact with the part of the object.

[0035] Example 15 is a software or hardware agent, in particular a robot, comprising a camera designed for providing image data of an object, a control device designed for implementing a machine learning model, and a training device designed for training the machine learning model by the method of any of examples 1 to 12.

[0036] Example 16 is the software or hardware agent according to example 15, comprising at least one effector, wherein the control device is designed for controlling the at least one effector with the output of the machine learning model.

[0037] Example 17 is a computer program comprising instructions which, when executed by a processor, cause the processor to perform the method according to any of examples 1 to 14.

[0038] Example 18 is a computer readable medium storing instructions that, when executed by a processor, cause the processor to perform a method according to any of examples 1 to 14. BRIEF DESCRIPTION OF DRAWINGS

[0039] In the drawings, like reference numerals refer to like parts throughout the various views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the application. In the following description, various aspects of the application are described with reference to the following drawings, in which:

[0040] Figure 1 a robot is shown,

[0041] Figure 2 training of a neural network according to one embodiment is explained,

[0042] Figure 3 an exemplary embedding of a 4-node graph for determining a descriptor is shown,

[0043] Figure 4 a definition of an angle at a vertex of a 3D model for determining an edge weight according to a Laplace-Beltrami operator is explained,

[0044] Figure 5 a geometric training target image generated according to one embodiment is shown,

[0045] Figure 6 a view-dependent training target image generated for a training target image according to a first variant form is shown, Figure 5

[0046] Figure 7 a channel of a view-dependent training target image generated for a geometric target image according to a second variant form is shown, Figure 5

[0047] Figure 8 a method for training a machine learning model for recognizing an object topology from an object image according to one embodiment is shown.

[0048] The following detailed description relates to the accompanying drawings, which illustrate specific details and aspects of the present disclosure, in which embodiments of the application can be practiced. Other aspects and features can be used, and structural, logical and electrical changes can be made without departing from the scope of the present application. The various aspects of the present disclosure are not necessarily mutually exclusive, as some aspects of the present disclosure can be combined with one or more other aspects of the present disclosure to form new aspects. DETAILED DESCRIPTION

[0049] ​​Various examples are described in more detail below.

[0050] Figure 1 A robot 100 is shown.

[0051] The robot 100 comprises a robot arm 101, e.g. an industrial robot arm for operating or assembling workpieces (or one or more other objects). The robot arm 101 comprises manipulators 102, 103, 104 and a base (or support) 105 by means of which the manipulators 102, 103, 104 are supported. The expression "manipulator" relates to a movable member of the robot arm 101, the operation of which effects a physical interaction with the environment in order to, for example, perform a task. For the purpose of control, the robot 100 comprises a (robot) control device 106, which is designed to perform the interaction with the environment in accordance with a control program. The last member 104 of the manipulators 102, 103, 104 (which is furthest from the support 105) is also referred to as an end effector 104 and can comprise one or more tools, such as a welding torch, a clamping equipment, a paint spraying device, etc.

[0052] The other manipulators 102, 103 (which are closer to the support 105) can form positioning devices, such that together with the end effector 104, at the end of the robot arm 101, a robot arm 101 with the end effector 104 is provided. The robot arm 101 is a machine arm which can provide similar to human arm functionality (possibly by means of a tool at its end).

[0053] The robot arm 101 can comprise joint elements 107, 108, 109 which connect the manipulators 102, 103, 104 to each other and to the support 105. The joint elements 107, 108, 109 can have one or more joints which can provide a rotatable movement (i.e. a rotational movement) and / or a translational movement (i.e. a displacement) of the associated manipulators relative to each other, respectively. The movement of the manipulators 102, 103, 104 can be initiated by means of actuators which are controlled by the control device 106.

[0054] The expression "actuator" can be understood as a component which is configured to react to a drive thereof to produce a mechanism or a process. The actuator can implement an instruction (a so-called activation) created by the control device 106 as a mechanical movement. The actuator, e.g. an electromechanical transducer, can be designed to transform electrical energy into mechanical energy as a reaction to a drive thereof.

[0055] The expression "control device" can be understood as any kind of entity implementing logic, which for example can comprise a circuit and / or a processor capable of executing software, firmware or a combination thereof stored in a storage medium and can issue instructions, for example to the actuators in the present example. For example, the control device can be configured by program code, for example software, in order to control the functioning of the system, in the present example the robot.

[0056] In the present example, the control device 106 comprises one or more processors 110 and a memory 111 storing code and data, wherein the processors 110 control the robot arm 101 based on the code and data. According to various embodiments, the control device 106 controls the robot arm 101 based on a machine learning model 112 stored in the memory 111.

[0057] According to various embodiments, the machine learning model 112 is designed and trained for the robot 100 to implement recognition of an object 113, for example placed in the workspace of the robot arm 101. For example, the robot can decide how to handle the object 113 depending on what the object is, for example the object type, or can also recognize and decide which part of the object 113 should be grasped by the end effector 109. The robot 100 can for example be equipped with one or more cameras 114, which implement for it recording images of its workspace. For example, the cameras 114 are fixed at the robot arm 101, so that the robot can record images of the object 113 from different angles by moving the robot arm 101 around it.

[0058] An example for a machine learning model 112 for object recognition is a dense object network. The dense object network maps an image, for example an RGB image provided by the camera 114, to a descriptor space image of any dimensionality (dimensionality D), as it is described in reference 1.

[0059] The dense object network is a neural network trained with self-supervised learning, which outputs a descriptor space image for an input image of an image. However, the effectiveness of the approach depends to a large extent on the quality of the training data collected and the sensors involved, for example the camera 114, in particular its calibration accuracy. Furthermore, the interpretation of the network predictions can be difficult.

[0060] According to various embodiments, a scheme for recognizing objects and their poses is provided, which assumes that a 3D model of the object, for example a CAD (Computer Aided Design) model, is known, which is a typical case for industrial assembly or machining tasks. An RGBD image (RGB + depth information) of the object can also be recorded and from it the 3D model of the object can be determined.

[0061] According to various embodiments, a non-linear dimensionality reduction technique is used to compute an optimal target image for an input image used to train a neural network. Thus, instead of using self-supervised training of a neural network, supervised training of a neural network is used.

[0062] According to one embodiment, for generating training data for training a machine learning model 112, first data collection is performed. In particular, registered RGB (red-green-blue) images are collected, for example. Here, a registered image denotes an RGB image with known intrinsic and extrinsic camera values. For example, in a real-world scenario, during a robot (e.g., a robot arm 101) loop-around movement, a camera 114 fixed at the robot (e.g., a camera fixed at the robot wrist) is used to scan an object, for example. Other extrinsic estimation techniques, such as ChArUco markers, can be used, i.e., an object can be placed in different positions and poses with respect to a ChArUco board and images of this arrangement (ChArUco board and object) are recorded. In a simulated scenario, RGB images generated with a photorealistic rendering using known object poses are used.

[0063] After collecting the RGB images, the neural network is trained supervisedly to render target images of the RGB images.

[0064] Assumption: The pose of each object in world coordinates is known in each collected RGB image. This is not complex for a simulated scenario, but requires manual coordination for a scene in the real world, e.g., placing objects at predefined positions. RGBD images can also be used to determine the position of an object.

[0065] With the information and in case of using the vertex descriptor computation technique as described in the following for each RGB image (i.e., training input image), a descriptor image (i.e., training output image, also called target image or ground truth image) is rendered.

[0066] When a target image is generated for each RGB image, i.e., a pair of RGB image and target image is formed, this pair of training input image and related target image can be used as training data for training a neural network, as explained in Figure 2

[0067] Figure 2 The training of a neural network 200 according to one embodiment is explained.

[0068] The neural network 200 is a fully convolutional network mapping h x w x 3 tensors (input images) to h x w x D tensors (output images).

[0069] ​The neural network comprises multiple levels of convolutional layers 204, followed by pooling layers, up-sampling layers 205 and skip connections 206 to combine the outputs of different layers.

[0070] For training, the neural network 200 receives a training input image 201 and outputs an output image 202 with pixel values in a descriptor space (e.g. in terms of color components according to descriptor vector components). A training loss is computed between the output image 202 and a target image 203 associated with the training input image. This can be done for a stack of training input images, and the training loss can be averaged over the training input images, and the weights of the neural network 200 are trained using stochastic gradient descent with the training loss. The training loss computed between the output image 202 and the target image 203 is e.g. an L2 loss function (to minimize the least squares error pixel-wise between the target image 203 and the output image 202).

[0071] The training input image 201 shows an object, and the target image as well as the output image contain vectors in a descriptor space. The vectors in the descriptor space can be mapped onto colors, such that the output image 202 (as well as the target image 203) resembles a heat map of the object.

[0072] The vectors in the descriptor space (also referred to as (dense) descriptors) are d-dimensional vectors (e.g. d is 1, 2 or 3) assigned to each pixel in the respective image (e.g. to each pixel of the input image 201, in case the input image 201 and the output image 202 have the same dimensions). The dense descriptors implicitly encode the surface topology of the object shown in the input image 201, which is invariant with respect to the object pose or camera position.

[0073] If a 3D model of a given object is given, an optimal (in the Riemannian sense) and unambiguous descriptor vector for each vertex of the object 3D model can be determined analytically. According to various embodiments, the target image for a registered RGB image is produced with the optimal descriptor (or an estimate of the descriptor determined by optimization), which leads to a fully supervised training of the neural network 200. Additionally, the descriptor space is interpretable and optimal, irrespective of the chosen descriptor dimension d.

[0074] In the following, during the observation geometry, the 3D model is considered to be embedded in a Riemannian manifold which leads to the computation of geodesics (shortest paths between vertices). By embedding the 3D model into a d-dimensional Euclidean descriptor space such that the geodesic distances between neighboring vertices are preserved as well as possible, an explicit encoding of the optimal surface topology is achieved. The Euclidean space is considered as the descriptor space, and a search for the optimal mapping is performed: ​According to one embodiment, a Laplacian is computed for the mesh and its eigenvalue decomposition is used to determine (or at least estimate) the optimal embedding of the vertices in the descriptor space. Thus, instead of separating the geodesic computation and the mapping optimization, the descriptors are extracted in a unique framework by computing the Laplacian of the 3D model.

[0075] According to the scheme described below, an embedding of a 3D object model in Euclidean space into a descriptor space is determined in order to preserve the distances (e.g. geodesic distances) between the vertices.

[0076] In order to reduce the dimensionality via the Laplacian, a set of points should correspond to the nodes in the undirected graph. should represent the proximity or connection strength between two nodes and , e.g. .

[0077] The goal is to find a d-dimensional embedding (typically applicable d < D) such that if and are close, their embeddings should also be close:

[0078] (1)

[0079] where .

[0080] The optimization problem (1) is equivalent to

[0081] (2)

[0082] where is a semi-positive definite Laplacian matrix. A is the adjacency matrix with elements and . It should be noted that the optimal solution can have any scale and trend. In order to remove this randomness, the weighted second moment can be normalized by which forces unit variance into the different dimensions. The resulting optimization problem then becomes

[0083] (3)

[0084] The optimization is solved with a constraint using a Lagrange parameter

[0085] (4)

[0086] ​This is a generalized eigenvalue problem, which can be solved using the standard library of linear algebra. Since L and D are positive (semi-)definite matrices, the eigenvalues ​​can be written as... .

[0087] Furthermore, the first eigenvector ( The first column of the vector is equal to 1 (all vectors are 1), which is a trivial solution that maps each vertex to a point. Additionally, any two eigenvectors are orthogonal to each other. The solution to the eigenvalue problem yields N eigenvalues ​​and corresponding eigenvectors of dimension N. However, in practice, only the first d eigenvectors corresponding to the lowest eigenvalue (excluding the trivial solution) are used.

[0088] Therefore, the i-th column of Y is the path from node i to R. d The embeddings in the data are represented by rows, where each row represents the embedding of each point in different orthogonal dimensions.

[0089] Figure 3 An example embedding of a 4-node graph is shown.

[0090] Eigenvalues ​​are significant in relation to the optimality of the embedding. In the optimal embedding... Under the condition that the restrictions are met And therefore applicable:

[0091] (5)

[0092] This means that the eigenvalues ​​correspond to embedding errors of different dimensions. For simplicity, d=1, in which case each x maps to the point y=1. In this case, (5) simplifies to:

[0093] (6)

[0094] because That is, if all vertices of an object are mapped to a single point, the embedding error is 0 because the distance between all points y is 0. This is useless for practical purposes, thus omitting the first eigenvalue and eigenvector. Using d=2 corresponds to mapping each point x to a line, and λ1 is the corresponding embedding error, and so on. Since the eigenvectors are orthogonal to each other, increasing d adds a new dimension to the embedding, which aims to minimize the error in the new orthogonal dimension. The same effect can be seen in (3): because Therefore, the original target setting can be transformed to minimize the embedding error in each dimension. Without considering the chosen d, the resulting descriptor vector is optimal.

[0095] In some cases, the subsequent eigenvalues ​​are the same, that is to say (See) Figure 3where the eigenvalues are the same for d = 2 and d = 3). This exploits some information about symmetry, where there are multiple orthogonal dimensions with the same embedding error. In fact, in the 4-node graph example above, if the graph is fully connected, the embedding in each dimension is symmetric, and all eigenvalues are the same except for the trivial solution. Figure 3

[0096] The graph embedding scheme above can be directly applied to meshes, point clouds, etc. For example, a K-Nearest-Neighbor (KNN) algorithm can be used to form local connections between vertices and create an adjacency matrix. The scheme is sufficient to create a graph Laplacian and compute an embedding for each vertex. However, the scheme is inherently based on Euclidean distance metrics and heuristic construction, which do not necessarily take into account the underlying Riemannian geometry of 3D object models. For example, some edges can traverse an object, or can connect non-adjacent vertices of a mesh. Even a few incorrect entries in the adjacency matrix can cause poor embedding performance. According to one embodiment, when using a model, it is therefore ensured that the geodesic distance between any two vertices is correct or has a minimal approximation error.

[0097] In general, an object model such as a mesh or point cloud can be represented as an embedding A Riemannian manifold M with a uniformly varying metric g can be considered "locally Euclidean", which detects the locally smooth nature of real-world objects. The Laplacian operator generalizes to Riemannian manifolds is the Laplace-Beltrami (LB) operator Δ. Similar to the Laplacian in Euclidean space, the LB operator applied to a function is the divergence of the gradient of the function. While the Laplacian is readily computable for graphs and in Euclidean space (from adjacency information or finite differences), the LB operator in differential geometry is built based on exterior calculus (exterior computations), and is generally not readily available for manifolds.

[0098] However, for known discrete manifolds such as meshes, the LB operator can be approximated. This provides an efficient and simple computational framework when working with meshes, point clouds, etc. Since the Riemannian equivalent of the Laplacian is the Laplace-Beltrami operator, the embedding scheme above can be directly applied to Δ. The eigenvectors Y of Δ represent the optimal d-dimensional Euclidean embedding of the mesh vertices.

[0099] ​For a mesh, Δ can be effectively computed as follows. Assume a mesh is given with N vertices V, faces F and edges E. In this case, Δ has size N x N. The i-th row of Δ describes the adjacency information of the i-th vertex to its connected vertices. φ should be an arbitrary function at the mesh. Then, the application of the discrete LB operator on said function is mapped onto Δφ. The i-th element of said function can be described by:

[0100] (7)

[0101] Figure 4 The definition of the angle and is explained.

[0102] The sum of the cotangent expressions is used as the connection weight w ij . According to one embodiment, these weights occurring in (7), i.e. the weights of the LB operator, are used as weights for determining D and A of equation (2) when used in a mesh.

[0103] It should be noted that due to negative connection weights w ij may occur, in particular if one angle is significantly larger than the other angles (bad face). To overcome said problem, the connection weights can be approximated by Edge Flipping.

[0104] The descriptor generation scheme described above univocally handles each vertex. That is, each vertex is assigned to a univocal descriptor. However, an object can be symmetric and thus assigning a univocal descriptor to seemingly identical vertices causes an asymmetric embedding.

[0105] To solve said problem, according to various embodiments, the intrinsic symmetry of a shape is detected and symmetric embeddings are compressed so that symmetric vertices are mapped onto the same descriptor. A shape can be shown to have intrinsic symmetry if the eigenfunctions of the Laplace-Beltrami operator appear symmetric in Euclidean space. In other words, if the geodesic distance preserving Euclidean embedding (descriptor space) of a mesh, point cloud, etc. shows Euclidean symmetry, then the symmetric feature of said mesh, point cloud, etc. is detected. A compact manifold has intrinsic symmetry if there exists a homeomorphism T that preserves the geodesic distance between each vertex of the manifold.

[0106] To compress symmetric descriptors, so-called global intrinsic symmetric invariants (GISIF) can be used. In case of assuming a global intrinsic symmetric homeomorphism and a function of the manifold f, where g stands for the geodesic distance, if it holds for each point p on the manifold that:

[0107] (8)

[0108] f is a GISIF. For example, on a torus, the homeomorphism should be arbitrary rotation around the z-axis. This means that if f is a GISIF, it must be invariant with respect to the rotation.

[0109] Furthermore, it can be shown that in the case of eigenvalues This GISIF is the sum of the squared components of the eigenvector of the point, i.e.

[0110] .

[0111] This is in line with the above analysis of identical eigenvalues, which is a necessary condition for symmetric embedding. Since in practice there are rarely identical eigenvalues due to numerical limitations, there one can use a heuristic approach where eigenvalues are considered identical if they lie within the same ε-ball (with a small ε), i.e. if the eigenvalues differ by less than a pre-set threshold, e.g. 0.1% or 0.01%. Since one has to find the symmetric dimension for a given object only once, this can be done manually.

[0112] For example, for a torus, the first 7 eigenvalues of the eigenvalue decomposition should be as follows:

[0113] .

[0114] — without considering the trivial solution — to The GISIF embedding into

[0115] .

[0116] In the case of multiple objects, this can represent multiple separate connected graphs. In this case, the adjacency matrix is block-diagonal. The symmetric positive definite Laplacian re- gains orthogonal eigenvectors. In comparison to the single graph embedding case, there are two differences in the result of the eigenvalue decomposition: First, the non-decreasing eigenvalues will be the unsorted embedding errors of all objects. Second, the eigenvectors have zero entries, since the corresponding eigenvalues remain orthogonal. This means that each dimension of the descriptor space will correspond to only one object embedding. Furthermore, the dimensions are ordered with respect to the embedding error of the corresponding object. Thus, if one should produce a 3-dimensional embedding of two objects, one uses d = 8, since there are two trivial solutions corresponding to λ = 0.

[0117] The uncomplicated approach operates on multiple objects independently, while there can be sub-optimal methods that thus provide a comparably good embedding with lower d that makes use of the relations between the objects.

[0118] Given the pose of the object, a target image can be produced by projecting the descriptors onto the image plane. The descriptor space - random image noise or a single descriptor mapped onto the farthest point in the descriptor space can be used as non-object (background).

[0119] To improve the robustness of the trained network 200, image augmentation methods such as domain randomization or perturbations, e.g. Gaussian soft focus, cropping or loss can be used.

[0120] In the above example, the target image is produced by mapping the descriptors associated with the vertices onto locations in the target image, wherein the locations in the target image have the vertices in the target input image. The target image produced in a geometric way is also referred to as geometric (descriptor) target image in the following.

[0121] According to various embodiments, in addition to the geometric target image, a target image with view-dependent descriptors is produced and used for training (alone or together with the geometric target image).

[0122] Figure 5 The geometric target image is shown.

[0123] Figure 6 The view-dependent target image produced from the geometric target image in Figure 5 The view-dependent target image produced from the geometric target image in

[0124] By this, the descriptor component values (e.g. per color channel) in the object edge region (e.g. at locations within the object and at a specific distance from the object edge) are equalized with the descriptor component values of the background of the target image (where the background can also only have unique descriptor component values). By this, intuitively a descriptor is produced in the target image which mimics a view-dependent edge detection descriptor separating the object and the background.

[0125] Figure 7 The channel of the view-dependent target image produced from the geometric target image in Figure 5 The channel of the view-dependent target image produced from the geometric target image in

[0126] In the described example, one color channel (e.g. the B channel of an RGB image) is used as a mask, where the interior of the object is emphasized with respect to the edge regions of the object, i.e. the descriptor component values at positions further away from the edge are more strongly distinguished from the background, i.e. more strongly distinguished from positions closer to the edge. It is also possible to simply set the descriptor component values of the channel to zero at positions in the interior of the object which are at a minimum distance from the edge of the object (in its view in the target image), while all other descriptor component values (i.e. at positions in the background and at positions closer to the edge than the minimum distance) are set to zero. In this way, the channel specifies a mask for the interior of the object. The descriptor component values can also gradually change from zero (from the edge) to 1 (at the center of the object in its view in the target image).

[0127] Figure 6 The processing of the geometry target image smoothes the geometry target image at the edges, while Figure 7 The processing of the geometry target image embeds an object mask (by replacing the descriptor component values of one channel or adding a channel).

[0128] The view-dependent target image can be automatically computed for the geometry target image by means of the object mask.

[0129] However, it should be noted that it is not necessary to generate a geometry target image in order to generate the view-dependent target image belonging thereto. Especially in the case of masking using a channel, as in the example of Figure 7 It is not necessary to create a geometry target image for the channel with geometry descriptor component values (i.e. geometry descriptor component values resulting from mapping the descriptors of the object vertices) in order to generate the view-dependent target image for the channel.

[0130] However, the view-dependent image belonging to the geometry target image can also be generated by re-rendering the geometry target image (e.g. in the case of edge blending as in the example of Figure 6 ).

[0131] It should also be noted that the edges of the object in the target image can correspond to different parts of the object, depending on how the object is turned in the target image. Thus, the descriptors of the parts of the object are view-dependent and are view-dependent in this sense.

[0132] In summary, according to various embodiments, a method as described in Figure 8 is provided.

[0133] Figure 8 A method for training a machine learning model for recognizing the object topology of an object from an object image is shown according to one embodiment.

[0134] In 801, a 3D model of an object is obtained, wherein the 3D model comprises a mesh of vertices.

[0135] In 802, descriptor component values for each vertex of the mesh are determined;

[0136] In 803, training data image pairs are generated, wherein each training data image pair comprises a training input image and a target image showing the object, and wherein generating the target image comprises the following:

[0137] determining vertex positions of vertices of an object model of the object, the vertices having the vertex positions in the training input image;

[0138] assigning, for each determined vertex position in the training input image, a descriptor component value determined for a vertex at the vertex position to a position in the target image; and

[0139] adapting at least some of the descriptor component values assigned to positions in the target image or adding descriptor component values to positions of the target image, such that the adapting or adding assigns descriptor component values to positions within the object and within a preset distance from edges of the object in the target image, which are closer to descriptor component values assigned to positions outside the object than to descriptor component values assigned to positions within the object, wherein the positions within the object are further away from the edges of the object than the preset distance.

[0140] In 804, the machine learning model is trained by supervised learning using the training data image pairs as training data.

[0141] In other words, according to various embodiments, the object is embedded into a descriptor space and a training target image for a training input image recorded with a camera is generated by projecting onto the image plane of the camera. Furthermore, the descriptor component values are adapted or added such that edges of the object in the target image are less sharp, at least for one channel.

[0142] It should be noted that the descriptor component values can be components of a descriptor vector. In the above example, for instance, the descriptor is a 3-dimensional vector, the first component of which is considered as a red value, the second component as a green value and the third component as a blue value, for instance, such that the target image is an RGB image.

[0143] In connection with Figure 8The term "descriptor component value" can relate to one component such that, for example, only the blue channel of the descriptor vector is adapted when adapting. A descriptor component value of a component of a first vector and a descriptor component value of a component of a second vector refers to the descriptor component value out of the first vector and the corresponding descriptor component value out of the second vector, i.e. the descriptor component value belonging to the same dimension of the first vector or having the same index or the same row number when the descriptor is represented as column vector or belonging to the same color channel when interpreted as color value.

[0144] In the above examples the machine learning network is described as a neural network, while other types of regressors mapping a 3D tensor to another 3D tensor can be used.

[0145] According to various embodiments, the machine learning model assigns descriptors to pixels of the object (in the image plane). This can be seen as an indirect encoding of the surface topology of the object. This connection between the descriptors and the surface topology can be performed explicitly by rendering in order to map the descriptors onto the image plane. It should be noted that the descriptor component values at a face of the object model (i.e. a point that is not a vertex) can be determined by means of interpolation. For example, if a face is given by 3 vertices of the object model with their respective descriptor component values y1, y2, y3, then at any point of this face the descriptor component values y can all be computed as a weighted sum w1 y1 + w2 y2 + w3 y3 of these values. In other words, the descriptor component values are interpolated at the vertices.

[0146] In order to produce image pairs for training data, an image (e.g. an RGB image) of an object including one (or more) objects with a known 3D (e.g. CAD) model and a known pose (in a global (i.e. world) coordinate system) is mapped onto a (dense) descriptor image which is produced by means of finding descriptors that minimize a geometric property (in particular the proximity of the object's points) between the object model and its representation (embedding) in the descriptor space. In practical use, the theoretically optimal solution for the minimization is usually not found, as the finding is restricted to a certain search space. Nevertheless, an estimate of the minimum is determined within the limits of practical application (available computational precision, maximum number of iterations, etc.).

[0147] Thus, by performing the minimization process of the sum of the distances between the descriptors of pairs of connected vertices in a way that the distances between the descriptors of pairs of vertices are weighted by the weights of the edges between the pairs of vertices, the descriptors of the vertices are found, wherein each descriptor is found for a respective one of the vertices.

[0148] Each training data image pair comprises a training input image and a target image of an object, wherein the target image is generated by projecting the descriptors of the vertices that are visible in the training input image onto the training input image plane according to the pose the object has in the training input image. Then, the descriptor values can also be adapted as described above, or the descriptor values can be added as described above.

[0149] The images are used for supervised training of a machine learning model together with their associated target images.

[0150] Thus, a machine learning model is trained in order to recognize unambiguous features of an object (or multiple objects). By means of evaluating the machine learning model in real time, the information can be used for various applications in robot control, e.g. predicting an object gripping pose for assembly. It should be noted that the supervised training scheme enables an explicit encoding of symmetry information.

[0151] Figure 8 The method of the present application can be executed by one or more computers comprising one or more data processing units. The expression “data processing unit” can be understood as any kind of entity that enables processing data or signals. For example, data or signals can be processed in accordance with at least one (i.e. one or more than one) specific function performed by the data processing unit. The data processing unit can comprise or be formed by an analog circuit, a digital circuit, a mixed signal circuit, a logic circuit, a microprocessor, a microcontroller, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), an integrated circuit, or any combination thereof. Any other means for implementing the respective functions described in more detail below can also be understood as a data processing unit or logic circuitry. It goes without saying that one or more of the method steps described in detail herein can be performed (e.g. implemented) by the data processing unit via one or more specific functions performed by the data processing unit.

[0152] The expression “robot” can be understood as relating to any physical system (having a mechanical part whose motion is controlled), such as a computer-controlled machine, a vehicle, a household appliance, a power tool, a personal assistant, or an access control system.

[0153] The various embodiments can receive various sensor signals, such as (e.g. RGB) camera, video, radar, LiDAR, ultrasound, thermal imaging technology, etc., and are used, for example, to obtain sensor data showing an object. The embodiments can be used to generate training data and train a machine learning system, for example, for autonomously controlling a robot, e.g. a robot manipulator, in order to implement various manipulation tasks in various scenarios. In particular, the embodiments can be used when implementing and supervising manipulation tasks, e.g. in an assembly line.

[0154] While particular embodiments are described and illustrated herein, it will be appreciated that those skilled in the art will contemplate a variety of alternatives to the specific embodiments described and illustrated herein without departing from the scope of the present application. The application intended to cover any adaptations or variations of the specific implementations discussed herein. Therefore, it is intended that the application be limited only by the claims and their equivalents.

Claims

1. A method for training a machine learning model for recognizing an object topology of an object from an object image, the method comprising the following: obtaining a 3D model of the object, wherein the 3D model comprises a mesh of vertices; determining descriptor component values for each vertex of the mesh; generating training data image pairs, wherein each training data image pair comprises a training input image and a target image showing the object, and wherein generating the target image comprises the following: determining vertex positions of vertices of an object model of the object, which vertices have the vertex positions in the training input image; and assigning, for each determined vertex position in the training input image, a descriptor component value determined for the vertex at the vertex position to a position in the target image; adapting at least some descriptor component values assigned to positions in the target image or adding descriptor component values to positions of the target image, such that descriptor component values assigned to positions inside the object and within a preset distance from an object edge in the target image are adapted or added to be closer to descriptor component values assigned to positions outside the object than to descriptor component values assigned to positions in the interior of the object, which are further away from the object edge than the preset distance; and training the machine learning model by supervised learning using the training data image pairs as training data.

2. The method according to claim 1, having determining descriptors for each vertex of the mesh, wherein the descriptors are vectors having a vector component for each of a plurality of channels, and adapting descriptor component values of a channel, such that descriptor component values of the channel assigned to positions inside the object and within a preset distance from an object edge in the target image are adapted to be closer to descriptor component values of the channel assigned to positions outside the object than to descriptor component values of the channel assigned to positions in the interior of the object, which are further away from the object edge than the preset distance; and / or adding a channel having descriptor component values, such that descriptor component values of the added channel assigned to positions inside the object and within a preset distance from an object edge in the target image are closer to descriptor component values of the added channel assigned to positions outside the object than to descriptor component values of the added channel assigned to positions in the interior of the object, which are further away from the object edge than the preset distance.

3. The method according to claim 2, wherein one channel of the plurality of channels is replaced by the added channel.

4. The method according to claim 2 or 3, wherein the descriptor component values of the added channel are chosen such that the descriptor component values monotonically vary when traversing positions from positions outside the object to positions in the center of the object.

5. The method according to any one of claims 1 to 3, wherein the descriptor component values are adapted such that descriptor component values assigned to positions inside the object within a preset distance and descriptor component values assigned to positions outside the object are equalized. ​ ​ 6. The method of any one of claims 1-3, wherein generating the training data image pairs comprises: obtaining a plurality of images of the object in different poses, and generating training data image pairs from each of the obtained images by generating a target image for the obtained image.

7. The method of any one of claims 1 to 3, comprising determining vertex positions of vertices of the object model of the object from respective poses the object has in the training input images, wherein the vertices have the vertex positions in the training input images.

8. The method of any one of claims 1 to 3, wherein vertices of the 3D model are connected by edges, wherein each edge has a weight that specifies a closeness of two vertices of the object connected by the edge, and wherein, The descriptor component value of each vertex of the mesh is determined by looking up the descriptor component value of the vertex of the pair of connected vertices that minimizes the distance between the descriptor component values of the pair of connected vertices.

9. The method of claim 8, wherein looking up the descriptor component values comprises determining an eigenvector of a Laplacian matrix of a graph formed by the vertices and edges of the 3D model and adopting components of the eigenvector as descriptor component values.

10. The method of claim 9, comprising associating each vertex with a component position in the eigenvector, and for each vertex adopting the component in the eigenvector at the component position associated with the vertex as the descriptor component value of the vertex.

11. The method of claim 10, wherein determining descriptor component values further comprises combining components of eigenvectors whose eigenvalues differ by less than a preset threshold into a single component.

12. The method of any one of claims 1-3, wherein obtaining a 3D model of the subject comprises: obtaining a 3D mesh of vertices and edges modeling the object, and assigning edge weights of a Laplace-Beltrami operator as weights of the edges, the Laplace-Beltrami operator being used at the mesh.

13. A method for controlling a robot, comprising the following: training a machine learning model with the method of any one of claims 1 to 12; obtaining an image showing the object; inputting the image into the machine learning model; determining a pose of the object from an output of the machine learning model; and controlling the robot in accordance with the determined pose of the object.

14. The method of claim 13, wherein determining an object pose comprises determining a position of a portion of the object, and wherein controlling the robot in accordance with the determined pose of the object comprises controlling an end effector of the robot to move to the position of the portion of the object and to interact with the portion of the object.

15. A robot, comprising the following: a camera designed for providing image data of an object; a control device designed for implementing a machine learning model; and a training device designed for training the machine learning model by means of the method of any one of claims 1 to 12.

16. The robot of claim 15, comprising at least one effector, wherein the control device is designed for controlling the at least one effector with an output of the machine learning model.

17. A computer program product comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 14.

18. A computer readable medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Visual recognition and positioning method for robot intelligent capture application

    CN108171748A

  • Method for determining a pose of an object in the surroundings of the object by means of multi-task learning, and control means

    CN111566700A