Method and apparatus for training a machine learning model

A self-supervised training method using two deep neural networks optimizes image feature extraction for robustness and efficiency by learning keypoint correspondence and descriptor similarity from image pairs, addressing inefficiencies in existing methods and enhancing localization capabilities.

DE102024201292A1Pending Publication Date: 2025-08-14ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024201292
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-13
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing image feature extraction methods lack robustness against perspective changes, capture conditions, and environmental variations, and require manual construction of cost functions with hyperparameter tuning, making them inefficient and suboptimal.

Method used

A self-supervised training method using two deep neural networks for image feature extraction, where image pairs with similar content but different perspectives are processed to learn keypoint correspondence and descriptor similarity, optimizing a cost function based on photometric differences without requiring ground truth camera poses.

Benefits of technology

The method achieves improved robustness and efficiency in image feature extraction, enabling accurate relative pose determination and localization tasks, particularly in autonomous vehicles and robotics, without manual hyperparameter tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention describes a method and an apparatus for training a machine learning model comprising a first and a second deep neural network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method and a device for training a machine learning model having a first and a second deep neural network. State of the art

[0002] In modern image processing and computer vision, methods for extracting local image features play a major role. These techniques are important for a wide range of applications, from automated image recognition to 3D reconstruction and robotics. To be effective, image feature extraction methods must be carefully engineered to possess specific properties, especially when determining the relative position of one image compared to another.

[0003] One of the key aspects in this area is the robustness of keypoint detection to various types of changes, such as perspective changes, changes in appearance due to different recording conditions, or changes in the environment. This robustness is crucial because it ensures that, despite these challenges, the important points, the so-called keypoints, can be detected consistently and reliably.

[0004] In addition to robustness, the uniform distribution of keypoints across the entire image is another important criterion. A clump-free, homogeneous distribution of keypoints enables a more comprehensive and representative capture of the information contained in an image.

[0005] Another critical factor is the quality of the keypoint descriptors. Descriptors describing corresponding keypoints should be highly similar, while descriptors for non-corresponding keypoints should be as different as possible. This enables effective feature matching, which forms the basis for subsequent applications such as camera pose estimation.

[0006] The integration of these properties by design into image feature extraction methods, such as SIFT (Lowe, 2004) or by learning appropriate cost functions in more recent methods such as SuperPoint (DeTone et al., 2018) and KP2D (Tang et al., 2020), has proven crucial. In particular, learned image feature extraction methods have often proven superior to manually constructed alternatives, primarily due to the more complex image processing operators developed during training and adapted to the specific data.

[0007] Furthermore, image reconstruction based on local image features, as explored, for example, by Jin et al. (2022) in "TUSK: Task-agnostic unsupervised keypoints," is an established method. This method uses local image features to reconstruct the original image.

[0008] Even though several approaches are already known from the state of the art, there is still potential for development.

[0009] It is therefore an object of the invention to provide an improved method and an improved apparatus for training a machine learning model.

[0010] The problem is solved by a method for training a machine learning model according to the features of patent claim 1. The problem is solved by an inference method for image feature extraction from images for relative pose determination of image content according to the features of patent claim 10. The problem is solved by a device for training a machine learning model according to the features of patent claim 12. Disclosure of the invention

[0011] According to a first aspect, a method for training a machine learning model having a first and a second deep neural network is provided. The method comprises the steps of: providing two training images that have at least partially identical or similar image content, wherein the two training images depict the image content from different perspectives and / or at different times; Processing the two training images by the first deep neural network such that image features are extracted from the two training images by setting corresponding keypoints in the two training images and describing the corresponding keypoints with the same or similar descriptors; Matching the extracted image features of the two training images, wherein the image features are matched with the same or similar descriptors to form image feature pairs; for each matched image feature pair, Extracting keypoint information from one of the two training images and descriptor information from the other of the two training images; Reconstructing one of the two training images, for which information of the respective keypoints of the matched image feature pairs was extracted, by the second deep neural network based on the extracted information of the respective keypoints and on the basis of the extracted information of the respective descriptors of the matched image feature pairs; Quantifying differences between the reconstructed training image and the provided one of the two training images using a cost function; Adjusting parameters of the first and / or second deep neural network based on the quantified differences to optimize the cost function to provide a trained machine learning model.

[0012] It is understood that the steps according to the invention, as well as other optional steps, do not necessarily have to be performed in the order shown, but can also be performed in a different order. Furthermore, additional intermediate steps can be provided. The individual steps can also comprise one or more substeps without thereby departing from the scope of the method according to the invention.

[0013] According to a second aspect, a device for training a machine learning model comprising a first and a second deep neural network is provided. The device comprises an evaluation and / or computing device configured to perform the following steps: providing two training images that at least partially have identical or similar image content, wherein the two training images depict the image content from different perspectives and / or at different times; Processing the two training images by the first deep neural network such that image features are extracted from the two training images by setting corresponding keypoints in the two training images and describing the corresponding keypoints with the same or similar descriptors; Matching the extracted image features of the two training images, wherein the image features are matched with the same or similar descriptors to form image feature pairs; for each matched image feature pair, Extracting keypoint information from one of the two training images and descriptor information from the other of the two training images; Reconstructing one of the two training images, for which information of the respective keypoints of the matched image feature pairs was extracted, by the second deep neural network based on the extracted information of the respective keypoints and on the basis of the extracted information of the respective descriptors of the matched image feature pairs; Quantifying differences between the reconstructed training image and the provided one of the two training images using a cost function; Adjusting parameters of the first and / or second deep neural network based on the quantified differences to optimize the cost function to provide a trained machine learning model.

[0014] The statements made for the method apply accordingly to the device. The device may be part of a system. It is understood that linguistic modifications of procedurally formulated features can be reformulated for the device according to common linguistic practice, without such formulations having to be explicitly listed here.

[0015] This article describes a method for training a machine learning method that is used primarily for the extraction of local image features. These features are particularly suitable for determining the relative pose of images or between multiple images. Pose refers to the position and orientation of an image feature that at least partially represents an image's content. In the following, the term "deep neural network" is used to represent any machine learning method.

[0016] Each image feature has at least one keypoint, which specifies the image coordinates of the image feature. Each image feature has at least one descriptor, which describes the local environment of the keypoint in the image.

[0017] By assigning or matching corresponding image features between the at least two training images, which at least partially show the same or a similar environment, a camera pose of one image relative to another image can be determined. Two image features are preferably considered corresponding if they characterize the same point in the real 3D world that is visible in the different training images. A model trained according to the method can be used purely for example for the self-localization of (autonomous) vehicles, drones, and / or robots.

[0018] For each matched image feature pair, the descriptor of the corresponding image feature is preferably extracted from one of the two images, and the keypoint of the matched image feature is preferably extracted from the other of the two images. The idea here is that geometric information can be determined from one of the two images in this way. Secondly, photometric information, in particular, can be determined from the other of the two images. Since the matched image features preferably represent the same physical locations in the environment, a descriptor of an image feature from the other of the two images is also a suitable descriptor for the corresponding image feature from one of the two images. The collected information about the matched image features is preferably passed on to the second deep neural network as input.The goal of this network is to create the most accurate reconstruction of one of the two images.

[0019] The method described here for training a machine learning model for image feature extraction for localization is particularly frugal in its requirements for the training data or training images used. Preferably, only image pairs are required that represent at least partially or in sections the same or a similar environment, particularly when viewed at a threshold value. For example, precise camera poses are not required for object localization based on the (training) images. This property facilitates the use or inference of the trained model and thus enables scaling for training on large datasets (e.g., > 1000 image pairs). This makes the present training approach suitable for training a foundation model (also referred to as a base model).

[0020] The method is based on using information about the matched image features between the two training images to reconstruct one of the two images. For this purpose, two deep neural networks are trained independently of each other. The first network extracts the image features, preferably local ones. The second network uses the matches between the image features of two images for image reconstruction.

[0021] Image reconstruction describes the task that is solved, in particular, in a self-supervised manner during training. Image reconstruction is therefore not necessarily the goal of training. The goal of training, or rather the ability of the trained model, is preferably to determine the image features in such a way that the resulting matches are suitable for image reconstruction. This is again the case if the image features are suitable for localization tasks. The trained model has preferably learned, in a supervised manner, to perform image feature extraction based on the matching. The invention is based on matching image features between two different images and subsequently using the information from these matches for reconstruction. The matching used is preferably differentiable in this case.

[0022] The object of the method described in the invention is to train at least one deep neural network to perform image feature extraction from an image, which enables relative pose determination between at least two images. The model or network trained here impresses with comparatively better performance in pose determination and can also be easily adapted to any pose determination application by appropriately selecting the training images.

[0023] Compared to other machine learning methods for training deep neural networks for image feature extraction, the features for image feature extraction are not learned through a suitable cost function, which, for example, forces descriptors of corresponding image features and / or keypoints to receive descriptors that are as similar as possible, as would be the case with contrastive loss, for example. Rather, the features arise implicitly from the task being solved during training.

[0024] A disadvantage of state-of-the-art explicit cost functions is that they are "manually constructed" by a human. There are numerous ways to construct the costs for the same goal. Thus, it is unlikely that the optimal variant of the cost function will be selected with this manual approach. Furthermore, manually created cost functions must be weighted. These are additional hyperparameters that typically have to be correctly adjusted manually. Therefore, using cost functions requires numerous decisions that depend on the user's intuition.

[0025] The present training method deviates from this. The task of image reconstruction based on matched, especially local, image features intrinsically requires that the image feature matching function well. Image feature matching works in this case if keypoints are found in at least two (training) images at the same locations in the two images, and if the descriptors of these corresponding keypoints are the same or similar, while the descriptors of non-corresponding keypoints are different. Furthermore, a uniform distribution of keypoints in the respective image is advantageous, as this makes more information available for reconstruction.

[0026] The present method, in which the desired properties of the trained model arise implicitly from the task solved during training, could alternatively be obtained by directly training a deep neural network for a localization task. The disadvantage of such an alternative approach, however, is that the correct solution—for example, a relative camera position and / or orientation between two images—must be known in order to use it as a ground truth value for the training procedure. However, such precise ground truth for camera poses is often not directly available.

[0027] Furthermore, it is technically complex to generate these camera poses, for example using structure-from-motion methods.

[0028] The method presented here does not require such camera poses because an error between an actual image and a reconstruction of this image is used.

[0029] The advantage over other state-of-the-art methods is that the described method is self-supervised, i.e., a precise solution for the localization task to be solved does not need to be known in advance, even though the deep neural network determines a correspondence between real and reconstructed image pairs during training. This also differs from other self-supervised methods that do not use real image pairs during training, but rather use artificially generated image pairs through homography warping, for which the true image feature correspondences can be determined based on the known homography. The use of real image pairs, in contrast, has the advantage that real variations in the appearance of the same or a similar environment or image content can be learned.

[0030] The (training) images are preferably captured and provided by an optical sensor. The images can be preprocessed or available as raw data captured by the sensor. The sensor can be a camera, a lidar sensor, a radar sensor, an ultrasonic sensor, an infrared sensor, or another imaging sensor.

[0031] The image pairs used for training preferably have at least some of the same image content, i.e., they show the same environment, e.g., images taken of tourists in front of the Eiffel Tower or in front of a specific mountain range. On the other hand, the image pairs used for training are preferably not identical, but rather show the environment from a different perspective and / or at a different time.

[0032] Instead of real image pairs, image pairs of which one or both were synthetically generated can also be used. Such image pairs can be part of a pre-training process. Synthetically generated image pairs can be generated through artificial transformations, such as homographies.

[0033] In one embodiment, to provide the trained machine learning model, the aforementioned steps of processing, matching, extracting, reconstructing and quantifying are at least partially repeated iteratively until a termination criterion or a threshold of the cost function is reached.

[0034] In one embodiment, the cost function determines an error between the reconstructed training image and the provided one of the two training images based on the quantified differences.

[0035] In one embodiment, the cost function determines the error as the sum of pixel differences between the reconstructed training image and the provided one of the two training images as a photometric error.

[0036] The cost function of the training method is therefore preferably based on the photometric difference between the actual and the reconstructed image, thus the true relative pose between the images is not required for training. The cost function is preferably evaluated by assessing the photometric difference between the reconstructed and the corresponding original image. The sum of the pixel value differences of pixels with the same image coordinates can be added together. Based on the costs of image reconstruction, the gradients for a gradient descent method (backpropagation) for training the first deep neural network are preferably calculated. Since both the first and second deep neural networks and / or the feature matching are preferably differentiable, this is a method that can be trained end-to-end. This means:The output costs are used to adjust the parameters of the first and second deep neural networks, and possibly also of the matching method, so that the costs for the same input (an image) are lower in the future. As an alternative to the photometric error, perceptual similarity can also be considered.

[0037] In one embodiment, the corresponding keypoints are placed at the same or similar locations of the two training images

[0038] In one embodiment, the first deep neural network comprises an encoder, and the second deep neural network comprises a decoder.

[0039] The first deep neural network, the encoder, preferentially determines the image features for the two training images. The encoder's architecture can be chosen similarly to a KP2D structure, as disclosed by Tang et al. (2020) "Neural outlier rejection for self-supervised keypoint learning." However, the KP2D training method is not adopted, or is only used as pre-training, allowing the encoder to already extract image features. Image feature extraction can then be further improved using the present training method.

[0040] In one embodiment, a plurality of two training images, each preferably representing image pairs of the same or similar image content, are provided.

[0041] The number of training images used for training can preferably be selected depending on the application and / or the required performance. After the described training method has been carried out with a large number of image pairs, the first deep neural network, in particular the encoder, can preferably be used independently of the other components for image feature extraction. For example, the trained first deep neural network can be used in an autonomous vehicle, which localizes itself in its environment using the image feature correspondences between a current image and, for example, a previously captured image. The described method can continue to be used for feature matching; it is also used during training and, if necessary, adapted during inference.

[0042] In one embodiment, processing the two training images by the first deep neural network further comprises calculating an image descriptor for at least one of the two training images. Reconstructing one of the two training images is preferably also performed based on the image descriptor associated with one or the same of the two training images.

[0043] In this embodiment, the first deep neural network preferably calculates an image descriptor, i.e. a compact d-dimensional representation of the image, in addition to the local image features. Such image descriptors are used, for example, for place recognition tasks, in that images with at least partially identical content have similar image descriptors. In contrast to local image features, image descriptors are not suitable for the precise relative pose determination between two images. The image descriptor of one of the two images determined by the first network is preferably additionally made available to the second network for generating the reconstruction. This represents one possibility for the second network to receive information about the style of the image to be reconstructed, e.g. that it is an image taken at night and / or in winter.Since the second network otherwise only receives photometric information from the other image, the second network can otherwise only use the style of the other image for reconstruction, which may, however, differ significantly from the actual style of the image to be reconstructed. This variant of the invention thus has the additional advantage that the image descriptors determined by the first network can be used for an initial rough localization before a precise localization is subsequently performed based on the local image features.

[0044] In one embodiment, each pixel or at least a plurality of pixels or at least a predetermined grid of pixels or every n-th pixel, where n may be > 1, of the two training images each defines a keypoint.

[0045] The present method can only highlight a small subset of pixels in an image using keypoints and consider only these keypoints during matching. Alternatively, a dense method can be used, in which every pixel or every nth number of pixels serves as a keypoint. In this case, the resolution of the (input) image is preferably reduced so that the computational load is not too great.

[0046] The method can also be used (e.g., in the dense variant) for self-supervised training of a foundation model. This model preferentially learns to extract local image features from large data sets that are suitable for subsequent 3D image processing tasks (such as video-based localization).

[0047] This article also presents an inference method for image feature extraction from images for relative pose determination of image content. The inference method includes: Providing two images, each preferably captured by an optical sensor, Extracting image features from the provided images by a machine learning model trained according to the present method; Relative pose determination of image contents of the two images based on the extracted image features by the machine learning model trained according to the present method.

[0048] Particularly preferably, the inference method is used for environmental localization and / or self-pose determination based on the relative pose determination of image content for an automated driving function of a motor vehicle, an automated function of a drone, or an automated function of a robot. The inference method can be used, for example, for video-based self-motion estimation for use in the field of autonomous parking.

[0049] The present invention also claims a control device which is used for a partially automated or automated driving function of a motor vehicle and / or a drone and / or in a robotics system and / or in an industrial machine and / or is used for optical inspection, and on which a machine learning model trained according to the invention can be executed.

[0050] The present invention also claims a computer program with program code for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention relates to a computer program (product) comprising instructions that, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.

[0051] The present invention also proposes a computer-readable data carrier containing program code of a computer program for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention relates to a computer-readable (storage) medium comprising instructions which, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.

[0052] The described designs and further training courses can be combined as desired.

[0053] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the embodiments that are not explicitly mentioned. Short description of the drawings

[0054] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate embodiments and, in conjunction with the description, serve to explain principles and concepts of the invention.

[0055] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements illustrated in the drawings are not necessarily drawn to scale.

[0056] They show: Fig. 1 a schematic flow diagram of the present training method; Fig. 2 is a schematic block diagram of the present training method according to an embodiment; and Fig. 3 a schematic block diagram of the present training method according to another embodiment.

[0057] In the figures of the drawings, the same reference symbols designate the same or functionally equivalent elements, parts or components, unless otherwise stated.

[0058] Fig. 1 shows a schematic flow diagram of an exemplary method for training a machine learning model having a first deep neural network and a second deep neural network.

[0059] In any embodiment, the method can be carried out at least partially by a device 100, which for this purpose can comprise several components not shown in detail, for example, one or more provision devices and / or at least one evaluation and computing device. It is understood that the provision device can be designed jointly with the evaluation and computing device or can be different from it. Furthermore, the system can comprise a storage device and / or an output device and / or a display device and / or an input device.

[0060] The preferred computer-implemented training method comprises at least the following steps: In a step S1, two training images are provided which at least partially have the same or similar image contents, wherein the two training images depict the image contents from different perspectives and / or at different times.

[0061] In a step S2, the two training images are processed by the first deep neural network in such a way that image features are extracted from the two training images by setting corresponding keypoints in the two training images and describing the corresponding keypoints with the same or similar descriptors.

[0062] In a step S3, the extracted image features of the two training images are matched, whereby the image features with the same or similar descriptors are matched to image feature pairs.

[0063] For each matched image feature pair, the following steps are preferably performed at least once, preferably iteratively: In a step S4, information of the key point is extracted from one of the two training images and information of the descriptor is extracted from the other of the two training images.

[0064] In a step S5, one of the two training images, for which information of the respective key points of the matched image feature pairs was extracted, is reconstructed by the second deep neural network on the basis of the extracted information of the respective key points and on the basis of the extracted information of the respective descriptors of the matched image feature pairs.

[0065] In a step S6, differences between the reconstructed training image and the provided one of the two training images are quantified using a cost function.

[0066] In a step S7, parameters of the first and / or the second deep neural network are adjusted based on the quantified differences to optimize the cost function to provide a trained machine learning model.

[0067] Fig. Figure 2 shows a schematic block diagram of an embodiment of the present method. Two training images A and B (hereinafter also referred to as images A and B) show at least partially the same or similar image content, for example from different perspectives and / or at different (capturing) times. The training images are each processed by the first deep neural network 200, an encoder. This network 200 extracts local image features from the training images with the aim of determining corresponding image features. In doing so, key points 202, 203 (circles in Fig. 2 and Fig. 3) placed at approximately the same positions in the training images A, B and described (per image) with similar descriptors D, E. The descriptor values ​​are in Fig. 2 and Fig. 3 for each of the circles per training image A, B. The Fig. 2 and Fig. 3 marked key points 202 are described per image A, B by an identical or similar descriptor D. The Fig. 2 and Fig. The key points 203 marked in 3 are described by an identical or similar descriptor E for each image A and B. The positions of the circles correspond to the key points 202 and 203.

[0068] In the next step, a differentiable matching is performed between the image features from image A and B. The corresponding features with similar descriptor values ​​D and E are matched into pairs. The matched pairs are connected by lines L in the figure.

[0069] For each matched image feature pair from images A, B, information about the corresponding keypoint 202, 203 in image B, as well as the descriptor D, E from image A, is extracted. This information is passed to a second deep neural network 204, a decoder.

[0070] The second network 204 reconstructs image B based on the extracted information (keypoints 202, 203 from image B, with descriptors D, E from image A).

[0071] A cost function 205 quantizes the error made based on the differences between the reconstruction B' of image B and the actual image B, e.g. as the sum of the pixel differences between B' and B (photometric error).

[0072] Fig. Figure 3 shows a schematic block diagram of an embodiment of the present method. In contrast to Fig.1, in this variant, additional information about the appearance of image B is made available for reconstruction. In this example, image B not only shows the environment depicted in image A from a different perspective, but the appearance has also changed, e.g., because the season and / or the time of recording has changed. Since the second network 204 only had keypoints 202, 203 from image B and descriptors D, E from image A available for reconstruction, there is no way to predict the change in appearance. For this purpose, according to this variant, the first network 200 determines, in addition to the local image features, an image descriptor 206 for image B, which is additionally made available to the second network 204 as input for reconstruction. QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature

[0000] Lowe, 2004

[0006] Jin et al. (2022) “TUSK: Task-agnostic unsupervised keypoints

[0007] Tang et al. (2020) “Neural outlier rejection for self-supervised keypoint learning

[0039]

Claims

[1] A method for training a machine learning model having a first and a second deep neural network (200, 204), the method comprising the steps of: Providing (S1) two training images (A, B) which at least partially have the same or similar image contents, wherein the two training images (A, B) depict the image contents from different perspectives and / or at different times; Processing (S2) the two training images (A, B) by the first deep neural network (200) such that image features are extracted from the two training images (A, B) by setting mutually corresponding key points (202, 203) in the two training images (A, B) and describing the mutually corresponding key points (202, 203) with the same or similar descriptors (D, E); Matching (S3) the extracted image features of the two training images (A, B), wherein the image features are matched with the same or similar descriptors (D, E) to form image feature pairs; for each matched image feature pair, Extracting (S4) information of the key point (202, 203) from one of the two training images (A, B) and information of the descriptor (D, E) from the other of the two training images (A, B); Reconstructing (S5) one of the two training images (B), for which information of the respective key points (202, 203) of the matched image feature pairs was extracted, by the second deep neural network (204) on the basis of the extracted information of the respective key points (202, 203) and on the basis of the extracted information of the respective descriptors (D, E) of the matched image feature pairs; Quantifying (S6) differences between the reconstructed training image (B') and the provided one of the two training images (B) by means of a cost function (205); Adjusting (S7) parameters of the first and / or second deep neural network (200, 204) based on the quantified differences to optimize the cost function (205) to provide a trained machine learning model. [2] Method according to claim 1, wherein, in order to provide the trained machine learning model, steps S2 to S7 are repeated at least partially iteratively until a termination criterion or a threshold value of the cost function (205) is reached. [3] Method according to claim 1 or 2, wherein the cost function (205) determines an error between the reconstructed training image (B') and the provided one of the two training images (B) on the basis of the quantified differences. [4] Method according to claim 3, wherein the cost function (205) determines the error as a sum of pixel differences between the reconstructed training image (B') and the provided one of the two training images (B) as a photometric error. [5] Method according to one of the preceding claims, wherein the mutually corresponding key points (202, 203) are set at the same or similar positions of the two training images (A, B). [6] The method of any preceding claim, wherein the first deep neural network (200) comprises an encoder and the second deep neural network (204) comprises a decoder. [7] Method according to one of the preceding claims, wherein a plurality of two training images (A, B), each preferably representing pairs of images of the same or similar image content, is provided. [8] Method according to one of the preceding claims, wherein the processing (S2) of the two training images (A, B) by the first deep neural network (200) further comprises: calculating an image descriptor (206) of the one of the two training images (B); and wherein the reconstruction (S5) of the one of the two training images (B) is also carried out on the basis of the image descriptor (206) associated with the one of the two training images (B). [9] A method according to any one of the preceding claims, wherein each pixel of the two training images defines a keypoint. [10] Inference method for image feature extraction from images for relative pose determination of image contents, comprising: Providing two images, each preferably captured by an optical sensor, Extracting image features from the provided images by a machine learning model trained according to a method according to one of claims 1 to 9; relative pose determination of image contents of the two images on the basis of the extracted image features by the machine learning model trained according to a method according to one of claims 1 to 9. [11] Inference method according to claim 10 for environmental localization and / or for self-pose determination based on the relative pose determination of image contents for an automated driving function of a motor vehicle, an automated function of a drone or an automated function of a robot. [12] Device (100) for training a machine learning model having a first and a second deep neural network, the device (100) comprising an evaluation and / or computing device which is designed to carry out the following steps: Providing (S1) two training images (A, B) which at least partially have the same or similar image contents, wherein the two training images (A, B) depict the image contents from different perspectives and / or at different times; Processing (S2) the two training images (A, B) by the first deep neural network (200) such that image features are extracted from the two training images (A, B) by setting mutually corresponding key points (202, 203) in the two training images (A, B) and describing the mutually corresponding key points (202, 203) with the same or similar descriptors (D, E); Matching (S3) the extracted image features of the two training images (A, B), wherein the image features are matched with the same or similar descriptors (D, E) to form image feature pairs; for each matched image feature pair, Extracting (S4) information of the key point (202, 203) from one of the two training images (A, B) and information of the descriptor (D, E) from the other of the two training images (A, B); Reconstructing (S5) one of the two training images (B), for which information of the respective key points (202, 203) of the matched image feature pairs was extracted, by the second deep neural network (204) on the basis of the extracted information of the respective key points (202, 203) and on the basis of the extracted information of the respective descriptors (D, E) of the matched image feature pairs; Quantifying (S6) differences between the reconstructed training image (B') and the provided one of the two training images (B) by means of a cost function (205); Adjusting (S7) parameters of the first and / or second deep neural network (200, 204) based on the quantified differences to optimize the cost function (205) to provide a trained machine learning model. [13] Control unit (1000) for an automated driving function of a motor vehicle, an automated function of a drone, a robot and / or for an automated optical inspection of components and / or samples, wherein the control unit is designed to execute a machine learning model trained according to a method according to one of claims 1 to 9. [14] Computer program with program code to carry out at least parts of a method according to one of claims 1 to 11 when the computer program is executed on a computer. [15] Computer-readable data carrier with program code of a computer program for carrying out at least parts of a method according to one of claims 1 to 11 when the computer program is executed on a computer.