Methods for estimating monocular depth in an image

DE102024135328B4Undetermined Publication Date: 2026-06-25CARIAD SE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
CARIAD SE
Filing Date
2024-11-28
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Existing methods for estimating monocular depth from 2D image data are computationally intensive, require additional sensors like LiDAR, and are prone to faults, leading to inefficiencies and increased system complexity, while training is time-consuming and relies on manual annotation.

Method used

A method involving a training phase that uses a camera to capture source and target images, employing a neural network to calculate depth estimates based on odometry, epipolar scaling, and structure matrix, followed by an application phase where a trained network provides depth estimates for autonomous driving, without relying on additional sensors or manual data labeling.

Benefits of technology

The method optimizes computational resources, reduces reliance on additional sensors, enhances precision and robustness, and simplifies training, enabling efficient and accurate monocular depth estimation for autonomous driving applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for estimating the monocular depth in an image of a scene (1), a computer program product, a computer-readable data storage device, a control unit (ECU) and a machine (100).
Need to check novelty before this filing date? Find Prior Art

Description

The invention relates to a method for estimating the monocular depth in an image of a scene, a computer program product, a computer-readable data storage device, a control unit and a machine. Image data, e.g., from cameras and / or videos, can be used to support and / or implement driver assistance systems in machines, particularly vehicles. Information derived from such image data can also be used for the (at least partially) automated and / or autonomous driving of such machines. The following documents are examples of such work: LIU, Feng, et al. Unsupervised monocular depth estimation for monocular visual slam systems. IEEE Transactions on Instrumentation and Measurement, 2023, Vol. 73, pp. 1-13; ZHAO, Chaoqiang, et al. GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor Scenes. In: 2023 IEEE / CVF International Conference on Computer Vision (ICCV). IEEE, 2023, pp. 16163-16174. However, established methods and systems have drawbacks. Computing power for real-time implementation may be insufficient. Assessing (absolute and / or monocular) depth from 2D image data may require additional (reference) data, for example, through the use of LiDAR and / or RGB-D cameras. This can increase the system's weight and / or complexity. It can also reduce robustness if part of the system malfunctions and / or provides faulty data. Training can be time-consuming and / or require manual annotation / labeling of data, for example, to generate ground truth data. Consequently, an objective object of the present invention may be to overcome at least one of the aforementioned disadvantages (at least partially). In particular, it may be an objective to provide a method with optimized costs, speed, efficiency, precision, robustness, training data requirements, reference data requirements, and / or simplicity. The aforementioned problem is solved by a method with the features of the independent method claim, a computer program product with the features of the independent claim relating to a computer program product, a computer-readable storage medium with the features of the independent claim relating to a computer-readable storage medium, an electronic control unit with the features of the independent claim relating to an electronic control unit, and a machine with the features of the independent machine claim. Further features and details of the invention are presented in the dependent claims, the description, and the drawings.Features and details described in connection with the method according to the invention naturally also apply in connection with the computer-readable storage medium according to the invention and / or in connection with the electronic control unit according to the invention and / or in connection with the machine according to the invention, and vice versa, so that with regard to the disclosure, reference is always made or can be made to the individual aspects of the invention. In particular, the advantages described in connection with the first, second, third, fourth and / or fifth aspect also apply to the first, second, third, fourth and / or fifth aspect. The above-mentioned task is solved, according to a first aspect, by a method for estimating the (monocular and / or geometric and / or computed) depth in a (2D) image of a (recorded) scene, in particular a street scene, in front of a (physical) machine, in particular a vehicle, e.g. a car, wherein the method comprises a training phase and (subsequently / later) an application phase, wherein the training phase comprises: - capturing a source image of a scene in front of a machine at a first time point and a target image of the scene in front of the machine at a second time point by a camera of the machine, the second time point following the first time point, - wherein an electronic control unit, in particular of the machine, performs the following steps: ◯ computing a depth estimate based on the target image by a first network,• Calculating an odometry of the scene based on the source image and the target image by a second network, • Calculating a distorted source image based on the source image and at least the odometry, • Calculating an epipolar scaling by a third network based on the distorted source image, • Calculating a structure matrix based on the odometry and the epipolar scaling, • (Simultaneously) training the first network, the second network, and the third network, wherein the first network is trained at least on the basis of the structure matrix, the application phase comprising: - Providing the trained first network to an application machine (in particular the machine), - Capturing an application input image of an application scene in front of the application machine (in particular the machine) by an application camera (e.g. the camera) of the application machine (in particular the machine),- Calculating an application depth estimate by an electronic (application) control unit using the trained first network based on the application input image, wherein the application depth estimate is specific for the distance to at least one depicted object in the application input image; - Controlling the application machine by the electronic control unit of the application machine based on the application depth estimate. The method according to the first aspect can be at least partially computer-implemented and / or repeated. Advantageously, at least one of the described steps can be performed in the method, wherein the steps are preferably performed sequentially in the specified order or alternatively in any other arbitrary order, and individual steps can also be repeated. Preferably, the method can be performed before and / or (preferably) during operation or use of a machine and / or an electronic control unit. Operation can include (manual) driving, (at least partially) autonomous driving, and / or automated driving. The method can be performed at least partially during commissioning and / or maintenance.An electronic control unit can implement the process (at least partially), for example by executing the steps mentioned above and / or controlling and / or regulating corresponding components (e.g., steering unit and / or drive unit). The process can be used to optimize costs, speed, efficiency, precision, robustness, training data requirements, reference data requirements, and / or simplicity. The method can be configured to estimate and / or calculate a (monocular and / or geometric and / or calculated) depth in a (2D) image of a (recorded) scene, in particular a street scene, in front of a (physical) machine, in particular a vehicle, e.g. a car. For example, the machine can drive along a street and / or through a (street) scene. The method can include a training phase, which is performed before the (subsequent and / or later) application phase. In the training phase, the first, second, and / or third network can be trained. In the application phase, the trained first network can be used. Preferably, only the trained first network is used in the application phase. Preferably, the second and / or third network are used only during training, in particular by using the second and / or third network (in combination and / or as a planar parallax pipeline) as a teacher for the first network (during training). In the training phase, the machine can drive along a scene (e.g., along a road) and / or a plurality of scenes to acquire a source image (in particular, a plurality of source images) and a target image (in particular, a plurality of target images), which can preferably be used as training data.The training (training phase) can provide a trained first, second, and / or third network (as a result). The trained first network can be configured to provide an application depth estimation based on an (application) input image. Preferably, the source image(s) and the corresponding target image(s) are aligned, in particular planarly aligned. For example, the image planes are (approximately) parallel. The (multiple) source image(s), target image(s), distorted source image(s), (application) input image(s), structure matrix(s), (application) depth estimation(s), (application) certainty mask(s) and / or (application) planar surfaces may include 2D information, 2D pixel data and / or a 2D matrix (etc.), especially with identical resolution, e.g. Full HD or 244x244 pixels. The scene and / or the application scene can include a street scene in front of the (application) machine and / or a (multiple) object(s) that can preferably be captured by the camera and / or represented by the source image(s), target image(s) and / or (application) input image. Preferably, the steps of the training phase can be repeated, especially for a large number of (corresponding) source images, target images, distorted source images, structure matrix(s), depth estimates, certainty masks and / or planar surfaces, etc. The training can be based on this large number of images and / or data. Capturing, by a camera of a machine, a (2D) source image (in particular a plurality of source images) of a scene in front of the machine at a first time point and a (2D) target image (in particular a plurality of target images) of the (3D) scene in front of the machine at a second time point, wherein the second time point follows the first time point, in particular (only) one time interval later (e.g. 1 second). The first image and the second image can preferably be geometrically aligned and / or include planar orientations. The target image can depict the (same and / or 3D) scene at a later time point than the source image. For example, the machine is driving on a street with parked cars on the left and right sides. The source image can depict this (street) scene at a first time point, in particular wherein one parked car on the right side is fully visible (e.g. trunk and left side of the parked car).The target image can depict this (same street) scene at a second (immediately later and / or subsequent) point in time, where, in particular, the parked car on the right is no longer fully visible due to the distance traveled between the first and second points in time (e.g., the trunk is no longer visible, while parts of the left side of the parked car are still visible). In other words, the target image can depict the scene at a later point in time. The camera can be a video camera. The source image(s) and the target image(s) can be successive frames of a video (of the scene) recorded by the camera, where, for example, the first and second points in time are defined by a frame rate of the (video) camera.During the training phase, a multitude of source images and a multitude of target images can be acquired, each source image comprising a subsequent target image (so that pairs of source and target images are used for further processing). For the sake of simplicity, the invention is described using only one source image and one target image. However, the training phase can be carried out for and / or based on a multitude of source images, target images, and / or pairs of source and target images, preferably with the training steps (by the electronic control unit) being performed (separately) for each (pair of) source image and a (corresponding) training image. The electronic control unit may include data processing equipment and / or a computer. Training can be performed by the electronic control unit and / or an external computer, such as a cloud computing system and / or a (large) backend server (to improve speed and / or performance). The calculation of a (2D) depth estimate by a first network based (only) on the target image can be performed exclusively on the basis of the target image. A training objective can preferably be to train the first network to calculate a (monocular) depth of a (single and / or 2D) input / target image. Consequently, the first network can be configured to provide (as output) a depth estimate (preferably in the form of a 2D matrix), in particular a disparity and / or structure, of the scene that is specific to the scene and / or objects within the scene represented by the input / target image. In other words, each pixel of the depth estimate includes an (absolute) depth, a metric scale, and / or a distance to an object represented at a corresponding pixel position in the input / target image.For example, a specific pixel, which can be specific to the (distance to a) trunk of a parked car, could be 5 or 7 meters. The computation of an odometry of the scene by a second network based on the source and target images can include the computation of a rotation (e.g., using a rotation matrix, in particular a [3x3] rotation matrix) and / or a translation (e.g., using a translation matrix, in particular a [3x3] translation matrix) between the source and target images, in particular between an optical center of the source image and an optical center of the target image. Consequently, the odometry can be specific to (the geometry between) the source and target images, in particular the source image can be transformed (at least partially) into the target image and / or vice versa using the odometry. The second network can comprise a convolutional neural network and / or a ResNet-18 network. For example, the input, in particular the source image and / or the target image (e.g.,(after scaling), have a resolution of 224 x 224 pixels. The second network can comprise 18 layers. For example, the second network can comprise a network, specifically a ResNet-18 network, as used in: “Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770-778, 2016” and / or “Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel J Brostow. Digging into self-supervised monocular depth estimation. In Proceedings of the IEEE / CVF international conference on computer vision, pages 3828-3838, 2019. 1, 2, 3, 5, 6, 7, 8”. The second network can be used, in particular (additionally), during the application phase to calculate application odometry based on an application input image, an application source image, and / or an application target image. For example, the application source image can be acquired at a first application time point, and the application target image at a second (subsequent) application time point, preferably analogous to the training phase. This makes it possible to extract application odometry (as additional information) during the application phase. The calculation of a (2D) distorted source image based on the source image and at least the odometry can be performed by the electronic control unit and / or include distortion of the source image based on at least the odometry (and in particular the source image and / or the camera position / height [see below]). In other words, the source image is transformed by distortion. Therefore, the source image and the target image can include (adjacent) planar orientations of the scene. The calculation of an (epipolar) scale based on the (2D) distorted source image by a third network can include the calculation of a scale for the previously distorted source image, specifically with respect to the target image. The epipolar scaling can be specific to the scaling between the distorted source image and the target image. For example, the epipolar scaling can include at least two vectors representing a scale along the two dimensions of a 2D matrix. The training phase, in particular the calculation of odometry, a distorted image, epipolar scaling, and / or a structure matrix, can be based on (and / or utilize) a planar parallax geometry (as described below). This can provide metrically scaled values ​​(depth estimation) that can be obtained relative to a correctly oriented real-world planar scene. The fundamentals and / or implementation details can be described in: “Michal Irani and Prabu Anandan. Parallax geometry of pairs of points for 3d scene analysis. In Computer Vision-ECCV'96: 4th European Conference on Computer Vision Cambridge, UK, April 15-18, 1996 Proceedings, Volume I 4, pages 17-30. Springer, 1996. 3, 4, 5” and / or “Hao Xing, Yifan Cao, Maximilian Biber, Mingchuan Zhou, and Darius Burschka. Joint prediction of monocular depth and structure using planar and parallax geometry. Pattern Recognition, 130:108806, 2022.”3 "and / or" Haobo Yuan, Teng Chen, Wei Sui, Jiafeng Xie, Lefei Zhang, Yuan Li, and Qian Zhang. Monocular road planar parallax estimation. IEEE Transactions on Image Processing, 2023. 3, 4, 5 ". The acquisition process can provide a source image and a target image. The source image Is can be aligned with the target image It, preferably with respect to a planar surface π, in particular using planar parallax geometry and / or the second and / or third network. The planar surface can comprise the surface of the scene and / or the application scene, e.g., the street of the scene. The alignment can be achieved by mapping the points / pixels from the source image ps to corresponding points in the target image pter, in particular by planar homography Hs→t and / or the following equation (1): This equation describes how each point psim in the source image is transformed into a new position in the distorted source image. Hs→t can be calculated by matching the correspondences between the views of this planar surface, or if the plane and the odometry (relative pose) between It and Is are known, it can be calculated using the following equation (2): Here, an intrinsic camera matrix K of the camera can be used. The camera matrix K can include a projection matrix that describes the figure of the (pinhole camera) from 3D points of the scene to 2D points in the (source and / or target) image. Rs→t can include a rotation matrix. Ts→c can include a translation matrix. Consequently, the odometry can include and / or be defined by the rotation matrix and the translation matrix. The odometry, in particular the rotation matrix and the translation matrix, can describe a transformation from the optical center of the source view Os (which is specific to the source image) to the optical center of the target view Ot (which is specific to the target image). NT can include a normal vector of the planar surface π. The planar surface can include the ground (the floor) and / or the road beneath the machine.hc can include a height of the camera (center), particularly in relation to a lowest support point of the machine and / or above a ground / planar surface located below the machine (where the machine is standing and / or driving on / above the ground). The calculation can include the calculation of a residual parallax between any point in the distorted source image and any corresponding point in the target image. This can include, provide, and / or describe the relative motion, the depth (estimation) of objects within the scene, and / or the height of objects relative to the planar surface. In particular, this can be defined and / or calculated according to formula (3): Here, the residual parallax between any point in the distorted source image and any corresponding point in the target image can be described by pt. Here, γ encompasses the (3D) structure of the points. In particular, γp describes the structure of a corresponding point in the equation above and / or corresponds to where hp encompasses the height of the point from the planar surface and dp encompasses the depth of this point from the target image / point. Here, can encompass the epipole of the target image. Tz can encompass a translation in the z-direction, particularly with respect to the translation matrix (z-component) and / or between the source image and the target image. Preferably, the calculation is performed for Tz ≠ 0. Preferably the epipolar scaling includes the first part of equation (3) and / or the right-hand side of the equation, in particular excluding (pt- et). The calculation of this residual parallax (also called residual parallax flow) can be based on the differences in apparent motion between planar-oriented points in Is and the same points in It, in particular by providing a mathematical framework for understanding and manipulating visual depth cues. This is also described in the aforementioned publication by Irani et al. and / or Sawhney. 3D geometry from planar parallax. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 1994, pp. 929-934. IEEE, 1994. 3, 5. The calculation of a (2D) structure matrix based on odometry and epipolar scaling can provide a (second) depth estimate of the scene, particularly in addition to and / or independently (since it is a result of the second and third networks) of the first network and / or the depth estimate. The structure matrix can comprise a 2D matrix specific to a depth (of pixels and / or objects in the source image) and / or a height (of pixels and / or objects in the source image) with respect to a depth. The depth estimate and the structure matrix can (both) comprise a result of an (estimated and / or calculated) depth, particularly of pixels and / or objects in the scene and / or objects represented in the source image, target image, and / or distorted source image.Consequently, the second and / or third network can be used as a teacher (using the structure matrix) for the first network (which can be the student), in particular to optimize the calculation of the depth estimation by the first network based on the structure matrix and / or based on at least one of the five losses (see below). The training of the first, second, and third networks, with the first network being trained at least on the basis of the structure matrix, can be performed simultaneously and / or sequentially. This allows for the optimization of the weights of the first, second, and / or third networks in a simultaneous, efficient, and / or sequential manner. The training can be self-supervised. Alternatively, it can be a teacher-student model, where the second and third networks (which together can form a planar parallax pipeline) act as teachers for the first network (which acts as the student). Based on this instruction, the first network can provide improved and / or more accurate depth estimation. Preferably, only the (trained) first network, and in particular exclusively, is used during the application phase. Preferably, the method, particularly the training phase, does not use fixed data-specific depth binning to transform the network output into the predicted depth. Instead, residual flow binning can be used, which is intrinsically linked to the scene structure and ensures the adaptability of the depth range per frame. Preferably, the method does not use cost volumes, transformers, and / or similar architectures (which are computationally intensive). The application phase can be carried out after the training phase is completed and / or a trained initial network has been provided (and / or transferred to the application machine). Providing the trained first network to an application machine can (simply) involve using the same machine as in the training phase. Alternatively, it can involve transferring the trained first network to an application machine (a different and / or distinct machine), in particular an electronic application control unit, e.g., in another machine (e.g., a vehicle), which is preferably identical and / or analogous in design. Preferably, the application machine and the machine (used during the training phase) include the same camera and / or camera position, in particular the same camera height. Capturing an application input image of an application scene in front of the application machine using an application camera of the application machine, wherein the application scene preferably comprises and / or represents at least one object (or a plurality of objects). The application scene can comprise a street scene, in particular comparable to (or identical with) the (street) scenes during the training phase. The application camera can be identical to the camera. The application machine can be identical to the machine (see above). The application input image can be comparable to (or even identical with) the target image (or source image). Preferably, the application input image is a 2D matrix of image data captured by the application camera, e.g., of a street scene traversed by the application machine, particularly during the application phase. Preferably, during the application phase, the trained first network is used exclusively to provide an (application) depth estimation based solely on the application input image. In other words, the entire pipeline from the training phase cannot be used during the application. Preferably, only the trained first network can be used during the application phase. This can optimize costs, speed (due to reduced computational complexity), efficiency, precision (compared to, for example, other single-shot methods), robustness, the need for (additional) reference data (e.g., LiDAR), and / or simplicity. The calculation of an application depth estimate by an electronic application control unit using the trained first network, based on the application input image (which serves as input for the trained first network), can include the calculation of the application depth estimate, which can be specific to the (absolute and / or monocular) distance to at least one displayed object in the application input image. Therefore, the application depth estimate can comprise a 2D matrix and / or pixels that are specific to (correspond to) a (monocular) depth, (absolute) distance, and / or height of a (displayed) object and / or a corresponding pixel of an object in the application input image. The control of the application machine by the electronic (application) control unit of the application machine, based on application depth estimation, can include the control of the application machine based on (at least partially) automated and / or autonomous driving. The control can include the provision of a (driving) assistance function, in particular a (at least partially) automated and / or autonomous assistance function. For example, the control can include control based on the output of the first network, in particular based on application depth estimation. The control can, for example, include the control of a hardware system, in particular a steering unit, a chassis damping unit, a distance control system, and / or a braking system, in particular an anti-lock braking system.Preferably, the electronic control unit can control the hardware system, for example, by transmitting a control signal to an actuator of the hardware system (e.g., via a specific data link configured for transmission). For example, emergency braking can be performed based on a braking distance estimate, e.g., when the distance to an object that can be detected (by the electronic control unit) based on the braking distance estimate is small (compared to the current speed). Alternatively or additionally, the control can include providing the braking distance estimate to mapping and / or topology and / or navigation systems. For example, the application depth estimate can be used to enhance a virtual road scene displayed to a driver and / or provided to a driver assistance system and / or other machines (e.g., via the internet and / or car-to-car communication). Within the scope of the present invention, a method can be provided, wherein the method particularly includes, after acquisition, the following step: - transferring the source image and the target image from the camera to an electronic control unit of the machine (or an external computer for training), particularly via an intermediate data connection, and / or the method includes the following step, particularly after acquisition: - transferring the application input image from the application camera to an electronic application control unit of the application machine (or another computer), particularly via an intermediate application data connection. Within the scope of the present invention, a method can be provided in which the calculation of a (calculated) depth estimate based on the target image by a first network comprises: - providing a depth, in particular comprising an absolute distance to a depicted object of the (street) scene, for each pixel in the target image, and / or - calculating the depth estimate exclusively on the basis of the target image. Therefore, only the target image can be used as input for the first network. The depth estimation can comprise a 2D matrix of pixels corresponding to objects in the scene and / or specific to a distance from (represented) objects in the source image, (preferably) the target image, and / or the distorted source image. Within the scope of the present invention, a method can be provided in which the training of the first network, the second network and the third network is based in particular exclusively on the following: - the source image, in particular a plurality of source images, - the target image, in particular a plurality of target images, which preferably correspond to the plurality of source images (with respect to a first time point for each of the source images and a corresponding, later, subsequent and / or following second time point for each of the [corresponding] target images), and - a height of the camera above a lowest support point of the machine and / or above a floor below the machine (where the machine is on / above the floor and / or moving). The multitude of source images and the multitude of target images can also be derived from a stream of images from a video camera. For example, a machine travels along a street scene, and the camera continuously records images. These images can be used as source and target images. For instance, a source image at a first time point t = 1 second can be used in combination with a target image at a second time point t = 2 seconds. Then, the target image (at t = 2 seconds) is used as a source image (at t = 2 seconds) in combination with a subsequent target image (at t = 3 seconds), and so on. Source images can thus also be used as target images, and vice versa. Within the scope of the present invention, particularly during the application phase, a method can be provided in which an application depth estimation is calculated using the trained first network based on the application input image by an electronic application control unit, wherein the application depth estimation is specific for the distance to at least one object depicted in the application input image and is used exclusively on the basis of the application input image as input for the (trained) first neural network, in particular without using the (trained) second network, the third network and / or additional reference data, for example LIDAR data, further camera data and / or ground truth data.In particular, the method provides a way to calculate the monocular metric depth (images and / or information) of an (application) scene, especially without ground truth, LIDAR and / or position information. In connection with the present invention, a method can be provided in which the calculation of a distorted source image is based on the source image and at least the odometry, and additionally on the basis of the training image and / or a camera position, in particular a height of the camera above a lowest support point of the machine and / or above a floor located below the machine, wherein in particular a transformation, in particular a homography (layer), of the source image is calculated, in particular to calculate / output the distorted source image. The calculation of a distorted source image can be performed as described above and / or by equations 1 and / or 2. Therefore, homography can be used based on the camera (matrix), the camera height, and / or odometry, in particular rotation (matrix) and / or translation (matrix). Within the scope of the present invention, a method can be provided in which the calculation of the depth estimation based on the target image is carried out by the first network by inputting the target image into a first part of the first network, which in particular comprises a first encoder [mono] which computes an output of the first part of the first network, which is used as input for a second part of the first network, which in particular comprises a first [depth] decoder which is configured to calculate the depth estimation. Within the scope of the present invention, a method can be provided in which the calculation of an epipolar scaling based on the distorted source image by the third network is carried out by inputting the distorted source image into a first part of the third network, which in particular comprises a second (distortion) encoder, wherein the output of the first part of the third network is used as input for a second part of the third network, which in particular comprises a second (flowscale) decoder, wherein the input preferably also includes the output of the first part of the first network, wherein the second part of the third network is configured to calculate the epipolar scaling. The first part of the third network, which in particular includes a second (distortion) encoder, can include a ResNet-50, especially as described in: “Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770-778, 2016.” The first part of the third network can be pre-trained on ImageNet, in particular as described in: “Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115:211-252, 2015. 5” The first part of the first network can preferably be identical to the first part of the third network. Preferably, the output of the first part of the first network and the output of the first part of the third network are combined, in particular concatenated, and forwarded to the second part of the third network. The second part of the third network can comprise a neural network, in particular a CNN. Preferably, the second part of the third network comprises a set of CNNs configured specifically for upsampling the input and predicting the output and / or for epipolar scaling. This can be implemented as described in “Godard et al.” (see above) and / or “Jamie Watson, Oisin Mac Aodha, Victor Prisacariu, Gabriel Brostow, and Michael Firman. The temporal opportunist: Self-supervised multi-frame monocular depth. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp. 1164–1174, 2021. 2, 5, 6, 7.” The second part of the first network can be identical to and / or configured accordingly with the second part of the third network. In particular, the second part of the first network can comprise a neural network, especially a CNN. Preferably, the second part of the first network comprises a set of CNNs configured specifically for upsampling the input and predicting a depth estimate. This can be done as described in Godard et al. (see above) and / or Jamie Watson et al. (see above). However, the second part of the first network can (exclusively) use the target image and / or the application input image (during the application phase) as input. The first network N1 can provide a depth estimation based on an input of the target image (training phase) and / or the application input image (application phase). Therefore, it can be defined and / or calculated by the following equation (4): Here, the calculated depth can be derived from the depth estimate generated by the first network. The depth estimate can include a disparity map, based on which the depth can be calculated. The second and third networks (combined) can provide a planar parallax pipeline Θpp. Therefore, it can be defined and / or calculated by the following equation (5): Here, st can provide the output of the third (and second) network. The output of the third (and second) network and / or Θpp (each) can include the epipolar scaling, in particular a single value per pixel, which comprises an (epipolar) scaling value. The epipolar scaling (value) can be a representative of the first term of equation (3). The second part of the third network, specifically the second (flow scaling) decoder, can output the epipolar scaling (values). Consequently, the epipolar scaling can comprise a 2D matrix with epipolar scaling values ​​for each pixel. The output of the second (flowscale) decoder, in particular the epipolar scaling values ​​st, can be assigned to specific bins that can be adapted for different image resolutions to provide scaled epipolar scaling values ​​St. The scaled epipolar scaling values ​​can be identical, representative, and / or correspond to the first term of equation (3). Using equation (3), the residual parallax can be calculated (6): and / or Using the residual parallax, the structure matrix y can be calculated. Therefore, for specific pixels / points encompassing a certain Tzund h, the following can be calculated (7): Consequently, the output of the planar parallax pipeline after scaling can be the pixel offset along the epipolar line. The structure (matrix) y and / or the components of the structure matrix can be derived via equation (7). An (absolute) depth (and / or height) can be calculated and / or derived using equation (8): Here, p can represent a specific pixel. K can encompass the camera matrix. This can be calculated and / or proven in: “Haobo Yuan, Teng Chen, Wei Sui, Jiafeng Xie, Lefei Zhang, Yuan Li, and Qian Zhang. Monocular road planar parallax estimation. IEEE Transactions on Image Processing, 2023. 3, 4, 5” Therefore, the depth estimation of the first network, in particular the calculated depths, can be compared with the structure matrix, in particular the depths derived from it, to calculate a loss and / or to use the second and third networks (combined) as teachers for the first network as students. Based on and / or, the loss(s) can be calculated. The method can provide novel synthesized images, e.g., from adjacent views. These can be calculated using equation (9), in particular to calculate the photometric reprojection losses: Here, `proj()` can include the resulting coordinates of the projected depths. Furthermore, 〈 〉 include a bilinear sampling operator that may be locally subdifferentiable. R and T may include the odometry, the rotation matrix, and / or the translation matrix, particularly between the synthesized image and The new views can (also) be constructed using residual parallax and / or using equation (10): Within the scope of the present invention, a method can be provided in which providing the trained first network to an application machine comprises transferring the trained first network, in particular a parameterization (e.g., the weights optimized during training) of the trained first network, from the machine to the application machine, particularly using a transmission data connection, preferably the internet. Tz can be derived from the output of the second network and / or the translation matrix (in particular the z-component). The height of the camera hcs can be measured and / or known and / or stored in the electronic control unit (e.g., during design). Within the scope of the present invention, a method can be provided in which the training includes calculating a certainty mask based on the structure matrix. The certainty mask can be used to reduce noise around the epipole, which can occur particularly when the term "(pt-et)" around the epipole may be minimal and the parallax negligible, potentially disrupting the learning process and / or leading to incorrect weights. The certainty mask can be used and / or applied pixel by pixel, specifically to the structure and / or the depth calculated from the structure. Consequently, the certainty mask can eliminate pixels resulting from errors around the epipole by including only those pixels reconstructed by the same method. The certainty mask Mcert can be defined and / or calculated using equation (11): Here, ε represents a threshold that can be estimated empirically. The last part ε can be used to measure the accuracy of the calculated homography, since the planar road surface can be detected in equation (2). A flat-area detector can be used for this purpose, inspired by the following source: “Zhenheng Yang, Peng Wang, Yang Wang, Wei Xu and Ram Nevatia. Lego: Learning edge with geometry all at once by watching videos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 225-234, 2018. 5, 7” The flat area can be identified by calculating the surface normal vectors of each pixel based on its neighboring points. For each pixel pi, the normal of this vector can be calculated by taking the cross product of the vector pi with its neighbors. First, pi can be projected into the three-dimensional world using the predicted depth Di, regardless of whether this is the student's or the teacher's perspective, since they converge in the static road area, so pi = Di · K⁻¹ · pi. Then, the average normal vector is calculated using the cross products of neighboring vectors. Alternatively and / or additionally, a road mask Mflat can be used and / or calculated using equation (12): Here, cossim can be the cosine similarity between vectors, and τ can be the threshold for determining whether it is a planar surface. Navg can be an average normal vector, particularly at a pixel position (i, j). To focus on the relevant area of ​​the image representing the road, a trapezoidal mask centered on the image can isolate the road by including only central pixels with near-zero structure, especially using |γ| ≤ 0.05. This mask can selectively exclude only the road plane from other flat areas. This can be used as a detector for flat surfaces, particularly when calculating the uncertainty mask. In connection with the present invention, a method can be provided in which the training comprises simultaneous training of the first network, the second network and the third network and / or training of the first network as a pupil based on the second network and / or the third network as a teacher. Therefore, the structure matrix, and in particular a depth estimation based on the structure matrix, can be used by the second and third networks (as teachers) to train the first network (as students), preferably enabling the first network to determine the depth information. In other words, the second and third networks preferably teach the first network about the geometry, static structure details, correct (epipolar) scaling, and / or masking (of dynamic objects, especially using the certainty mask). This approach can provide improved robustness and / or adaptability to different scenarios. Five different losses can be used for training, including Lhomo, Lmono, Lpp, Lcons, and Lres. These losses can be minimized and / or optimized (simultaneously) during training. Firstly, a photometric reprojection loss, which is used in some losses, L1 and SSIM according to equation (13): Dies wird auch beschrieben in:„ Clément Godard, Oisin Mac Aodha, and Gabriel J Brostow. Unsupervised monocular depth estimation with left-right consistency. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 270-279, 2017. 1, 2, 3, 6, 7 “,und / oder„ Clément Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow. Digging into self-supervised monocular depth estimation. In Proceedings of the IEEE / CVF international conference on computer vision, pages 3828-3838, 2019. 1, 2, 3, 5, 6, 7, 8 “und / oder„ Jamie Watson, Oisin Mac Aodha, Victor Prisacariu, Gabriel Brostow, and Michael Firman. The temporal opportunist: Self-supervised multi-frame monocular depth. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 1164-1174, 2021. 2, 5, 6, 7 “und / oder„ Hang Zhao, Orazio Gallo, luri Frosio, and Jan Kautz. Loss functions for image restoration with neural networks.IEEE Transactions on computational imaging, 3(1):47-57, 2016. 6”. Homography loss can be the loss responsible for an accurate estimation of road planar homography, and it is also the one that helps ensure the pose scale is as close as possible to the ground-truth pose information. It helps the planar surface resemble the planar surface of the road. This forces the correct orientation of the road planar surface. The loss can be calculated using equation (14): Here, pe can be the projection loss from equation (13). N can be the number of pixels. The loss can be calculated for all source images and / or distorted source images. This can also be done based on the following: “Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel J Brostow. Digging into self-supervised monocular depth estimation. In Proceedings of the IEEE / CVF international conference on computer vision, pages 3828-3838, 2019. 1, 2, 3, 5, 6, 7, 8” The losses Lmono and Lpp can (similarly) minimize the reprojection of the depth estimation, whether by the first network or the planar parallax pipeline. A key difference, however, can be that the depth is masked by the certainty mask due to planar parallax, as is preferably done for Lmono by equation (15): Here, automatic masking is used with the mask Mauto, which is described in particular in: “Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel J Brostow. Digging into self-supervised monocular depth estimation. In Proceedings of the IEEE / CVF international conference on computer vision, pages 3828-3838, 2019. 1, 2, 3, 5, 6, 7, 8”. Lpp can be defined and / or calculated according to equation (16): A residual loss Lres can quantify a computational error between the third network and the planar parallax pipeline. Minimizing this residual loss ensures that the epipolar scaling from the third network and / or the planar parallax pipeline represents the correct epipolar scaling. The residual loss Lres can be calculated using equation (17): The consistency loss Lcons can be configured to optimize the depth estimation by the first network by adjusting it to better match the structure matrix, particularly the depth calculated from the structure matrix. Therefore, a consistency mask can be used, inspired in particular by: “Jamie Watson, Oisin Mac Aodha, Victor Prisacariu, Gabriel Brostow and Michael Firman. The temporal opportunist: Self-supervised multi-frame monocular depth. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 1164-1174, 2021. 2, 5, 6, 7” The consistency mask can be used to check the pixels where the depth estimate predicted by the first network and the depth based on the structure matrix agree up to a certain threshold δ. Therefore, equation (18) can be computed: Furthermore, equation (19) can be calculated and / or minimized during training: Preferably, a normalized depth D can be used. This can optimize accuracy and / or prevent inaccuracies due to possible differences in scale. The training can involve using a total loss of Lmono + Lres + Lpps for the first 5 epochs. Odometry can be optimized in conjunction with this total loss. Then, the nomography loss Lhomo can be added to correctly align the planar surface and / or road. After 20 epochs, the second and third networks (planar parallax pipeline) can be frozen. Then, Lcons can be used additionally, specifically to mask Lmono with Mstatic to eliminate dynamic objects, so that the total loss becomes Lmono + Lhomo + Lcons. The procedure and / or the training can be performed and / or validated based on publicly available datasets, e.g.: - KITTI: “Andreas Geiger, Philip Lenz and Raquel Urtasun. Are we ready for autonomous driving? The Kitti Vision Benchmark Suite. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3354-3361, 2012. 1, 6, 7, 8” and / or - Cityscapes The first, second, and / or third network can be trained with an input and / or output resolution of 640x192. The same procedure and / or training, in particular the same augmentation, can be used as described in: “Clément Godard, Oisin Mac Aodha, and Gabriel J Brostow. Unsupervised monocular depth estimation with left-right consistency. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 270–279, 2017. 1, 2, 3, 6, 7”, and / or “Clément Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow. Digging into self-supervised monocular depth estimation. In Proceedings of the IEEE / CVF international conference on computer vision, pages 3828–3838, 2019. 1, 2, 3, 5, 6, 7, 8”. An Adam optimizer can be used for all epochs with a learning rate of 10⁻⁴. The following values ​​can be used for the hyperparameters: α = 0.85, δ = 0.2, ε = 5, and τ = cos(3°). To map the sigmoid output to the correct flow scales for the second and / or third network, fmin can be used. = -100 and f max A value of 100 can be used for the input resolution. For the first network, a sigmoid output can be used, where the output can be mapped to the depth range of the teacher and, in particular, limited to 80 m. This can reduce the data and / or also provide comparability with other methods. According to a second aspect of the invention, the above-mentioned problem is solved by a computer program product comprising instructions which, when the computer program product is executed by a computer, cause the computer to implement the method according to the first aspect of the invention. This means that, with regard to the computer program product according to the second aspect, the same technical advantages can be realized as have already been described above for the method according to the first aspect of the invention. According to a third aspect of the invention, the above-mentioned problem is solved by a computer-readable storage medium in which instructions are stored which, when executed by a computer, cause the computer to implement the method according to the first aspect of the invention. This means that, with regard to the computer-readable storage medium according to the third aspect, the same technical advantages can be realized as have already been described above for the method according to the first aspect of the invention and / or the computer program product according to the second aspect of the invention. According to a fourth aspect of the invention, the above-mentioned problem is solved by an electronic control unit comprising a computing unit and / or a storage unit in which instructions are stored which, when executed at least partially by the computing unit, implement a method according to the first aspect of the invention. This means that with regard to the electronic control unit according to the fourth aspect, the same technical advantages can be realized as have already been described above for the method according to the first aspect of the invention and / or the computer program product according to the second aspect of the invention and / or the computer-readable storage medium according to the third aspect of the invention. According to a fifth aspect of the invention, the above-mentioned problem is solved by a machine comprising an electronic control unit according to the fourth aspect. This means that, with regard to the machine according to the fifth aspect, the same technical advantages can be realized as have already been described above for the method according to the first aspect of the invention and / or the computer program product according to the second aspect of the invention and / or the computer-readable storage medium according to the third aspect of the invention and / or the electronic control unit according to the fourth aspect of the invention. Further technical features, advantages, and details of the invention are disclosed by the following description of the figures. The figures provide a detailed description of possible embodiments of the present invention. Therefore, the features described by the claims and the description can be implemented individually or in any combination. The following exemplary description includes: Fig. 1 a method, Fig. 2 an overview of the method / pipeline(s), Fig. 3 an overview of the method / pipeline for application, Fig. 4 a machine, and Fig. 5 an application machine. In the following figures, identical reference numerals are used for identical (or corresponding) features, in particular for different embodiments of the invention. Fig. 1 shows a method for estimating the monocular depth in an image of a scene 1, in particular a street scene, in front of a machine 100, in particular a vehicle, wherein the method comprises a training phase I and an application phase II, wherein the training phase I comprises: - capturing by a camera 10 of the machine 100 a source image I_s of a scene 1 in front of a machine 100 at a first time t1 and a target image I_t of the scene 1 in front of the machine 100 at a second time t2, wherein the second time t2 follows the first time t1, - wherein an electronic control unit ECU, in particular of the machine 100, performs the following steps: o calculating by a first network N1 a depth estimate I_depth based on the target image I_t, o calculating by a second network N2 an odometry Odo of the scene 1 based on the source image I_s and the target image I_t,• Compute 140 a distorted source image I_s_warp based on the source image I_s and at least the odometry Odo, • Compute 150 by a third network N3 an epipolar scaling scal_epi based on the distorted source image I_s_warp, • Compute 160 a structure matrix γ based on the odometry Odo and the epipolar scaling scal_epi, • Train 170 the first network N1, the second network N2 and the third network N3, wherein the first network N1 is trained at least on the basis of the structure matrix γ, wherein the application phase II comprises: - Provide 210 the trained first network N1 to an application machine 100', - Capture 220 an application input image I_in' of an application scene 1' in front of the application machine 100' by an application camera 10' of the application machine 100',- Calculate 230 an application depth estimate I_depth' by an electronic application control unit ECU' using the trained first network N1 based on the application input image I_in', where the application depth estimate I_depth' is specific for the distance to at least one imaged object in the application input image I_in', - Control 240 the application machine 100' by the electronic control unit ECU of the application machine 100' based on the application depth estimate I_depth'., Within the scope of the present invention, a method can be provided, wherein the method, in particular after acquisition 110, comprises the following step: - Transferring 115 of the source image I_s and the target image I_t from the camera 10 to an electronic control unit ECU of the machine 100, in particular via an intermediate data connection, and / or the method comprises the following step, in particular after acquisition 220: - Transferring 225 of the application input image I_in' from the application camera 10' to an electronic application control unit ECU' of the application machine 100', in particular via an intermediate application data connection. Within the scope of the present invention, a method can be provided in which the calculation of a depth estimate I_depth based on the target image I_t by a first network N1 comprises: - providing a depth, in particular comprising an absolute distance to a depicted object of the scene 1, for each pixel in the target image I_t, and / or - calculating the depth estimate I_depth exclusively on the basis of the target image I_t. Within the scope of the present invention, a method can be provided in which the training 170 of the first network N1, the second network N2 and the third network N3 is based in particular exclusively on the following: - the source image I_s, in particular a plurality of source images I_s, - the target image I_t, in particular a plurality of target images I_t, which preferably correspond to the plurality of source images I_s, and - a height h_s of the camera 10 above a lowest support point of the machine 100 and / or above a floor located below the machine 100. Within the scope of the present invention, a method can be provided in which an application depth estimation I_depth' is calculated by an electronic application control unit ECU' using the trained first network N1 on the basis of the application input image I_in', wherein the application depth estimation I_depth' is specific for the distance to at least one imaged object in the application input image I_in', is performed exclusively on the basis of the application input image I_in' used as input for the first neural network N1, in particular without using the second network N2, a third network N3 and / or additional reference data, for example LIDAR data, further camera data and / or ground truth data. Fig. 2 shows an overview of the process and / or the pipeline(s), especially for training phase I. Fig. 2 shows that within the scope of the present invention a method can be provided in which the calculation 140 of a distorted source image I_s_warp is based on the source image I_s and at least the odometry Odo, as well as additionally on the basis of the training image I_t and / or a camera position, in particular a height h_s of the camera 10 above a lowest support point of the machine 100 and / or above a floor located below the machine 100, wherein in particular a transformation, in particular a homography, of the source image I_s is calculated in order to output the distorted source image I_s_warp. Within the scope of the present invention, a method can be provided in which the calculation 120 of the depth estimate I_depth by the first network N1 on the basis of the target image I_t is carried out by inputting the target image I_t into a first part N1_1 of the first network N1, which in particular comprises a first encoder which computes an output of the first part N1_1 of the first network N1, which is used as input for a second part N1_2 of the first network N1, which in particular comprises a first decoder which is configured to compute the depth estimate I_depth. Within the scope of the present invention, a method can be provided in which the calculation of an epipolar scaling scal_epi by the third network N3 on the basis of the distorted source image I_s_warp is carried out by inputting the distorted source image I_s_warp into a first part N3_1 of the third network N3, which in particular comprises a second encoder, wherein the output of the first part N3_1 of the third network N3 is used as input for a second part N3_2 of the third network N3, which in particular comprises a second decoder, wherein the input preferably also includes the output of the first part N1_1 of the first network N1, wherein the second part N3_2 of the third network N3 is configured to calculate the epipolar scaling scal_epi. Within the scope of the present invention, a method can be provided wherein the provision 210 of the trained first network N1 for an application machine 100' comprises the transfer 209 of the trained first network N1, in particular a parameterization of the trained first network N1, from the machine 100 to the application machine 100', in particular using a transmission data connection, preferably the Internet. Within the scope of the present invention, a method can be provided in which the training 170 comprises: - simultaneous training of the first network N1, the second network N2 and the third network N3 and / or - training of the first network N1 as a pupil based on the second network N2 and / or the third network N3 as a teacher. Within the scope of the present invention, a method can be provided in which the training 170 comprises the calculation 171 of a certainty mask M_cert based on the structure matrix γ. Fig. 3 shows the process and / or the pipeline, particularly in application phase II. As can be seen from Fig. 3, particularly in comparison to Fig. 2 and / or training phase I, only the (trained) first network N1 is used during application phase II. Consequently, an application input image I_in', which can be captured by an application camera 10', is used as input for the (trained) first network N1 (which could be referred to as N1'). List of reference symbols 1 Scene 1' Application scene 10 Camera 10' Application camera 100 Machine 100' Application machine 110 Acquisition of a source image 115 Transmission of the source and target images 120 Calculation of a depth estimate 130 Calculation of an odometry 140 Calculation of a distorted image 150 Calculation of an epipolar scaling 160 Calculation of a structure matrix 170 Training of the first network, the second network, and the third network 171 Calculation of a certainty mask 209 Transmission of the trained first network to an application machine 210 Provision of the trained first network 220 Acquisition of an application input image 225 Transmission of the application input image 230 Calculation of an application depth estimate 240 Control of the application machine based on the application depth estimate ECU Electronic application control unit ECU' Electronic control unit for the application CU Computing unit MU Storage unit h_s Height of the camera Hs→t HomographyI Training Phase II Application Phase I_depth Depth Estimation I_depth' Application Depth Estimation I_in' Application Input Image I_s Source Image I_s_warp Distorted Source Image I_t Target Image M_cert Certainty Mask N1 First Network N1_1 First Part of First Network N1_2 Second Part of First Network N2 Second Network N3 Third Network N3_1 First Part of Third Network N3_2 Second Part of Third Network Odo Odometry scal_epi Epipolar Scaling t1 First Time Point t2 Second Time Point Residual parallax γ structure matrix π planar surface

Claims

Method for estimating the monocular depth in an image of a scene (1), in particular a street scene, in front of a machine (100), in particular a vehicle, wherein the method comprises a training phase (I) and an application phase (II), wherein the training phase (I) comprises: - capturing (110) by a camera (10) of the machine (100) a source image (I_s) of a scene (1) in front of a machine (100) at a first time (t1) and a target image (I_t) of the scene (1) in front of the machine (100) at a second time (t2), wherein the second time (t2) follows the first time (t1), - wherein an electronic control unit (ECU), in particular of the machine (100), performs the following steps: o Computing (120) a depth estimate (I_depth) by a first network (N1) based on the target image (I_t), o Computing (130) an odometry (Odo) of the scene (1) by a second network (N2) based on the source image (I_s) and the target image (I_t),o Computation (140) of a distorted source image (I_s_warp) based on the source image (I_s) and at least the odometry (Odo), o Computation (150) of an epipolar scaling (scal_epi) by a third network (N3) based on the distorted source image (I_s_warp), o Computation (160) of a structure matrix (γ) based on the odometry (Odo) and the epipolar scaling (scal_epi), o Training (170) of the first network (N1), the second network (N2) and the third network (N3), wherein the first network (N1) is trained at least on the basis of the structure matrix (γ), wherein the application phase (II) comprises: - Provisioning (210) of the trained first network (N1) to an application machine (100'), - Acquisition (220) of an application input image (I_in') of an application scene (1') in front of the application machine (100') by an application camera (10') of the application machine (100'),- Calculating (230) an application depth estimate (I_depth') by an electronic application control unit (ECU') using the trained first network (N1) based on the application input image (I_in'), wherein the application depth estimate (I_depth') is specific for the distance to at least one imaged object in the application input image (I_in'), - Controlling (240) the application machine (100') by the electronic control unit (ECU) of the application machine (100') based on the application depth estimate (I_depth'). The method according to claim 1, characterized in that the method, in particular after acquisition (110), comprises the following step: - transmitting (115) the source image (I_s) and the target image (I_t) from the camera (10) to an electronic control unit (ECU) of the machine (100), in particular via an intermediate data connection, and / or the method, in particular after acquisition (220), comprises the following step: - transmitting (225) the application input image (I_in') from the application camera (10') to an electronic application control unit (ECU') of the application machine (100'), in particular via an intermediate application data connection. Method according to claim 1 or 2, characterized in that the calculation (120) of a depth estimate (I_depth) on the basis of the target image (I_t) by a first network (N1) comprising: - providing a depth, in particular comprising an absolute distance to a depicted object of the scene (1), for each pixel in the target image (I_t), and / or - calculating the depth estimate (I_depth) exclusively on the basis of the target image (I_t). Method according to one of the preceding claims, characterized in that the training (170) of the first network (N1), the second network (N2) and the third network (N3) is in particular based exclusively on: - the source image (I_s), in particular a plurality of source images (I_s), - the target image (I_t), in particular a plurality of target images (I_t), which preferably correspond to the plurality of source images (I_s), and - a height (h_s) of the camera (10) above a lowest support point of the machine (100) and / or above a floor located below the machine (100). A method according to one of the preceding claims, characterized in that an application depth estimation (I_depth') is calculated by an electronic application control unit (ECU') using the trained first network (N1) on the basis of the application input image (I_in'), wherein the application depth estimation (I_depth') is specific for the distance to at least one imaged object in the application input image (I_in') and is performed exclusively on the basis of the application input image (I_in') (N1) used as input for the first neural network, in particular without using the second network (N2), the third network (N3) and / or additional reference data, for example LIDAR data, further camera data and / or ground truth data. Method according to one of the preceding claims, characterized in that the calculation (140) of a distorted source image (I_s_warp) is based on the source image (I_s) and at least the odometry (Odo) and additionally on the basis of the training image (I_t) and / or a camera position, in particular a height (h_s) of the camera (10) above a lowest support point of the machine (100) and / or above a floor located below the machine (100), wherein in particular a transformation, in particular a homography, of the source image (I_s) is calculated in order to output the distorted source image (I_s_warp). Method according to one of the preceding claims, characterized in that the calculation (120) of the depth estimation (I_depth) by the first network (N1) on the basis of the target image (I_t) is carried out by an input of the target image (I_t) into a first part (N1_1) of the first network (N1), which in particular comprises a first encoder which computes an output of the first part (N1_1) of the first network (N1), which is used as input for a second part (N1_2) of the first network (N1), which in particular comprises a first decoder which is configured to compute the depth estimation (I_depth). A method according to one of the preceding claims, characterized in that the calculation (150) of an epipolar scaling (scal_epi) by the third network (N3) on the basis of the distorted source image (I_s_warp) is carried out by inputting the distorted source image (I_s_warp) into a first part (N3_1) of the third network (N3), which in particular comprises a second encoder, wherein the output of the first part (N3_1) of the third network (N3) is used as input for a second part (N3_2) of the third network (N3), which in particular comprises a second decoder, wherein the input preferably also includes the output of the first part (N1_1) of the first network (N1), wherein the second part (N3_2) of the third network (N3) is configured to calculate the epipolar scaling (scal_epi). Method according to one of the preceding claims, characterized in that the provision (210) of the trained first network (N1) for an application machine (100') comprises the transfer (209) of the trained first network (N1), in particular a parameterization of the trained first network (N1), from the machine (100) to the application machine (100'), in particular using a transmission data connection, preferably the Internet. Method according to one of the preceding claims, characterized in that the training (170) comprises the calculation (171) of a certainty mask (M_cert) based on the structure matrix (γ). Method according to one of the preceding claims, characterized in that the training (170) comprises the following: simultaneous training of the first network (N1), the second network (N2) and the third network (N3) and / or training of the first network (N1) as a student based on the second network (N2) and / or the third network (N3) as a teacher. Computer program product comprising instructions which, when the computer program product is executed by a computer, cause the computer to execute the method according to any one of the preceding claims. A computer-readable storage medium on which instructions are stored which, when executed by a computer, cause the computer to execute the method according to any of the preceding claims. Electronic control unit (ECU) comprising a computing unit (CU) and / or a storage unit (MU) in which instructions are stored which, when executed at least partially by the computing unit (CU), implement a method according to any of the preceding claims. Machine (100) with an electronic control unit (ECU) according to the preceding claim.