Method for determining three-dimensional dimension information of a target object, motor vehicle, computer program and electronically readable data carrier
By determining the orientation of a target object's vertically oriented side in a camera image and using a recording geometry model and ground plane, three-dimensional cuboid representations are efficiently reconstructed from two-dimensional images, addressing inefficiencies in existing methods and enabling accurate automated driving applications.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- AUDI AG
- Filing Date
- 2022-05-17
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for deriving three-dimensional extent information from two-dimensional camera images in automotive applications, such as for partially automated driving, either fail to provide sufficient information or require complex three-dimensional annotations and additional sensors, making them inefficient and costly.
Determine the orientation of a target object's vertically oriented side in the camera image and use this information, along with a recording geometry model and ground plane, to reconstruct a three-dimensional cuboid representation of the object, utilizing intrinsic and extrinsic camera calibration and optional additional sensors for precise ground plane determination.
Enables robust and cost-effective generation of three-dimensional representations of target objects from two-dimensional images with minimal computational effort, providing accurate position, orientation, and size information for automated driving functions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for determining three-dimensional dimension information of a target object, in particular a vehicle, from a two-dimensional camera image showing the target object, which is recorded with a camera that is intrinsically and extrinsically calibrated with respect to a camera-carrying ego-object, in particular a motor vehicle, wherein the three-dimensional dimension information is determined from ground information relating to the three-dimensional position of the ground plane under the target object relative to the camera and from two-dimensional evaluation information, determined by evaluating the camera image using an evaluation algorithm, which describes at least one rectangle enclosing the target object in the camera image. The invention further relates to a motor vehicle, a computer program, and an electronically readable data carrier.
[0002] For a variety of sensor-based applications, particularly in the automotive sector, information about the position and extent of target objects within the sensor's detection range is to be derived from sensor data. A cost-effective option for such a sensor is a two-dimensional camera. However, since this only provides two-dimensional images, research and development also focuses on methods for deriving three-dimensional extent information for at least one target object from these two-dimensional images.
[0003] The detection of target objects in road traffic, such as other vehicles and their position, orientation, and size, is a particularly important area of application, especially with regard to applications that aim to enable at least partially automated driving of vehicles, particularly motor vehicles. If the vehicle has only a single, outward-facing camera, such as a front camera, current approaches can locate the target objects either in the two-dimensional camera image or in three-dimensional space. For both approaches, the use of machine learning methods, such as trained functions like neural networks, has already been proposed. However, this requires large amounts of data to achieve the necessary performance and reliability.
[0004] Regarding many driving functions, especially those enabling at least partially automated vehicle control, existing methods and approaches either fail to provide sufficient information (2D target object recognition) or require a large number of precise annotations of target objects in three-dimensional space during development (3D object recognition). Such three-dimensional annotations for training data are very complex and often only obtainable by incorporating additional sensors, such as laser scanners. Two-dimensional annotations, on the other hand, are simpler and faster to create. In this context, it has been proposed to use annotated rectangles around the target objects or to perform a pixel-by-pixel mapping to target objects (instance segmentation). However, these 2D annotations initially provide no information about the position and orientation of the target object in three-dimensional space.
[0005] US 2020 / 0160033 A1 describes systems and methods for generating three-dimensional representations from monocular two-dimensional images. A monocular 2D image is captured by a camera and processed to generate one or more property maps whose properties can include these properties or object labels.
[0006] Regions of interest that correspond to vehicles in the image are identified, and a lifting function is applied to each region to determine values such as height and width, camera distance, and orientation. An eight-point box, or cuboid, is generated, which is a 3D representation of the vehicle contained within the region of interest. These 3D representations can be used, for example, for route planning, collision avoidance, or as training data. Information about the subsurface, terrain, roads, surfaces, and other properties can be incorporated.
[0007] A method for estimating the relative position of an object in the vicinity of a vehicle is described in EP 3 594 902 A1. This method determines the relative position based on a two-dimensional camera image. First, an object contour of the object is determined within the camera image, and at least one digital object template is generated that represents the object based on this contour. The at least one object template is projected from different positions onto an image plane of the camera image, with each forward-projected object template generating a corresponding two-dimensional contour suggestion. The correct position is then determined by comparing these contour suggestions with the object's actual contour. The object contour can be a two-dimensional bounding box, i.e., a rectangle. The object template can be a three-dimensional bounding box, i.e., a cuboid, representing the object.It is also conceivable that each object template represents a specific object type, size, and spatial orientation, such as a vehicle, pedestrian, or cyclist. The number of possible positions for forward projections can be reduced by including a ground plane.
[0008] US 2016 / 0140400 A1 discloses systems and methods for providing an advanced warning system (AWS) to a driver of a vehicle in which traffic scenes are recorded using a single camera video. The single camera video is analyzed to generate monocular structure-from-motion (SFM) and two-dimensional object detection in real time, from which a ground plane is determined. A dense 3D estimation is performed using the monocular SFM and the two-dimensional object detection to generate a joint 3D object localization from the ground plane and the dense three-dimensional estimation. In particular, two-dimensional boundary boxes can be determined from the monocular video and then tracked across multiple images to derive 3D information.
[0009] In an article by Buyu Li et al., “Gs3d: An efficient 3D object detection framework for autonomous driving,” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 1019–1028, an efficient 3D object detection framework based on a single RGB image is described for use in autonomous driving scenarios. The framework aims to accurately determine the 3D boundaries of an object without a point cloud or stereo data. This is achieved by refining a rough box based on features of visible surfaces.
[0010] An article by Shichio Yang and Sebastian Scherer, “CubeSLAM: Monocular 3-D Object SLAM,” IEEE Transactions on Robotics, 2019, Vol. 35, No. 4, pp. 925–938, describes a method for single-frame object detection, determination of three-dimensional cuboids, simultaneous localization of an object from different views, and mapping in both static and dynamic environments. For single-frame object detection, high-quality cuboid suggestions are generated from two-dimensional bounding rectangles and vanishing point scanning, which are then evaluated and selected by comparison with image edges.
[0011] The invention is based on the objective of providing an improved, in particular easily implementable and robust, determination of dimension information describing the three-dimensional position and extent of a target object from two-dimensional camera images.
[0012] This problem is solved by a method, a motor vehicle, a computer program, and an electronically readable data carrier according to the independent claims. Advantageous embodiments are described in the dependent claims.
[0013] In a method of the type mentioned at the outset, it is provided according to the invention that - as part of the evaluation information, additionally an orientation information indicating the orientation of the target object for at least one vertically running side of the target object at its lower edge in the camera image, in particular an orientation line or an orientation point depending on the image content of the camera image, is determined, and - for at least one computation point of the camera image determined by means of the orientation information and the rectangle, the three-dimensional position of an associated ground corner point lying on the ground plane of the base surface of a cuboid that at least partially encloses the target object is determined using a recording geometry model of the camera, and from this the extent information is reconstructed as at least the base surface of the cuboid, in particular as the cuboid that at least partially encloses the target object.
[0014] The term "vertically oriented side" refers to a side of the target object that is at least substantially vertical in three-dimensional space. This means, in particular, the front, back, and lateral sides (i.e., right and left sides) with respect to the enclosing cuboid. It is generally assumed that the target object can be represented by a cuboid that at least partially, and especially completely, encloses it. This is conveniently possible, for example, with vehicles, especially motor vehicles, which can easily be assigned a top, a bottom, and vertically oriented sides such as the front, back, right, and left sides.
[0015] In other words, the cuboid that at least partially encloses the target object, and which in the corresponding dimension information can be described in particular by its eight vertices (four base vertices and four roof vertices), can be considered a representation of the target object that reproduces its three-dimensional position, orientation, and extent with sufficient accuracy. Here, the rectangle and / or cuboid are preferably defined such that they enclose the target object as closely as possible, in particular as the smallest possible two-dimensional or three-dimensional bounding box.It should be noted that the preferred, tightest possible enclosure of the cuboid in three-dimensional space does not necessarily imply the tightest possible enclosure of the projection in the camera image. Particularly on the underside, for example, if the mounting points on the ground plane are offset from the corners, the rectangle may extend beyond the cuboid. This can be taken into account when determining the rectangle, as will be explained later. Furthermore, it is conceivable to define the target object more narrowly in some applications, for example, excluding the side mirrors of a motor vehicle, since these only protrude locally in a very small area.
[0016] The evaluation information can therefore be understood as an annotation in the two-dimensional space of the camera image. Accordingly, the present invention proposes that, instead of working with annotations in three-dimensional space or merely with rectangles in the camera image, the annotated rectangles of the evaluation information should be supplemented with additional information. This allows for the generation of three-dimensional representations of the target objects, specifically floor surfaces or, preferably, entire cuboids, based on knowledge of a floor plane. It has been shown that information regarding the orientation in the camera image, i.e., in two dimensions, is particularly easy to determine and, provided it relates to the lower edge of the target object, together with the rectangle, already provides essential information for determining floor vertices—in this case, computational points that can be derived through simple geometric considerations.At least in the case that two adjacent, vertically running sides are completely visible in the camera image and the orientation information, in particular as an orientation line, has been determined for at least one of these sides, it is possible, by simple geometric considerations, to derive the complete base surface of the cuboid, described by four base vertices, solely from the evaluation information, in particular the rectangle and the orientation information, a recording geometry model of the camera, in particular a simple pinhole camera model, and the knowledge of the ground plane on which the target object stands, and, if necessary, to complete it to a complete cuboid with the help of the rectangle.
[0017] The reconstruction of the entire floor area or preferably of the entire cuboid based on at least one floor corner point, which has already been determined solely on the basis of the evaluation information, the floor plane and the recording geometry model, is preferably carried out solely on the basis of the evaluation information, the floor plane and the recording geometry model; however, at least in the case of image content that does not completely encompass two vertically running sides of the target object in perspective, additional information, in particular dimensional information contained in or associated with the evaluation information, can also be taken into account.In particular, the base surface can first be completed by determining the remaining base corners, after which the upper boundary surface of the cuboid parallel to the base surface (hereinafter referred to as the roof surface) can then be determined, if desired, using the base corners, the rectangle and the recording geometry model, possibly taking additional and / or alternative consideration of the dimensional information (for example, if the target object is "cut off" at the top of the camera image).
[0018] The evaluation information can be generated cost-effectively and with minimal effort. Furthermore, deriving the surface area, particularly of the cuboid, as a three-dimensional representation of the target object requires little computational effort, making it cost-effective and enabling robust implementation. Thus, the determination of the target object's position, orientation, and size in three-dimensional space—described by the surface area or the entire cuboid as dimension information—is made possible in a simple, cost-effective, and robust manner by considering orientation information already present in the camera image.
[0019] The camera, which could be, for example, a front camera on a vehicle as the ego object, provides a camera image of the captured scene containing at least one target object, such as road traffic when considering vehicles as target objects. Preferably, the roll angle of the camera relative to the ground or to the ego object, especially the vehicle, can be at least substantially zero. This makes it particularly easy to identify the vertical in the two-dimensional camera image. Preferably, the camera images are rectified, as is generally known, so that a pinhole camera model can be used as the image geometry model.
[0020] The camera is intrinsically and extrinsically calibrated. The intrinsic calibration of the camera enables the projection of points in the three-dimensional camera coordinate system into the rectified camera image using the image geometry model. The extrinsic calibration enables the transformation of points in the camera coordinate system into the ego-object coordinate system, in particular the vehicle coordinate system, and vice versa. Optionally, other sensors provided on the ego-object, especially in the vehicle, particularly distance sensors such as radar sensors and / or lidar sensors and / or laser scanners, can preferably also be calibrated so that the transformation of points in the respective sensor coordinate system into the ego-object coordinate system and vice versa is possible. The ego-object geometry is known, so a description of a (reference) ground plane can be provided, especially when the ego-object is stationary.
[0021] The ground plane is preferably determined as precisely as possible so that it coincides with the base of the respective target object on which it stands. Within the scope of the present invention, it can be provided that the ground plane under the target object is determined using the extrinsic calibration of the camera by assuming the same position as the ground plane under the ego object. In particular, in the case of possible pitching and / or rolling movements, position data from a tilt sensor and / or other sensor data can also be taken into account to account for changes in the camera's position relative to the ground plane caused by the ego object. However, changes in the ground plane from the ego object to target objects in its vicinity would either not be considered or would be assumed to be small.Therefore, within the scope of the present invention, it is preferred to determine an individual ground plane for each target object that describes the ground on which the target object stands as accurately as possible. This is always advantageous when the entire surrounding ground cannot be approximated sufficiently by a single plane, for example, in the case of undulating roads or branching ramps. In this case, negligible changes in the ground plane compared to the (reference) ground plane of the ego object would no longer be practical. In this regard, an embodiment of the present invention may preferably provide that the ground plane under the target object is determined, again using the extrinsic calibration of the camera, from sensor data of another distance sensor, in particular a laser scanner and / or radar sensor and / or lidar sensor, which is also extrinsically registered with the ego object.Using such distance sensors, the terrain around the ego object can be scanned using known methods and used accordingly, for example, by using the scanned terrain around the target object to determine the ground plane. Additionally or alternatively, map data stored in the ego object, such as a vehicle's navigation system, can also be used to determine target-specific ground planes. This data should describe the topology in the ego object's environment, particularly the vehicle. For navigation environments, especially parking environments, it has also been proposed to provide digital map data in the vehicles operating them.
[0022] As already mentioned, a pinhole camera model is preferred as the image geometry model. With intrinsic calibration and appropriate rectification of the camera images, this offers an extremely simple and computationally efficient way to determine the rays along which features / points located at a specific pixel in the camera image lie. Furthermore, straight lines in three-dimensional space are mapped onto straight lines in the image, which also results in the rectangle providing a more precise impression of the cuboid.
[0023] According to the invention, the evaluation information determined by the evaluation algorithm comprises at least one rectangle that encloses the projection of the cuboid surrounding the object as closely as possible in the camera image. It should be noted here that the rectangle does not necessarily have to precisely enclose the two-dimensional, visible silhouette of the target object, but in exemplary embodiments can also be chosen to be larger, particularly in the area of the lower edge of the rectangle, in order, for example, to compensate for the stroke of the wheels, which are offset along the longitudinal direction, when vehicles are the target objects.
[0024] If the target object is partially obscured in the camera image, the evaluation algorithm can expedite an estimation of the rectangle's actual dimensions. Such an estimation can also be performed with respect to other components of the evaluation information when the object is partially obscured.
[0025] It is expedient and particularly advantageous to determine, as part of the evaluation information, a separation information describing the position of at least one vertically oriented side of the target object visible in the camera image, in particular a vertical dividing line between different vertically oriented sides of the target object in the camera image. The separation information thus indicates, firstly, how many vertically oriented sides of the target object were detected in the camera image, and secondly, in which area of the camera image, in particular within the rectangle, these sides are located. Therefore, if several vertically oriented sides of the target object are visible, the separation information allows the rectangle to be divided according to where each of the vertically oriented sides of the target object is located.A particularly advantageous method is to determine a dividing line as separation information. This line runs vertically through the rectangle in the camera image, dividing it into areas, each containing one of the vertical sides of the target object. In other words, a straight, vertical dividing line can be determined that, if more than one vertical side of the target object is visible in the camera image, divides the rectangle into distinct parts or areas, each containing exactly one vertical side of the target object. If only one vertical side of the target object is visible, the dividing line conveniently coincides with the right or left edge of the rectangle. In other words, if only one vertical side of the target object is visible in the camera image, the dividing line can be determined as a vertical edge of the rectangle.
[0026] In this context, it is particularly advantageous if, for at least one vertically oriented side of the target object, especially at least the side to which the orientation information is assigned, a side classification as front, back, or lateral side, particularly right side or left side, is determined as part of the evaluation information. Specifically, for each vertically oriented side of the target object visible in the camera image and described by the separation information, a side classification as front, back, right, or left side can be performed, whereby it is also conceivable to combine the right side and left side as a lateral side. Together with the separation information, it thus becomes clear, in the case of several visible, vertically oriented sides of the target object, which part of the rectangle shows which side to be classified and how.It should be noted that, in principle, assigning one side or area of the rectangle to a side classification is sufficient, since adjacent sides that run vertically can be deduced from this, assuming that the target object is standing with its bottom side on the ground plane. For example, it can be assumed that a vehicle, as the target object, is standing with its wheels on the ground.
[0027] In principle, within the scope of the present invention, the specific evaluation information also includes orientation information, in particular either as an orientation line or as an orientation point. An orientation line is advantageously used when two vertically extending sides of the target object are perspectively visible in the camera image, and an orientation point is used when only one vertically extending side of the target object is visible in the camera image (the orientation line then, so to speak, degenerates into the orientation point). In other words, whenever a vertically extending side of the target object, in particular a lateral side, is perspectively visible in the camera image, the orientation line can be defined as an oblique line along the visible lower edge of the side of the target object, in particular the right or left side of the target object.For vehicles as target objects, the orientation line can therefore run along the points of contact between the wheels and the ground. If only one vertical side is visible, especially only the rear or the front, the orientation line degenerates into a point. In other words, it can be provided that, when more than one vertical side of the target object is visible in the camera image, the orientation information is determined as an orientation line along the lower edge of one of the sides visible in perspective in the camera image, especially a lateral side of the target object; in the case of a vehicle as the target object, specifically as an orientation line touching the lower edge of the vehicle's wheels. Furthermore, it can be provided that, when only one vertical side of the target object is visible in the camera image, the orientation information is determined as an orientation point indicating an orientation extending into the image plane.
[0028] In an advantageous embodiment of the present invention, a class of the target object can also be determined as part of the evaluation information. This class can be assigned, for example, dimensional information describing at least one expected dimension of a target object of that class. The class of the target object describes, for example, a type of target object. For example, the class information, which can also be usefully determined independently of the separation information, can include, for vehicles as target objects, at least an assignment to at least one two-wheeler class and / or at least one passenger car class and / or at least one truck class and / or at least one bus class. However, a more detailed subdivision into classes according to the actual size and shape of the vehicles is also particularly advantageous.For example, with regard to passenger cars, one can consider classes such as "passenger car, upper mid-range, station wagon," "passenger car, small car, hatchback," "passenger car, luxury car, SUV," and the like. All these classes or types can expediently be assigned dimensional information that describes, in particular, the expected dimensions of the corresponding target objects, for example, as the expected width and / or length and / or height. The dimensional information can also include ranges for permissible dimensions within the corresponding class of the target object, which will be discussed in more detail below.
[0029] If a side classification is present in the evaluation information, it can be useful to assign sides of the cuboid in the extent information to corresponding sides of the target object. For example, if the cuboid is described by its vertices (base corners and roof corners), the assignments in the side classification can be used to determine which base corner corresponds to the front / rear left / right lower corner of the target object. A corresponding assignment can be made for roof corners to the front / rear left / right upper corners. This directly results in the assignments of sides of the cuboid to sides of the target object. This is useful, for example, when assessing the direction and / or possibilities of travel for vehicles as target objects, as the orientation can then be assigned to a corresponding alignment.
[0030] If the evaluation information comprehensively determines the separation information, a simple case distinction can be made regarding how exactly the extent information should be determined. In the first case, the separation information indicates that two vertically oriented sides of the target object are visible within the rectangle. In other words, the separation line divides the rectangle into two rectangles, each with a positive area. Then, the entire base area, and thus the entire cuboid, can be determined solely using the rectangle, the orientation line of the orientation information, the base plane, and the acquisition geometry model. This is particularly advantageous because no external information or information from other time steps is required.In a second case, which, like the first, concerns target objects fully contained within the camera image, only one vertical side of the target object is visible. Therefore, in the specific case of a dividing line, this side coincides with the left or right edge of the rectangle. In such a case, as will be explained, it is advantageous to use the aforementioned dimensional information, which can also be derived, at least in part, from analyses of at least one previous time step in which more vertical sides of the target object were visible.
[0031] In a specific, advantageous embodiment of the present invention, it can therefore be provided that, in the case of two vertically extending sides of the target object visible in the camera image, separating information (case 1) and an orientation line relating to one of the visible sides, in particular a lateral side, as orientation information for determining a floor area covered by the target object on the ground. - a first calculation point is determined as the intersection of the orientation line with the lower edge of the rectangle in the camera image, and a second calculation point is determined as the intersection of the orientation line with the vertical edge of the rectangle assigned to the side to which the orientation line is related, according to the separation information. - using the recording geometry model for the first and second calculation points, rays on which the calculation points lie are determined, and the respective intersection points of the rays with the ground plane are determined as the first and second ground corners in three-dimensional space, - a search line perpendicular to a connecting line in the plane connecting the first and second corner points of the ground, passing through the first corner point of the ground, and a third corner point lying on the search line, whose projection into the camera image, determined according to the recording geometry model, lies on the vertical edge of the rectangle in the camera image opposite the edge of the side to which the orientation line refers, are determined in three-dimensional space, and - a fourth corner point of the floor, completing the rectangular floor area, is determined from the first, second and third corner points, in particular by adding the vectors belonging to the third and second corner points and subtracting the vector belonging to the first corner point.
[0032] It should be noted here that if orientation lines have been determined for both sides, they should ideally intersect at the first calculation point. This can be achieved, for example, by averaging the results (weighted, particularly with regard to the robustness of the orientation line determination) and then shifting the orientation lines parallel to each other. Alternatively, the evaluation algorithm can be designed to determine both orientation lines accordingly. If the orientation lines intersect at the first calculation point, a third calculation point can be determined analogously to the second. This third point directly yields the third vertex, analogous to the first and second vertices, and the fourth vertex can be derived directly from this third point. The procedure described above, using a search line, can then be applied to verify the results.However, in the case of vehicles as target objects, it has been shown that for lateral sides, due to the determinability of the contact points of the wheels, the orientation information can be determined much more robustly and easily than for rear or front sides, which are also significantly shorter and do not touch the ground. Therefore, it is then preferable to determine an orientation line as orientation information, which is preferably related to a lateral side.
[0033] This case, where only one orientation line is used for the orientation information, will now be explained in more detail.
[0034] First, the intersection points of the orientation line with the edges of the rectangle are determined. The intersection of the orientation line with the bottom edge of the rectangle forms the first calculation point, and the intersection of the orientation line with the side edge of the rectangle that is intersected by the orientation line (and thus also associated with the vertically oriented side of the target object visible in the camera image) forms the second calculation point. According to the image geometry model, particularly in the preferred example of the pinhole camera model, all points in the three-dimensional camera coordinate system that are projected onto a point in the camera image lie on a ray. This ray can be calculated using the image geometry model based on the camera's intrinsic calibration. This results in a first ray for the first calculation point and a second ray for the second calculation point.The points where these rays intersect the ground plane in the three-dimensional camera coordinate system can now be calculated. These points form the first ground vertex (intersection of the first ray and the ground plane) and the second ground vertex (intersection of the second ray and the ground plane). A unique search line exists that lies in the ground plane and through the first ground vertex, forming a right angle with the line segment connecting the first and second ground vertices. On this search line, there is a unique third ground vertex, whose projection into the camera image lies on the line formed by extending the side edge of the rectangle that is not intersected by the orientation line. If no intersection points of the first and second rays with the ground plane exist, the ground plane can be corrected, particularly by raising the horizon.
[0035] It should be noted here that if the third corner point of the ground is determined such that its projection into the camera image lies on an extension of the corresponding edge outside the rectangle, the ground plane and / or the orientation line and / or the rectangle and / or the first and / or the second calculation point can be corrected so that the projection lies on the edge. For example, information about the direction in which the projection lies outside the rectangle can be used, and / or correction algorithms employing optimization methods can be used to achieve a plausible result with the least possible adjustments.
[0036] If the third vertex of the base lies on the side edge of the rectangle that is not intersected by the orientation line, it is considered to be correctly determined. After calculating the third vertex, the fourth vertex is determined accordingly, especially if the vectors to the corresponding vertices are denoted as V1, V2, V3, and V4, using V4 = V3 + V2 - V1.
[0037] At this point, it can be usefully checked whether the projection of the fourth ground corner into the camera image lies within the rectangle. If this is not the case, the ground plane and / or the orientation line and / or the rectangle and / or the first and / or second calculation point can be adjusted so that the projection lies within the rectangle. In summary, regarding this plausibility check, it can be said that if the projection is onto an extension of the corresponding edge of a given third ground corner that lies outside the rectangle, and / or if the projection is onto a given fourth ground corner that lies outside the rectangle, the ground plane and / or the orientation line and / or the rectangle and / or the first and / or second calculation point are corrected so that the respective projections lie on the edge or within the rectangle, respectively.In both cases, as well as in the combined case, information about the location outside the rectangle can be used, as mentioned previously, along with optimization methods to minimize changes. The evaluation information, i.e., the two-dimensional annotation, can be considered preferentially more accurate, and the primary goal should be to correct the ground plane, which can be inaccurate, especially for distant target objects, for example, when determined using lower-resolution LiDAR.
[0038] The ground corner points describe a rectangular ground area that lies entirely within the ground plane. They therefore describe the estimated ground area covered by the target object.
[0039] If a complete cuboid is to be obtained, its height must still be determined. For this purpose, it can be provided that a roof surface of the cuboid, corresponding to the base surface, is determined by the set of roof vertices obtained by adding a common height value multiplied by the unit normal vector of the base plane to the base vertices, whose projection into the camera image lies within the rectangle when the maximum height value is reached for at least one roof vertex or all roof vertices.If we again denote the vectors of the base vertices as V1, V2, V3, and V4, and the vectors of the roof vertices as T1, T2, T3, and T4, and furthermore, if h denotes the height and N the unit normal vector (i.e., the vector orthogonal to the base surface, pointing upwards and with a length of 1), then it can be said that the height h should be chosen as large as possible and such that the projections of all points Tk = S1k + hx V with k=1, 2, 3, 4 into the camera image are all contained within the rectangle. The rectangular roof surface formed by T1, T2, T3, and T4 at this height forms the upper surface of the cuboid. For reasons of plausibility, it is expedient to require that all projections of the roof vertices into the camera image should lie within the rectangle.
[0040] It should be noted at this point that, due to the extrinsic calibration of the camera, the position of the ground corner points and the roof corner points, and thus of the cuboid, is known not only in the camera coordinate system, but also in the ego object coordinate system, in particular the motor vehicle coordinate system.
[0041] In the second case mentioned above, it is expedient to also consider dimensional information in order to estimate a dimension not visible in the camera image, such as length or width, as reliably as possible. For this purpose, it can be provided that, as already mentioned, a class of the target object is determined as part of the evaluation information. This class is assigned at least one expected dimension of the target object within that class. Additionally or alternatively, at least one expected dimension of the target object can be assigned to the target object based on a previously determined extent information.For example, if dimension information regarding the target object was previously determined, particularly based on a previously recorded camera image where the first case applied (i.e., two vertical sides of the target object were visible in the camera image), then the width and / or length of the target object was also determined with a high degree of reliability at that time and can be provided as dimensional information. This information can also be continuously updated and refined, for example, by averaging, using additional camera images for which a determination according to the first case is possible. Especially in cases where no such dimensional information is available from a previous time step, an external dimensional information can also be used, which can be expediently selected based on the class of the target object, which was determined as part of the evaluation information.In this particular case, it is especially advantageous to use the most precise possible subdivision of the target object into different classes based on its size and / or shape. Examples of this have already been given above for vehicles as target objects, so that, for instance, smaller expected dimensions can be assumed for a small car than for an SUV. If the dimension information is to be used for a safety-relevant vehicle function, it can also be useful to deliberately set the expected dimensions within a maximum range of conceivable values.
[0042] Then, in the second case mentioned above, i.e., when only one vertically running side of the target object is visible in the camera image, indicated by the separation information, it may be provided that the following is used to determine a ground area covered by the target object: - a first calculation point is determined as the lower left corner point of the rectangle in the camera image and a second calculation point as the lower right corner point of the rectangle in the camera image, - using the recording geometry model for the first and second calculation points, rays on which the calculation points lie are determined, and the respective intersection points of the rays with the ground plane are determined as the first and second ground corners in three-dimensional space, and - two search lines perpendicular to a connecting line connecting the first and second corner points of the ground in the ground plane through the respective corner points and third and fourth corner points along the respective search line in a direction away from the camera when the camera image is taken at a distance in three-dimensional space determined according to the dimensional information.
[0043] If, in turn, the entire cuboid is to be determined, a roof surface of the cuboid associated with the base surface can again be determined by the set of roof vertices resulting from the addition of a common height value multiplied by the normal unit vector of the base plane to the base vertices, whose projection into the camera image at maximum height value for at least one roof vertex, in particular all roof vertex, lie within the rectangle.
[0044] In this second case, the lower right and left corners of the rectangle in the camera image are used as calculation points, as they correspond to ground vertices. As in the first case, the corresponding ground vertices can be determined in the three-dimensional camera coordinate system by finding the intersection of the corresponding rays with the ground plane. If no intersection points of the first and second rays with the ground plane exist, the ground plane can be corrected, particularly by raising the horizon. However, two search lines are now used, which are perpendicular to and pass through the line connecting the first and second ground vertices. In other words, they can also be understood as search rays that extend away from the camera at the time the image is captured and originate at the first and second ground vertices, respectively.A distance can now be derived from the dimensional information, for example, based on a result obtained in at least one previous time step or based on a corresponding expected dimension. If, for example, only the front or back of the target object is visible, the distance can be understood as a reference length of the target object; if only a lateral side, such as the left or right side of the target object, is visible, the distance can be understood as a reference width. In this context, the side classification as part of the evaluation information proves particularly useful, since the visible side of the target object is then classified, and thus it is immediately known which dimension is missing.
[0045] In other words, it can be provided that, if a side classification is available as part of the evaluation information, the expected dimension in the direction perpendicular to the vertically oriented, classified side visible in the camera image is selected from the dimensional information as a distance. As already mentioned, dimensional information associated with a class of the target object can be used; however, if a length or width of the target object, observed at other times, is known as the expected extent, this should be used preferentially. The third and fourth ground corners are then determined by moving along the search line away from the camera at the time the camera image was taken, by the distance, i.e., the reference length or reference width.
[0046] In this context, it should be noted that errors are conceivable in which at least one of the search lines initially extends beyond the rectangle in its projection into the camera image. In such cases, it may be expedient to correct the ground plane and / or the rectangle so that the projections of the search lines remain within the camera image at least up to a certain distance. In other words, if the projection of a search ray into the camera image immediately extends beyond the rectangle, a correction step is performed, and the ground plane and / or the rectangle are adjusted so that, after recalculating the first and second corner points of the ground and the search line, the projections of the search line, starting from the first and second corner points respectively, initially remain within the rectangle.
[0047] In this way, the base area of the cuboid is determined, so that if the entire cuboid is to be determined as expansion information, the roof area can also be determined accordingly, as described in the first case.
[0048] Particularly when the target object is not fully contained within the camera image, further cases beyond the first and second are conceivable, which can also be addressed within the scope of the present invention. In such special cases, the target object, whose enclosing cuboid is to be determined, is not fully visible in the camera image because it is cut off by the image edge. However, in many such special cases, at least a very good estimate can still be achieved, at least with the addition of dimensional information. In other words, if the target object is not fully captured in the camera image, especially if only one base corner can be determined, the missing information for reconstructing the base surface or the cuboid can be extracted from the dimensional information.
[0049] In these special cases, depending on visibility, only partial information from the camera image is used, for example, only two or three edges of the rectangle or only one of the first or second corner of the floor. Depending on availability, parts of the desired cuboid can then be calculated using the methods described for the first and second cases. Corners of the floor that cannot be deduced from the camera image can be derived using dimensional information, i.e., by using expected dimensions observed at another time and / or inferred from a class of the target object. For example, if only the second calculation point and part of the orientation line are visible in the camera image, the second corner of the floor can be determined. A ray can be calculated from the orientation line that runs in the floor plane from the second corner of the floor.One follows this search line along the expected dimension to reach the first corner point of the ground.
[0050] Furthermore, in exemplary embodiments of the present invention, it may also be provided that the dimensional information is used for plausibility checks. For example, the dimensional information may additionally include an interval of permissible dimensions for the respective class of the target object. In a plausibility check step, certain dimensions of the target object from the camera image are compared with this interval, and if the results are implausible, the dimension information and / or its underlying data are recalculated and / or corrected, and / or the dimension information is discarded. Based on the class of the target object, which may already be determined as part of the evaluation information, the obtained dimensions of the cuboid in the dimension information can be further plausibly checked and, if necessary, corrected.For this purpose, reference intervals for the length, width, and, if applicable, height of target objects of the respective class can be used. If the dimensions of the base or the specific cuboid lie outside these intervals, a correction can be made. For example, it can be stipulated that if the dimensions fall below the lowest value of the interval, they are corrected to that lower value, and if they exceed the highest value of the interval, they are corrected to that highest value. The corrections for length and width are preferably applied symmetrically. However, particularly in cases of significant deviations, it is also conceivable to discard the extent information for that time step and use a camera image from the next time step to attempt a new determination. For example, the extent information from a previous time step can then be used initially.Of course, a plausibility check using time steps is also conceivable within the framework of the present invention.
[0051] In a further advantageous embodiment of the present invention, it can be provided that a trained function, in particular a neural network, is used as the evaluation algorithm, which has been trained, in particular, with manually annotated camera images as training data. Within the scope of the present invention, it has been shown that for the minimum desired evaluation information (rectangle, orientation information), but also with the addition of optional, particularly preferred, further evaluation information such as separation information, side classification, and the class of the target object, a robust, fast, and manageable amount of training data-required trained function of artificial intelligence, for example, a convolutional neural network (CNN), can be used.Machine learning techniques can therefore be used to train the evaluation algorithm and derive the evaluation information quickly and efficiently from the camera image data, and in particular, exclusively from the camera image. Manual annotations can be used to provide suitable training data, but it is also conceivable to use at least partially automatically generated annotations to create the training data. For example, image processing algorithms that function without artificial intelligence are often capable of calculating rectangles and orientation lines that closely enclose target objects.Using such training data, an artificial intelligence function is trained through machine learning, which recognizes relevant target objects in the camera image and outputs exactly the aforementioned evaluation information for all recognized target objects, for example for vehicles.
[0052] In this context, it should be noted that, with regard to target objects not fully contained within the camera image, it is also conceivable to design the evaluation algorithm in such a way that it can estimate the missing information, at least in some cases. For example, it could be provided that, for at least some of the target objects not fully captured in the camera image, the evaluation algorithm, particularly through appropriate training, is designed to estimate the complete rectangle extending beyond the camera image. Specifically, an evaluation algorithm is conceivable that first determines a class of target object and then uses the dimensional information assigned to that class to determine the evaluation information for the expected dimensions.It is of course also conceivable that the evaluation algorithm first identifies the target object as having already been processed at a previous time, so that the dimension information from an earlier time step can then be used accordingly.
[0053] In addition to the method, the present invention also relates to a motor vehicle comprising a control unit configured for carrying out the method according to the invention and a camera. In this case, vehicles are preferably used as the target objects. The camera can, for example, be the front camera of the motor vehicle. All aspects relating to the method according to the invention can be applied analogously to the motor vehicle according to the invention, so that the aforementioned advantages can also be obtained with it.
[0054] The control unit comprises, in particular, at least one storage medium and at least one processor, wherein the processor can provide functional units for carrying out the method according to the invention. The control unit can, in particular, include an interface for receiving the camera image captured by the camera, an evaluation unit for executing the evaluation algorithm, a ground plane determination unit for determining the ground plane / ground information, and an extent information determination unit for determining the extent information.
[0055] The obtained dimension information can be used in a variety of ways, generally and universally, including in other applications, for example, to generate training data for other artificial intelligence algorithms. However, the dimension information is particularly advantageous for use in the operation of the motor vehicle, especially for vehicle functions. Thus, a further advantageous embodiment of the motor vehicle according to the invention provides that the control unit, or a further control unit, is designed to use the dimension information in the execution of at least one vehicle function, in particular a vehicle guidance function that includes at least partially automatic steering of the motor vehicle.
[0056] In particular, if information from distance sensors is to be used to determine the ground plane for the target object, it is also conceivable that the motor vehicle also has at least one distance sensor.
[0057] A computer program according to the invention can be directly loaded into a storage medium of a computing device, for example, a control unit of a motor vehicle, and includes program means for carrying out the steps of a method according to the invention when the computer program is executed on the computing device. The computer program can be stored on an electronically readable data carrier according to the invention, which therefore includes control information stored thereon, comprising at least one computer program according to the invention and, when the data carrier is used in a computing device, enabling the device to carry out a method according to the invention. The data carrier can, for example, be a non-transient data carrier, such as a CD-ROM.
[0058] Further advantages and details of the present invention will become apparent from the exemplary embodiments described below and from the drawing. The drawings show: Fig. 1 a flowchart of an embodiment of the method according to the invention, Fig. 2. A first camera image with annotations, Fig. 3 a second camera image with annotations, Fig. 4 a third camera image with annotations, and Fig. 5 a schematic diagram of a motor vehicle according to the invention.
[0059] Fig. Figure 1 shows a general flowchart of exemplary embodiments of the method according to the invention. The following describes exemplary embodiments in which a motor vehicle, as the ego object, has a camera, for example, a front camera. However, other ego objects are also conceivable, particularly when it comes to generating 3D annotations as training data for an artificial intelligence algorithm, such as tripods or mobile units with cameras. The camera is intrinsically and extrinsically calibrated, with the camera also performing, in particular, rectification of recorded two-dimensional camera images, so that rays can be calculated using a pinhole camera model as the recording geometry model, on which visible features, for example, points, must lie at certain pixels of the camera image.Through external camera calibration, the three-dimensional camera coordinate system is registered with the three-dimensional ego-object coordinate system, in this example the vehicle coordinate system. This means that a conversion of corresponding three-dimensional positions and directions is possible. In the vehicle coordinate system, the position of the ground plane traversed by the vehicle is also known, particularly due to the known position of the wheels.
[0060] The camera now provides two-dimensional images of a traffic situation in which a multitude of objects are visible. For specific target objects, in this case vehicles of various types and classes, the method described here is intended to determine dimension information that describes the position, orientation, and extent of the target objects, i.e., vehicles. In the embodiments described here, the extent is described by a cuboid enclosing the respective vehicle as the target object, and the position and orientation are described by the position and / or orientation of the cuboid. The cuboid is therefore a three-dimensional bounding box.
[0061] In step S1, a two-dimensional camera image is captured, primarily controlled by the computing unit executing the procedure. In the case of a motor vehicle as the ego object, the computing unit could be the vehicle's control unit.
[0062] In step S2, the camera image is passed to an evaluation algorithm, which outputs evaluation information in the form of two-dimensional annotations for each of the target objects detectable in the camera image, in this case, vehicles. The evaluation algorithm is a trained function of artificial intelligence, which was trained using appropriate training data (machine learning). The training data can include manually and / or automatically annotated two-dimensional camera images. However, implementation examples are also conceivable in which at least part of the evaluation information is determined "conventionally," for example, by appropriate image processing algorithms without machine learning.
[0063] The evaluation information comprises five components. First, a rectangle is defined for each target object in the camera image. This rectangle encloses the target object within the camera image and thus also includes the projection of the cuboid that precisely surrounds the real, three-dimensional target object. The rectangle can, but does not necessarily have to, precisely enclose the two-dimensional visible silhouette of the target object; it can be larger, particularly at the bottom. This is especially relevant in the case of vehicles as target objects discussed here, since the wheels, defining the lowest areas, are offset from the front and rear.
[0064] It should be noted at this point that, since the roll angle of the camera relative to the ground or vehicle is essentially zero, the vertical direction in the camera images as well as "top" and "bottom" are clearly defined and recognizable.
[0065] In the present case, the evaluation algorithm is also trained to estimate actual missing dimensions when the target object is partially obscured in the camera image, provided the rectangle and other evaluation information are annotated. Furthermore, in embodiments, it may also be provided that, in the case of target objects not fully visible in the camera image (i.e., cut off by the image edge), missing portions outside the image area of the camera image can be filled in, particularly using prior information, especially dimensional information of the target object, so that the rectangle extends beyond the camera image. However, the specific examples discussed below do not assume this.
[0066] In addition to the rectangle, the evaluation algorithm in step S2 also determines a vertical, straight dividing line as evaluation information. This line divides rectangle B into two sub-rectangles if two vertical sides of the target object are visible. Each sub-rectangle contains exactly one of these vertical sides of the target object, i.e., a lateral side (right or left side), front or rear of the vehicle. If only one of these vertical sides of the target object is visible in the camera image, the dividing line can also coincide with the right or left vertical edge of the rectangle.
[0067] As evaluation information, a side classification is determined for each vertically oriented side of the target object visible in the camera image, particularly for sides that can be assigned based on their position using the separation information. This means that for each of these visible, vertically oriented sides, it is known whether it is the front, the back, the left side, or the right side. It should be noted that the assignment of a side class to a visible, vertically oriented side of the target object necessarily implies other sides, assuming here that the vehicle as the target object is standing with its wheels on the ground.
[0068] The evaluation algorithm also determines orientation information within the two-dimensional camera image as evaluation information. Specifically, if the target object has two perspectively visible, vertically oriented sides, it determines an orientation line; if only one visible, vertically oriented side, it determines an orientation point to which the orientation line has degenerated. In such a (second) case, the orientation information is only indirectly considered for the actual derivation of the cuboid.
[0069] Finally, the evaluation algorithm also determines the class of the target object as evaluation information. Since the target objects are vehicles, the class always describes whether it is a two-wheeler / motorcycle, a passenger car, a truck, a bus, a tractor, or the like. However, in this case, a more detailed subdivision of the classes is provided according to the actual size and shape of the vehicles. For example, for passenger cars, a distinction can be made between small cars, lower mid-range cars, upper mid-range cars, and full-size cars, and a specific type of passenger car can also be determined, such as station wagon, hatchback, SUV, etc. The same procedure can, of course, be applied to trucks, buses, two-wheelers, construction vehicles, tractors, and the like.
[0070] Each class of target objects, for example, "passenger car, small car, station wagon," is assigned dimensional information. This information includes, on the one hand, expected dimensions for target objects of the class, such as an expected reference height, an expected reference width, and an expected reference length. On the other hand, it also contains intervals within which the corresponding dimensions typically lie, specifically intervals with minimum and maximum values for height, length, and width. It should be noted that dimensional information, for specific target objects, can also be assigned to them from previous time steps or points in time. Especially with a motor vehicle, it can be assumed that camera images are regularly captured and evaluated, particularly at intervals.Furthermore, it is known and common practice to track target objects between these time steps. If the extent information, in particular the cuboid, is determined for a target object in a given time step, its dimensions already describe expected dimensions for subsequent time steps, specifically an expected height, an expected width, and an expected length. These can remain assigned to the target object as dimensional information, provided it is tracked, and can be statistically aggregated over time to obtain increasingly reliable and refined values for the expected dimensions.When it comes to the question of whether to use dimensional information assigned to the target object from previous time steps or dimensional information assigned to the class of the target object, the dimensional information from the previous time steps is preferable if it is sufficiently reliable, especially if it was determined solely from the camera image, i.e., without any further external assumptions regarding the target object.
[0071] Examples of camera images with annotations provided by the evaluation information are in the Fig. 2 to 4 are shown. This shows Fig. Figure 2 schematically shows a camera image 1 in which a motor vehicle is visible on a road 3 as the target object 2. Additionally, the rectangle 4, the dividing line 5 (defined as separation information), and the orientation line 6 (defined as orientation information) are shown as evaluation information. The orientation line 6 was defined as the lateral side for the right side 7 of the target object 2, visible in the right area of the rectangle 4 divided by the dividing line 5, because particularly clear and robust detection is possible there due to the contact points of the wheels 8 of the target object 2. The vertically oriented side of the target object 2 visible on the left side is accordingly the rear side 9. The side classification thus indicates "rear side" for the left area of the rectangle 4 and "right side" for the right area of the rectangle 4.
[0072] In the present example, rectangle 4 has been extended downwards so that the dividing line 5 and the orientation line 6 intersect at the lower edge 10 of rectangle 4. This is intentional, as the wheels 8 also form the lowest point of the target object 2, which is higher at its rear, thus at least partially accommodating this variation. This illustrates that rectangle 4 does not necessarily have to precisely enclose the two-dimensional, visible silhouette of the target object 2, but can be larger if circumstances make this seem appropriate.
[0073] It should also be noted that in Fig. 2. The class of the target object was determined to be "passenger car, upper middle class, station wagon".
[0074] In the example of the Fig. Figure 3 shows another camera image 11, in which a target object 12, again a motor vehicle, is visible, located on an area 13. The evaluation information again includes the rectangle 14 and the dividing line 15, which, since only one vertical side of the target object 12 is visible, coincides with the right edge of the rectangle 14 and is highlighted by a dashed line. Accordingly, only one orientation point 16 is shown as orientation information, indicating that the orientation of the lateral side, in this case the right side 17 of the target object 12, extends into the image plane. The wheels 18 are also more difficult to identify due to the rims not being visible. The side classification contains the information that the front side 19 is visible in the rectangle 14.Furthermore, since the dividing line 15 coincides with the right edge of rectangle 14, it can contain the information that the right side 17 would follow there. If the dividing line 15 were located on the opposite, left edge of rectangle 14, which is also possible, then, logically, the left side 20 of the target object 12 would follow.
[0075] The class of target object 12 was defined here as "passenger car, small car, station wagon".
[0076] Fig. Figure 4 shows a third example of a camera image 21, which again shows a target object 22, designed as a vehicle, on a road 23. The rear end of the target object 22 is clearly cut off by the image edge. This means that the evaluation information here is incomplete and only partially determined. In fact, only two edges of the rectangle 24 could be identified. There is no separation information in this sense, since the right side 27 of the target object 22 is also not fully visible (see section edge 25). However, the orientation line 26 could be readily determined as orientation information due to the visibility of the wheels 28. The side classification at least contains the information that the partially visible, vertically running side of the target object 22 is the right side 27.That the front 29 would be attached on the left and the (not visible) back on the right is immediately apparent from the fact that the target object 22 is standing on the ground with its wheels 28.
[0077] The class of target object 22 was identified as "passenger car, upper class, SUV".
[0078] Returning to Fig. In step S3, the ground plane on which the target object 2, 12, 22 is located is determined, described by corresponding ground information. Fig. For the sake of simplicity, the ground plane in sections 2 to 4 is consistently indicated by the reference symbol 30. There are several ways to determine it.
[0079] One possibility is to define the ground plane 30 so that it coincides with the reference ground surface of the ego object, in this case, the user's own vehicle. This reference ground plane is also known in the three-dimensional camera coordinate system due to external calibration. However, this assumption is less suitable, particularly when the entire surrounding ground cannot be approximated accurately enough by a single plane, for example, in the presence of ramps.
[0080] Therefore, another variant may also provide for the use of sensor data from a distance sensor of the ego object, which is also extrinsically calibrated to it, and / or map data, for example map data from a navigation system of the vehicle itself as the ego object, to determine a ground plane specifically relating to the target object 2, 12, 22. Such a distance sensor that scans the ground could, for example, be a radar sensor, a lidar sensor and / or a laser scanner.
[0081] After the evaluation information and the ground level 30 have been determined, S4 is determined in one step, see again. Fig. 1, the extent information is determined. For this purpose, two cases are first distinguished using the separation information when the target object 2, 12 is completely visible in the camera image 1, 11 (cf. Fig. 2 and Fig. 3) In a first case, see: Fig. 2, are two vertically oriented sides of the target object 2, in the example of the Fig. 2. The right side 7 and the back 9 are visible in perspective in camera image 1, so that an orientation line 6 could also be determined. In a second case, however, only one vertically running side is visible, in the example of the Fig. 3 the front side 19, which can be seen in the camera image 11, so that the orientation information was determined as orientation point 16.
[0082] Is the first case recognizable from the separation information, in which the separation line 5 (cf. Fig. 2) If rectangle 4 is divided into two sub-rectangles, each with a positive area, the intersection of orientation line 6 with the lower edge 10 of rectangle 4 is first determined as the first calculation point 31, and the intersection of orientation line 6 with the corresponding side edge 32 of rectangle 4 is determined as the second calculation point 33. Orientation line 6 intersects the right edge 32, since this edge bounds the sub-rectangle containing the right side 7 for which orientation line 6 was determined.
[0083] According to the pinhole camera model, which is used here as the image geometry model, all points in the three-dimensional camera coordinate system that are projected onto a point in camera image 1 lie on a single ray. This ray can be calculated using the camera's intrinsic calibration. This is done here for calculation points 31 and 32. The points of intersection of these rays with the base plane 30 in the three-dimensional camera coordinate system are then calculated and form a first base vertex of the desired cuboid for the first calculation point 31 and a second base vertex of the desired cuboid for the second calculation point 33. The base vertices are connected by a straight line. A unique search line is then determined that runs in the base plane 30 and through the first base vertex, forming a right angle with the connecting line.
[0084] On the search line, there is a unique third ground corner whose projection into the camera image 1 lies on the line obtained by extending the left side edge 34 of rectangle 4, which is not intersected by the orientation line 6. If the projection of the third ground corner determined in this way does not lie on the left edge 34, but outside of rectangle 4, a correction must be made to the ground plane 30, rectangle 4, and / or the calculation points 31 and / or 33 such that, after recalculating the first, second, and third ground corners, the projection of the third ground corner into the camera image 1 lies on edge 34.
[0085] Once the third corner point of the base, or more precisely its three-dimensional position in the camera coordinate system, has been determined in this way, the fourth corner point, which completes the corner points of the base of the cuboid, can also be easily determined by summing the vectors to the third corner point and the second corner point and subtracting the vector to the first corner point from this sum.
[0086] Here, it is also checked whether the projection of the fourth ground corner lies within rectangle 4. If this is not the case, a correction can be made to the ground plane 30, rectangle 4, the first calculation point 31 and / or the second calculation point 33, so that after recalculating the ground corners, the third and fourth ground corners meet the described projection requirements.
[0087] The rectangular base of the cuboid, defined by its vertices and lying entirely within the plane 30, is now present. To obtain the complete cuboid, its height is determined, chosen to be as large as possible, such that the projections of the set of all roof points—derived from the base vertices by a shift of the height in the direction of the upward-pointing unit normal vector of the plane 30—are all contained within rectangle 4 in the camera image 1. The roof surface formed by this set of roof vertices constitutes the upper surface of the cuboid.
[0088] In the second case, see: Fig. 3. If the dividing line 15 coincides with the left or right edge of rectangle 14, then the lower left corner of rectangle 14 is chosen as the first calculation point 35, and the lower right corner of rectangle 14, which also corresponds to the orientation point 16, is chosen as the second calculation point 36. Analogous to the first, using Fig. In the case discussed, the corresponding rays and the points of intersection with the ground plane 30, i.e., the first and second ground vertices, can again be determined using the survey geometry model. Subsequently, two search lines are determined that pass through the respective first and second ground vertices in the ground plane 30 and are perpendicular to the line connecting the first and second ground vertices.
[0089] Should the projection of one of these search lines from the viewer, i.e., the position of the camera at the time the camera image 11 is taken, away from the first calculation point 35 or the second calculation point 36, immediately run out of the rectangle 14, a correction step is carried out and the ground plane 30 and / or the rectangle 14 are corrected so that, after recalculation of the first and second ground corner points and the search line, the projections of the search line starting from the first calculation point 35 or second calculation point 36 initially run within the rectangle 14.
[0090] As mentioned at the outset, dimensional information with expected extents is now known, either based on the class of the target object 12 according to the evaluation information or from previous time steps. Since, in the present case, only the front face 19 of the target object 12 is visible in the camera image 11, the perpendicular reference length or expected length is chosen as the expected dimension in the dimensional information. This is the distance along which the search line must be traversed away from the viewer in order to find the third and fourth ground corners with this assumed reference length. A dimensional information is preferably used that was determined from a previous determination of the extent information, when two vertically extending sides of the target object 12 were visible in the camera image from a perspective standpoint, as already described.
[0091] If only the right or left side of the target object 12, i.e. a lateral side, is visible in a camera image 11, the reference width is derived from the dimensional information and used accordingly.
[0092] The determination of the roof corner points for the roof surface is carried out as in the first case.
[0093] The side classification makes it possible to complete the extent information by assigning sides of the target object 2, 12, 22 to the corresponding sides of the cuboid, in particular to determine which of the bottom corner points corresponds to the front / rear left / right lower corner point of the cuboid, and similarly for the roof corner points.
[0094] It should be noted that if the dimensional information assigned to the target object class 2, 12, 22 also contains the described intervals, further plausibility checks of the obtained dimensions of the cuboid and, if necessary, a correction can be carried out. For this purpose, the dimensions of the specific cuboid are compared with the respective reference intervals for the length, width, and height of target objects for the class, and appropriate measures are taken if the dimensions are outside the intervals. For example, the dimension, especially if symmetrical, is set to the minimum or maximum value of the interval, or, in the case of significant deviations, the result is even discarded.
[0095] Even in the special case of Fig. 4, in which the target object 22 is not fully visible in the camera image 21, can be in step S4 of the Fig. 1. Extent information can be determined, since in such cases partial information can be used, which can then be supplemented by suitable dimensional information. Thus, in Fig. Only the second calculation point 37 is visible in the camera image 21, as the lower edge of the rectangle 24 is missing. Nevertheless, the corresponding second ground corner can be determined for the second calculation point 37. Using the orientation line 26, a search ray can be calculated that extends from the ground plane 30 away from the second ground corner. A distance corresponding to a reference length from the dimensional information can then be traversed along this search ray to determine the first ground corner. The procedure can then be continued as in the second case, with the distance along the search line then corresponding to a reference width.
[0096] Returning to Fig. In step S5, the specified extent information can be used, which describes the position, orientation, and extent of the corresponding target object 2, 12, 22. While such specified extent information can be used, for example, as training data in the sense of a three-dimensional annotation for trained functions, the extent information can also preferably be used in a vehicle function of the user's own vehicle as an ego object.
[0097] Fig.Figure 5 schematically shows a motor vehicle 38 according to the invention in the form of a schematic diagram. The motor vehicle 38 has a camera 39, which is directed towards the area in front of the motor vehicle 38 and provides the camera images. This camera is connected to a control unit 40, which is configured to carry out the method according to the invention. The control unit or another control unit can also be configured to perform a vehicle function that, as described in step S5, uses the dimension information, in particular in an at least partially automatic vehicle guidance function. Further components of the motor vehicle 38 can include a navigation system 41, which can provide digital map data for determining the ground plane 30, and distance sensors 42, which can also provide information in this regard.
[0098] The control unit 40 comprises, in addition to at least one storage means for carrying out the method according to the invention, in particular an interface and / or control unit for controlling the camera 39 or for receiving the two-dimensional camera image 1, 11, 21 (step S1), an evaluation unit for applying the evaluation algorithm to the camera image 1, 11, 21 to determine the evaluation information according to step S2, a ground plane determination unit for determining the ground information describing the ground plane 30, as described for step S3, and an extent information determination unit for determining the extent information according to step S4.
Claims
[1] Method for determining three-dimensional extent information of a target object (2, 12, 22), in particular a vehicle, from a two-dimensional camera image (1, 11, 21) showing the target object (2, 12, 22), which is recorded with a camera (39) that is intrinsically and extrinsically calibrated with respect to an ego-object, in particular a motor vehicle (38), carrying the camera (39), wherein at least one rectangle (4, 14, 24) describing the target object (2, 12, 22) in the camera image (1, 11, 21) is determined from ground information about the three-dimensional position of the ground plane (30) under the target object (2, 12, 22) relative to the camera (39) and from two-dimensional evaluation information describing the target object (2, 12, 22) in the camera image (1, 11, 21) by means of an evaluation algorithm. Extent information is determined, characterized by , that - as part of the evaluation information, additional orientation information indicating the orientation of the target object (2, 12, 22) for at least one vertically extending side of the target object (2, 12, 22) at its lower edge in the camera image (1, 11, 21), in particular an orientation line (6, 26) or an orientation point (16) depending on the image content of the camera image (1, 11, 21), is determined, and - for at least one computation point (31, 33, 35, 36, 37) of the camera image (1, 11, 21) determined by means of the orientation information and the rectangle (4, 14, 24), the three-dimensional position of an associated ground corner point of the ground surface of a cuboid at least partially enclosing the target object (2, 12, 22) lying on the ground plane (30) is determined using a recording geometry model of the camera (39) and the extent information is reconstructed from this as at least the ground surface of the cuboid covered by the target object (2, 12, 22), in particular as the cuboid. [2] Method according to claim 1, characterized by , that the ground plane (30) under the target object (2, 12, 22) using the extrinsic calibration of the camera (39) - by assuming the same position as the ground plane beneath the ego-object or - from sensor data from another distance sensor (42) that is also extrinsically registered with the ego object, in particular a laser scanner and / or radar sensor, and / or determined using map data. [3] Method according to claim 1 or 2, characterized by , that a pinhole camera model is used as the recording geometry model. [4] Method according to any of the preceding claims, characterized by , - that for at least one vertically extending side of the target object (2, 12, 22), in particular at least the at least one side to which the orientation information is assigned, a side classification into front (19, 29), back (9) or lateral side, in particular right side (7, 17) or left side (20, 27), is determined as part of the evaluation information and / or - that the orientation information is determined as an orientation line (6, 26) along the lower edge of one of the sides visible in perspective in the camera image (1, 11, 21), in particular a lateral side, of the target object (2, 12, 22) when more than one vertically extending side of the target object (2, 12, 22) is visible in the camera image (1, 11, 21), and in the case of a vehicle as the target object (2, 12, 22), in particular as an orientation line (6, 26) touching the lower edge of the wheels (8, 18, 28) of the vehicle, and / or - that if only one vertically extending side of the target object (2, 12, 22) is visible in the camera image (1, 11, 21), the orientation information is determined as an orientation point (16) indicating an orientation extending into the image plane. [5] Method according to any of the preceding claims, characterized by, that the evaluation information comprehensively determines a separation information describing the position of at least one vertically extending side of the target object (2, 12, 22) visible in the camera image (1, 11, 21), in particular a vertically extending dividing line (5, 15) between different vertically extending sides of the target object (2, 12, 22) in the camera image (1, 11, 21). [6] Method according to claim 5, characterized by , that in the case of two vertically extending sides of the target object (2, 12, 22) visible in the camera image (1, 11, 21) indicating separation information and an orientation line (6, 26) relating to one of the visible sides, in particular a lateral side, as orientation information for determining the ground area covered by the target object (2, 12, 22) on the ground plane (30) - a first calculation point (31) as the intersection of the orientation line (6, 26) with the lower edge (10) of the rectangle (4, 14, 24) in the camera image (1, 11, 21) and a second calculation point (33) as the intersection of the orientation line (6, 26) with the vertical edge (32) of the rectangle (4, 14, 24) assigned to the side to which the orientation line (6, 26) is related, according to the separation information, - using the recording geometry model for the first and second calculation points (31, 33), rays on which the calculation points (31, 33) lie are determined and the respective intersection points of the rays with the ground plane (30) are determined as the first and second ground corners in three-dimensional space, - a search line perpendicular to a connecting line in the plane (30) connecting the first and second corner points of the ground and passing through the first corner point of the ground, as well as a third corner point of the ground lying on the search line, whose projection into the camera image (1, 11, 21), which can be determined based on the recording geometry model, lies on the vertical edge (34) of the rectangle (4, 14, 24) in the camera image (1, 11, 21) opposite the edge (32) of the side to which the orientation line (6, 26) is referred, and - a fourth corner point of the floor, completing the rectangular floor area, is determined from the first, second and third corner points, in particular by adding the vectors belonging to the third and second corner points and subtracting the vector belonging to the first corner point. [7] Method according to claim 6, characterized by, that in the case of a projection on an extension of the corresponding edge (34) outside the rectangle (4, 14, 24) having a determined third ground corner and / or in the case of a projection outside the rectangle (4, 14, 24) having a determined fourth ground corner, the ground plane (30) and / or the orientation line (6, 26) and / or the rectangle (4, 14, 24) and / or the first and / or second calculation point (31, 33) are corrected in such a way that the respective projections lie on the edge (34) or inside the rectangle (4, 14, 24). [8] Method according to any one of claims 5 to 7, characterized by, that as part of the evaluation information a class of the target object (2, 12, 22) is determined, to which at least one expectation dimension of a target object (2, 12, 22) of the class is assigned dimensional information describing dimension information, and / or the target object (2, 12, 22) is assigned at least one expectation dimension of the target object (2, 12, 22) from a temporally previous determination of the extent information. [9] Method according to claim 8, characterized by , that if only one vertically extending side of the target object (2, 12, 22) is visible in the camera image (1, 11, 21), the separation information is used to determine the ground area covered by the target object (2, 12, 22) on the ground plane (30). - a first calculation point (35) as the lower left corner point of the rectangle (4, 14, 24) in the camera image (1, 11, 21) and a second calculation point (36) as the lower right corner point of the rectangle (4, 14, 24) in the camera image (1, 11, 21) are determined, - using the recording geometry model for the first and second calculation points (35, 36), rays on which the calculation points (35, 36) lie are determined, and the respective intersection points of the rays with the ground plane (30) are determined as the first and second ground corners in three-dimensional space, and - two search lines perpendicular to a connecting line connecting the first and second corner points of the ground in the ground plane (30) through the respective corner points of the ground and third and fourth corner points of the ground along the respective search line in a direction away from the camera (39) when the camera image (1, 11, 21) is taken at a distance in three-dimensional space determined according to the dimensional information. [10] Method according to claim 9, characterized by , that for search lines that initially extend directly out of the rectangle (4, 14, 24) in their projection into the camera image (1, 11, 21), the ground plane (30) and / or the rectangle (4, 14, 24) are corrected in such a way that the projections of the search lines run at least up to the distance within the rectangle (4, 14, 24) in the camera image (1, 11, 21). [11] Method according to claim 9 or 10, characterized by, that if a side classification is present as part of the evaluation information, the expected dimension in the direction perpendicular to the vertically running, classified side visible in the camera image (1, 11, 21) is selected from the dimension information as a distance. [12] Method according to any one of claims 8 to 11, characterized by , if the target object (2, 12, 22) is not fully captured in the camera image (1, 11, 21), especially if there is only one identifiable corner of the ground, missing information for the reconstruction of the ground surface is taken from the dimensional information. [13] Method according to any of the preceding claims, characterized by, that a roof surface of the cuboid associated with the ground surface is determined by the set of roof vertices resulting from the addition of a common height value multiplied by the normal unit vector of the ground plane (30) to the ground vertices, whose projections into the camera image (1, 11, 21) at maximum height value for at least one roof vertex, in particular all roof vertex, lie within the rectangle (4, 14, 24). [14] Method according to any of the preceding claims, characterized by , that the evaluation algorithm uses a trained function, in particular a neural network, which was specifically trained with manually annotated camera images as training data. [15] Method according to any of the preceding claims, characterized by, that for at least some of the target objects (2, 12, 22) not fully captured in the camera image (1, 11, 21) the evaluation algorithm is designed, in particular due to appropriate training, to estimate the complete rectangle (4, 14, 24) extending out of the camera image (1, 11, 21). [16] Motor vehicle (38) comprising a control unit (40) designed to carry out a method according to one of the preceding claims and the camera (39). [17] Computer program which performs the steps of a method according to any one of claims 1 to 15 when executed on a computing device. [18] Electronically readable data carrier on which a computer program according to claim 17 is stored.
Citation Information
Patent Citations
Method for estimating a relative position of an object in the surroundings of a vehicle and electronic control unit for a vehicle and vehicle
EP3594902A1
Atomic scenes for scalable traffic scene recognition in monocular videos
US20160140400A1
System and method for lifting 3D representations from monocular images
US20200160033A1