Device, vehicle and method for determining three-dimensional data of an environment of a vehicle
By evaluating the height and depth values of image points using camera units and artificial neural networks, the complex problem of 3D data acquisition in existing technologies is solved, achieving efficient 3D data acquisition and obstacle recognition.
Patent Information
- Application Number
- CN202211536013.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-14
- Filing Date
- 2022-12-01
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-12-01
AI Technical Summary
Existing technologies require the identification of candidate objects before calculating their height and depth, resulting in a complex process with a large computational load, making it difficult to efficiently determine the three-dimensional data of the vehicle environment.
An environmental image is acquired using a camera unit, and the height value of the image points is evaluated by a first artificial neural network. The depth value of the image points is evaluated by a second artificial neural network. The allocation of computing resources is optimized using a multi-task learning training method, and the resolution is adjusted according to the image precision region to achieve the acquisition of three-dimensional data of individual environmental images.
It enables the rapid and simple acquisition of vehicle environment height and depth data from individual environmental images, improving computational efficiency and resource utilization, and supporting 3D data acquisition and obstacle recognition of the vehicle environment.
Smart Images

Figure CN116263324B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a device for determining three-dimensional data / topography of an environment of a vehicle, a vehicle having a device for determining three-dimensional data of an environment of a vehicle and a method for determining three-dimensional data of an environment of a vehicle. BACKGROUND
[0002] It is essential for autonomous driving to understand the three-dimensional structure of the vehicle's environment. For this purpose, for example, classic camera-based methods, such as motion recovery structure algorithms or stereo camera methods, are used: these methods rely on a plurality of environment images of the vehicle's environment and are computationally intensive. In recent years, methods based on machine learning have been developed which are able to determine an image point-precise estimate of the distance between a world point of the vehicle's environment, which is the basis for a corresponding image point of an environment image, and a camera unit. This method includes so-called monocular depth methods.
[0003] For autonomous driving, it is important to determine information about the three-dimensional data of the vehicle's environment, so as to be able to distinguish, for example, the drivable area of a road ("free space") from the non-drivable area of a road, or to recognize objects on the road, such as a fallen bumper. Here, the three-dimensional data describes the total set of world point positions in the vehicle's environment of the vehicle. The respective position can include a respective height value of the world point relative to the ground of the vehicle's environment and / or the distance of the world point from the camera unit.
[0004] EP 3 627 378 A1 describes a system and method for automatically identifying unstructured objects on a lane of a vehicle, in particular a vehicle having a driving assistance system. The system is configured to classify the objects as drivable or non-drivable by determining the height of the object relative to the lane, which is assumed to be the ground plane. The system comprises a determination module configured to determine apparent portions of an image as candidate objects and a depth module configured to determine a depth value for the candidate objects. What is required for the calculation of the depth value is the pre-identification of the candidate objects before the determination of the depth value. The depth value is calculated by a trained neural network or a method for motion recovery structure.
[0005] US 2020 / 0380316A1 describes a system and method for estimating the height of an object from a monocular image. The system is designed to determine the size of an object, particularly its height, using a trained machine learning algorithm. The method includes the steps of: identifying candidate objects; processing portions of the image relevant to the candidate objects by cropping and / or rescaling the portion; and feeding the thus-processed image of the candidate object to the trained machine learning algorithm to determine the size of the object. The machine learning algorithm is trained based on depth data from a depth sensor, such as LiDAR (Light Imaging, Detection, and Ranging).
[0006] In the above method, candidate objects must first be identified before their size and height can be calculated. Summary of the Invention
[0007] The object of this invention is to provide a method that can more easily determine depth or height data of a vehicle's environment.
[0008] This objective is achieved through the subject matter of the independent claims. Advantageous improvements of the invention are derived from the features of the dependent claims, the following description, and the accompanying drawings.
[0009] A first aspect of the invention relates to an acquisition device for determining three-dimensional data of a vehicle environment. The acquisition device includes a camera unit designed to acquire at least one environmental image of the vehicle environment. The corresponding environmental image has image points located at corresponding image coordinates. In other words, the camera unit of the acquisition device is designed to acquire at least one region of the vehicle environment using the corresponding environmental image. The environmental image has image points. The image points acquired by the camera unit can be projections of the corresponding world point of the vehicle environment onto the image plane of the environmental image. Therefore, the image points can be projections of the corresponding world point from three-dimensional space in the world coordinate system onto the two-dimensional image plane of the corresponding sensor unit in the image coordinate system. The environmental image can be, for example, a graph of image points. The vehicle's coordinate system is also referred to as the world coordinate system. To enable transformations between the two coordinate systems, a rotation / transformation between the acquisition camera coordinate system and the world coordinate system is required. Rotation can be described by a rotation matrix or Euler angles. Here, when using Euler angles, nutation angle, precession angle, and rotation angle are widely used. They can be related to the longitudinal, transverse, or vertical axes of the camera and / or the vehicle.
[0010] This invention proposes that an acquisition device includes an evaluation unit designed to process a corresponding environmental image according to a predetermined height determination method. The environmental image is input into a first artificial neural network in the predetermined height determination method, which is trained to assign corresponding image points of the corresponding environmental image to corresponding height values relative to a predetermined horizontal plane in the world coordinate system of the vehicle environment. In other words, this invention proposes that the evaluation unit of the acquisition device evaluates the acquired corresponding environmental image. This evaluation unit is designed to execute the predetermined height determination method to determine the corresponding height values of corresponding image points in the environmental image.
[0011] Here, the height value is described in a world coordinate system related to the vehicle's environment. The height value determined for the corresponding image point describes the height of the world point in the vehicle environment that serves as the basis for the image point. In the predetermined height determination method, an environment image with the image points is fed to a first artificial neural network. The first artificial neural network can be a neural network that can be trained according to a predetermined training method to determine the corresponding height value for the image point. The first artificial neural network can assign a corresponding height value to each image point, which represents the distance of the image point in the world coordinate system relative to a predetermined horizontal plane of the vehicle environment. For example, the horizontal plane can be the horizontal plane below the vehicle. The height value can also be defined relative to another plane acquired in the vehicle environment.
[0012] In this context and hereinafter, an artificial neural network can be understood as software code stored on a computer-readable storage medium and representing one or more networked artificial neurons or capable of reproducing their functionality. The software code may also contain multiple software code components, which may, for example, have different functions. In particular, an artificial neural network can implement a nonlinear model or nonlinear algorithm that maps inputs to outputs, wherein the inputs are given by input feature vectors or input sequences / numbers, and the outputs may, for example, include the category of the output for a classification task, one or more predicted values, or a predicted sequence.
[0013] The advantage of this invention is that the height value of the vehicle environment can be obtained by evaluating a unique environmental image.
[0014] The present invention also includes improvements that provide additional advantages.
[0015] An improved embodiment of the present invention specifies that the evaluation unit is designed to acquire image points with corresponding lateral resolutions defined by predetermined precision regions. In other words, the environmental image includes predetermined precision regions, wherein the precision regions define corresponding distances between image points. Alternatively, the environmental image includes precision regions whose lateral resolutions differ from one another. Lateral resolution refers to the distance between image points within a corresponding precision region, where a higher lateral resolution describes a smaller distance between image points. It can be specified that the precision region including the lower half of the environmental image predefines smaller distances between image points compared to the precision region including the upper half of the environmental image. This results in the advantage that precision regions of the environmental image that are less important for determining height and / or depth values can be assigned lower resolutions, such that these precision regions require fewer resources in the evaluation than more important image regions. For example, it can be specified that, to determine road orientation, the precision region showing roads is more important than the precision region showing buildings. To save computational resources, it can be specified that the precision region showing buildings is acquired when the distance between image points is greater than the precision region showing roads in the environmental image.
[0016] An improved embodiment of the present invention specifies that an acquisition device is designed to determine the height value of an image point within a predetermined precision region of an environmental image at a height value resolution corresponding to the height value defined by the predetermined precision region. The improved embodiment of the present invention specifies that the predetermined precision region defines a corresponding height value resolution for the height value. In other words, it specifies that the precision region of the environmental image defines the height value determined for the image point within the corresponding precision region. For example, it can be specified that the height value resolution defined within the precision region of an environmental image showing a road is determined with centimeter precision. The precision region showing a building can be specified with decimeter or meter precision for determining the height value. This results in the advantage that the computational resources of the acquisition device can be focused on the precision region relevant to the evaluation.
[0017] An improved embodiment of the invention specifies that the evaluation unit is designed to determine a height value determined according to the predetermined height determination method with a corresponding height value resolution, wherein the corresponding height value resolution depends on the magnitude of the corresponding height value. For example, it can be specified that the evaluation unit determines and / or outputs height values with different height value resolutions. For example, it can be specified that the higher the corresponding height value, the lower the height value resolution of the corresponding height value. This results in advantages such as outputting heights with lower accuracy that are less correlated with route determination compared to other height values. It can also be specified that the height determination method continues with lower accuracy whenever, for example, a height value exceeding a predetermined threshold is determined during a method step. Thus, compared to objects far from the ground, for example, a smaller height value can be determined and / or output with a higher height value resolution, for example, a smaller height value that might be determined for objects close to the ground, such as curb edges or potholes. The height value resolution below the threshold can, for example, be precisely predetermined to be one centimeter. The height value resolution above the threshold can be precise to one decimeter or one meter.
[0018] An improved embodiment of the present invention specifies that the acquisition device is designed to process a corresponding environmental image according to a predetermined depth determination method. To this end, the evaluation unit may be designed to input the environmental image into a second artificial neural network in the predetermined depth determination method. The second artificial neural network is trained to assign corresponding distance or depth values in the world coordinate system of the vehicle environment to image points of the corresponding environmental image. The depth value is described in the world coordinate system, which is related to the vehicle's environment. Therefore, the depth value determined for the image point describes the distance / depth of a world point in the vehicle environment that serves as the basis for the corresponding image point. Here, depth can be the distance from the corresponding world point, which serves as the basis for the corresponding image point, to the camera unit, particularly the distance extending in the horizontal plane. In other words, the acquisition device is designed to evaluate the corresponding environmental image according to a predetermined distance determination method or depth determination method in order to determine the corresponding depth value for the corresponding image point. For example, the depth value can describe the distance from the corresponding world point in the world coordinate system of the vehicle environment to the camera unit. For example, a predetermined second artificial neural network can be trained according to a predetermined training method, wherein the environmental image and the associated acquired depth value can be fed to the second artificial neural network during the training method. The second artificial neural network can determine the corresponding depth value based on the training performed on the corresponding image points. This improved scheme has the advantage that, in addition to determining the height value, the depth value of the image points can also be determined by the acquisition device.
[0019] An improved embodiment of the present invention specifies that the evaluation unit is designed to determine a depth value determined according to a predetermined depth determination method with a corresponding depth value resolution, wherein the corresponding depth value resolution depends on the magnitude of the corresponding depth value. For example, it can be specified that the evaluation unit determines or outputs depth values with different depth value resolutions. For example, it can be specified that the larger the corresponding depth value, the lower the depth value resolution of the corresponding depth value. This results in advantages such as allowing depth values that are less relevant to route determination compared to other depth values to be output with lower precision. It can also be specified that the depth determination method continues with lower precision as long as, for example, a depth value exceeding a corresponding predetermined threshold is determined during a method step. The depth value resolution below the predetermined threshold (depth) can be precisely predetermined to one centimeter, for example. The depth value resolution above the threshold (depth) can be precise to one decimeter or one meter.
[0020] An improved embodiment of the present invention specifies that the evaluation unit is designed to determine the depth value of an image point within a predetermined precision region at a depth value resolution corresponding to the depth value defined by the predetermined precision region. The improved embodiment of the present invention specifies that the predetermined precision region defines a corresponding depth value resolution for the depth value. In other words, it specifies the precision at which the depth value of the image point located within the corresponding precision region of the environmental image is determined. For example, it can be specified that a higher precision of the depth value is predetermined in a precision region including a parked car, so that the distance to the car can be reliably determined, for example, for parking assistance. A lower depth value precision can be assigned to a precision region showing a distant object, because it is not important to determine the distance value more accurately for such a relevant precision region when performing a parking maneuver.
[0021] An improved embodiment of the present invention specifies that the acquisition device is designed to determine the vehicle's travel path using height and / or depth values according to a predetermined path determination method. In other words, the acquisition device is designed to determine a path for the vehicle within a drivable area of the vehicle environment. For example, the acquisition device is designed to determine drivable surfaces based on specific three-dimensional data. This allows for the determination of drivable surfaces independently of common methods such as semantic segmentation or object recognition. Therefore, the acquisition device is designed to determine areas where no obstacles are present. Obstacles can be identified using the acquired three-dimensional data.
[0022] An improved embodiment of the present invention specifies that the evaluation unit is designed to check the consistency between the height value and the depth value of a corresponding image point according to a predetermined inspection method. In other words, the evaluation unit checks the compatibility between the height values and depth values of at least some corresponding image points. For example, it can be specified that the coordinates of an image point in the world coordinate system are determined based on the position of one of the image points in the environmental image and the determined height value of that image point. Similarly, the coordinates of the image point in the world coordinate system can be determined from the position of the image point in the environmental image and the determined depth value of that image point. In the predetermined inspection method, the coordinates of the image point in the world coordinate system determined by the height value can be compared with the coordinates of the image point in the world coordinate system determined by the depth value. Here, for example, the distance between the two determined coordinates can be determined. If the distance is less than a predetermined threshold, the height value and depth value can be evaluated as compatible with each other.
[0023] An improved embodiment of the present invention specifies that the evaluation unit is designed to evaluate the height and / or depth values of image points according to a predetermined object recognition method to identify predetermined objects in a vehicle environment. The predetermined object recognition method may be, for example, a predetermined semantic segmentation method designed to assign image points of an environmental image to predetermined semantic categories based on determined height and / or determined depth values. The predetermined object recognition method may also be a method according to existing technology designed to identify predetermined objects in a point cloud. The predetermined semantic categories may, for example, relate to houses, vehicles, drivable surfaces, curb edges, or other objects. This results in the advantage that objects in the vehicle environment can be acquired considering both height and / or depth values.
[0024] A second aspect of the invention relates to a vehicle having a device for acquiring three-dimensional data for determining the vehicle's environment. The vehicle according to the invention is preferably designed as a motor vehicle, particularly as a passenger car or commercial vehicle, or as a bus or motorcycle.
[0025] A third aspect of the invention relates to a method for operating an acquisition device for determining three-dimensional data of a vehicle environment. In this method, at least one camera unit of the acquisition device acquires at least one corresponding environmental image of the vehicle environment. The corresponding environmental image has image points located at corresponding image coordinates of the environmental image. It is specified that an evaluation unit of the acquisition device processes the corresponding environmental image according to a predetermined height determination method, wherein the environmental image is guided to a first artificial neural network in the predetermined height determination method, the first artificial neural network being trained to assign corresponding height values for the corresponding image points of the corresponding environmental image relative to a predetermined horizontal plane in the world coordinate system of the vehicle environment. It is specified that corresponding height values are assigned to the corresponding image points of the environmental image, and the relevant height values are output by the acquisition device.
[0026] A fourth aspect of the invention relates to a method for training a first artificial neural network. It specifies that a camera unit of an acquisition device acquires at least one environmental image of a vehicle environment. The corresponding environmental image includes image points located at corresponding image coordinates of the environmental image. An environment detection unit acquires the corresponding height values of world points of the vehicle environment in the world coordinate system of the vehicle environment relative to a predetermined horizontal plane of the vehicle environment. An evaluation unit assigns a corresponding world point to each image point. Thus, a corresponding height value of the corresponding world point can be assigned to each image point. The world point to which the image point is based can be determined based on geometric calculations. The environmental image including the image points and the determined height values are input into the first artificial neural network. The first artificial neural network enables the model for determining the height values of the image points to be updated.
[0027] The improved scheme specifies that the environment detection unit acquires the corresponding depth values of world points in the vehicle environment within the vehicle's world coordinate system, which describe the distance between the image points and the camera unit in the world coordinates. The evaluation unit assigns the corresponding world points upon which the image points are based. Thus, the corresponding depth values of the corresponding world points can be assigned to the corresponding image points. The world points upon which the corresponding image points are based can be determined based on geometric calculations. The environmental image including the image points and the determined depth values are input into a second artificial neural network. The second artificial neural network enables the model for determining the depth values of the corresponding image points to be updated.
[0028] The improved scheme specifies that the first and second neural networks are trained in parallel using a multi-task learning training method. In other words, it specifies that a multi-task learning training method is used to train the first and second neural networks. The first and second neural networks can, for example, have the same base layer in their respective models. The base layer of the respective models can be non-specific, allowing specific layers to be built on top of the base layer, which can be configured to determine height or depth values. For example, the first and second neural networks can differ from each other only in specific layers. Therefore, in training the first neural network, the base layer of the first neural network, which can also be used by the second neural network, and the specific layers used only by the first neural network can be trained. Similarly, in training the second neural network, the base layer of the second neural network, which can also be used by the first neural network, and the specific layers used only by the second neural network can be trained.
[0029] The present invention also relates to an evaluation unit for acquiring a device. The evaluation unit may have a data processing device or a processor device designed to perform embodiments of the method according to the invention. The processor device may therefore have at least one microprocessor and / or at least one microcontroller and / or at least one FPGA (Field Programmable Gate Array) and / or at least one DSP (Digital Signal Processor). Furthermore, the processor device may have program code designed to perform embodiments of the method according to the invention when implemented by the processor device. The program code may be stored in the data memory of the processor device.
[0030] The present invention also relates to improvements to the vehicle and method according to the invention, which have features already described in conjunction with improvements to the acquisition device according to the invention. For this reason, corresponding improvements to the vehicle and method according to the invention will not be described here.
[0031] The present invention also includes combinations of features of the described embodiments. Therefore, the present invention also includes implementations having combinations of features of a plurality of described embodiments, provided that these embodiments are not described as mutually exclusive. Attached Figure Description
[0032] Embodiments of the present invention are described below. Therefore:
[0033] Figure 1 A schematic diagram of a vehicle equipped with an acquisition device is shown.
[0034] Figure 2 A schematic diagram of the flow for running a method to acquire a device is shown.
[0035] Figure 3A schematic diagram of the process for training first and second artificial neural networks is shown.
[0036] Figure 4 A schematic diagram of the environment image is shown, along with the depth and height values determined for image points in the environment image. Detailed Implementation
[0037] The embodiments described below are preferred embodiments of the present invention. In these embodiments, the components described are each considered independent features of the present invention, and each also independently improves upon the present invention. Therefore, this disclosure should also include combinations different from those shown in the embodiments. Furthermore, the described embodiments may be supplemented by other features of the present invention already described.
[0038] In the accompanying drawings, the same reference numerals denote elements that have the same function.
[0039] Figure 1 A schematic diagram of a vehicle equipped with an acquisition device is shown. Vehicle 1 may be, for example, a passenger car or a commercial vehicle. Acquisition device 2 may include, for example, a camera unit 3 and an environment detection unit 4. Camera unit 3 may be, for example, a monocular camera. Camera unit 3 may be designed to acquire a corresponding environmental image 5 of the vehicle environment 6 of vehicle 1. Environmental image 5 may be designed as a pixel graphic / image point graphic and has image points 7, which correspond to world points 8 in the vehicle environment 6. Acquisition device 2 may include an evaluation unit 9, which may be designed to receive at least one corresponding environmental image 5 from at least one camera unit 3 and evaluate the environmental image according to a predetermined height determination method.
[0040] The height determination method may include a height calculation method for determining the height value 10 of a corresponding image point 7 and a depth determination method for determining the depth value 11. The height value 10 of the corresponding image point 7 may refer to a predetermined plane 12 in the vehicle environment 6. This plane may, for example, be a predetermined area below the vehicle 1. The depth value 11 may be the distance between the image point 7 and the camera unit 3 or a general sensor device. To determine the depth value 11 in the predetermined depth determination method, the environmental image 5 may be assigned to a second artificial neural network 14, which may be trained to determine the depth value 11 of the image point 7. To determine the height value 10, the environmental image 5 may be assigned to a first neural network 13, which may be trained to determine the height value 10 of the corresponding image point 7.
[0041] The determined height value 10 and / or depth value 11 can be used to identify objects 15 in the vehicle environment 6. The predetermined objects 15 can be, for example, buildings, pedestrians, other traffic participants, vehicles 1, or curbs. Object identification 15 can be performed, for example, by semantic segmentation methods. It can be specified that the height value 10 or depth value 11 is used to determine the area 16 drivable by vehicle 1 and to determine the driving path 17 for vehicle 1. For training the first artificial neural network 13 and / or the second artificial neural network 14, it can be specified that the acquisition device 2 may have an environment detection unit 4. The environment detection unit can be designed, for example, as a lidar sensor unit, a stereo camera unit 3, an ultrasonic sensor unit, or a monocular camera, designed to determine the height value 10 or depth value 11 using a motion reconstruction structure method. The environment detection unit 4 can be designed to assign corresponding height values 10 or depth values 11 to image points 7 of the environment image 5 of the vehicle environment 6. The determined height value 10 and / or depth value 11 can be transmitted along with the captured environmental image 5 to the first and / or second artificial neural networks 13, 14, wherein the height value 10 and / or depth value 11 can be assigned to image points 7 of the environmental image 5. For example, this assignment is also referred to as labeling. Therefore, it is possible to train the model or layer 12 of the neural networks 13, 14 to determine the height value 10 and / or depth value 11 for the corresponding image points 7 of the environmental image 5 by performing training multiple times.
[0042] Figure 2 A schematic diagram of the method flow for operating the acquisition device is shown. In the first step S1, it can be specified that the camera unit 3 of the acquisition device 2 acquires an environmental image 5 of the vehicle environment 6.
[0043] In step S2A, the environmental image 5 can be transmitted to the evaluation unit 9, which can perform a predetermined height determination method to determine the corresponding height value 10 of the corresponding image point 7 of the environmental image 5. The predetermined height determination method may specify that the environmental image 5 with image point 7 is transmitted to a first neural network 13, which can be trained to assign height values 10 to the image points 7 of the environmental image 5. The first neural network 13 can determine the height value 10 of the image point 7, for example, relative to a predetermined plane 12.
[0044] Additionally or alternatively to step S2A, in step S2B, which operates in parallel with step S2A, the environmental image 5 having image point 7 is transmitted to the evaluation unit 9. The evaluation unit executes a predetermined depth determination method to assign a corresponding depth value 11 to the corresponding image point 7. The depth value 11 may, for example, describe the distance between the world point 8 assigned to the corresponding image point 7 and the camera unit 3. In the predetermined depth determination method, the environmental image 5 having image point 7 can be input into a second artificial neural network 14, which can be trained to assign depth values 11. The determined depth value 11 can be output by the acquisition device 2.
[0045] In step S3, for example, the coordinates of the world point 8 assigned to the image point 7 can be derived based on the height value 10 and / or the depth value 11, thereby determining the three-dimensional data of the vehicle environment 6.
[0046] In step S4, it can be specified that object 15 is determined in a predetermined object recognition method. This can be done, for example, by a semantic segmentation method. In step S4, a drivable surface 16 that vehicle 1 can traverse can be determined. It can be specified that, additionally, a travel path 17 of vehicle 1 can be determined, which can extend within the drivable surface.
[0047] Vehicle 1 can be controlled according to the determined driving path 17 to guide the vehicle along the driving path 17.
[0048] Figure 3 A schematic diagram of the process for training the first and second artificial neural networks is shown.
[0049] In the first step L1A, the camera unit 3 of the acquisition device 2 of vehicle 1 can acquire an environmental image 5 of the vehicle environment 6, which includes image points 7.
[0050] In parallel with this, in step L1B, the environment detection unit 4 of the acquisition device 2 can determine the height value 10 and depth value 11 of the world point 8 of the vehicle environment 6. This can be achieved, for example, by lidar, radar, ultrasound, or stereo imaging technology. Thus, the corresponding height value 10 and / or the corresponding depth value 11 can be acquired for each world point 8 of the vehicle environment 6.
[0051] In step L2, the environment image 5 with image point 7 can be merged together with the acquired world point 8.
[0052] In step L3A, the environment image 5 with image point 7 is assigned world point 8 and corresponding height value 10 according to a predetermined allocation method. In step L4A, image point 7 and associated height value 10 can be input into the first neural network 13, which can then update the model or network layer of the first neural network 13 in the height determination method, or in the sense of the height determination method.
[0053] In step L3B, the world point 8 of the vehicle environment 6 and its corresponding depth value 11 can be assigned to the corresponding image point 7 of the environment image 5. The corresponding depth value 11 and the corresponding image point 7 can be input into the second neural network 14 in step L4B, which can then update the model or network layer of the second neural network 14 in the depth determination method or in the sense of the depth determination method.
[0054] It can be stipulated that so-called multi-task learning is used. The following fact can be applied here: since the differences between the two neural networks 13 and 14 are small, the base layers of neural networks 13 and 14 can be trained together.
[0055] Figure 4 A schematic diagram of an environmental image and the depth and height values determined for image points in the environmental image are shown. The upper image shows an environmental image 5 captured by camera unit 3 of acquisition device 2 of vehicle environment 6 of vehicle 1. The middle image shows the depth value 11 determined for environmental image 5. The lower image shows the height value 10 determined for environmental image 5. Environmental image 5 may include two precision regions 17a and 17b. Through the two precision regions 17a and 17b, it is possible to predetermine the depth value 11 of image point 7 and / or the height value 10 of image point 7 determined by evaluation unit 9. The upper precision region 17a may specify a lower resolution because the corresponding area of environmental image 5 does not include objects related to vehicle guidance. The lower precision region 17b may specify a higher resolution because the corresponding area of environmental image 5 includes objects related to vehicle guidance, such as curb stones. In order to reliably identify objects by evaluation unit 9 in the object recognition method, a resolution within centimeter precision may be predetermined for height value 10 and depth value 11, for example.
[0056] The described method allows the height of each image point 7 in a single image to be determined, regardless of whether that image point is associated with object 15. The height value 10 can be used, in particular, to estimate the passable path of vehicle 1 by determining the height of any possible curb stones. Additionally or alternatively, the height value 10 can be sent as supplementary information to other systems, such as object recognition systems or semantic segmentation systems. In the case of an object recognition system, the height value 10 can, for example, help define the bounding box and / or size of object 15. In a semantic segmentation system, the height value 10 can be used to better classify passable paths.
[0057] Overall, this example demonstrates how a method can be provided to determine the height values of individual image points from a unique environmental image.
Claims
1. A device (2) for acquiring three-dimensional data to determine a vehicle environment (6), wherein, The acquisition device (2) has a camera unit (3) designed to acquire at least one environmental image (5) of the vehicle environment (6). The corresponding environmental image (5) has image points (7), which are located at the corresponding image coordinates of the environmental image (5). Its features are, The acquisition device (2) has an evaluation unit (9) designed to determine, for the corresponding image point (7), the corresponding height value (10) relative to a predetermined horizontal plane (12) in the world coordinate system of the vehicle environment (6) according to a predetermined height determination method. The evaluation unit (9) is designed to input an environmental image (5) into a first artificial neural network (13) in a predetermined height determination method. The first artificial neural network has been trained to determine the corresponding height value (10) of the corresponding image point (7) of the corresponding environmental image (5) in the world coordinate system of the vehicle environment (6) relative to a predetermined horizontal plane (12). The evaluation unit (9) is designed to determine, for a given image point (7), the corresponding depth value (11) in the world coordinate system of the vehicle environment (6) according to a predetermined depth determination method. This depth value describes the horizontal distance between the image point (7) and the camera unit (3) in the world coordinate system. The evaluation unit (9) is designed to input an environmental image (5) into a second artificial neural network (14) in a predetermined depth determination method. The second artificial neural network has been trained to determine the corresponding depth value (11) in the world coordinate system of the vehicle environment (6) for the corresponding image point (7) of the corresponding environmental image (5). The evaluation unit (9) is designed to check the consistency between the corresponding height value (10) and the corresponding depth value (11) according to a predetermined inspection method.
2. The acquisition device (2) according to claim 1, characterized in that, The acquisition device (2) is designed to acquire image points (7) in a predetermined precision area (17a, 17b) of an environmental image (5) at a corresponding lateral resolution defined by the predetermined precision area (17a, 17b).
3. The acquisition device (2) according to claim 2, characterized in that, The acquisition device (2) is designed to determine the height value (10) of the image point (7) in the predetermined precision region (17a, 17b) of the environmental image with a height value resolution of the corresponding height value (10) defined by the predetermined precision region (17a, 17b).
4. The acquisition device (2) according to claim 1 or 2, characterized in that, The evaluation unit (9) is designed to determine the height value (10) determined according to the predetermined height determination method with a corresponding height value resolution, wherein the corresponding height value resolution depends on the size of the corresponding height value (10).
5. The acquisition device (2) according to claim 1, characterized in that, The evaluation unit (9) is designed to determine a depth value (11) determined according to the predetermined depth determination method with a corresponding depth value resolution, wherein the corresponding depth value resolution depends on the magnitude of the corresponding depth value (11).
6. The acquisition device (2) according to claim 2, characterized in that, The acquisition device (2) is designed to determine the depth value (11) of an image point (7) in a predetermined precision region (17a, 17b) with a depth value resolution of a depth value (11) defined by the predetermined precision region (17a, 17b).
7. The acquisition device (2) according to claim 1 or 2, characterized in that, The evaluation unit (9) is designed to determine the driving path (17) of the vehicle (1) by height value (10) and / or depth value (11) according to a predetermined path determination method.
8. The acquisition device (2) according to claim 1 or 2, characterized in that, The evaluation unit (9) is designed to evaluate the height value (10) and / or depth value (11) of image points according to a predetermined object recognition method in order to identify a predetermined object (15) in the vehicle environment (6).
9. A vehicle (1) comprising an acquisition device (2) according to any one of claims 1 to 8.
10. A method for operating an acquisition device according to any one of claims 1 to 8, wherein, At least one environmental image (5) of the vehicle environment (6) is acquired by the camera unit (3) of the acquisition device (2). The corresponding environmental image (5) has image points (7), which are located at the corresponding image coordinates of the environmental image (5). Its features are, The evaluation unit (9) of the acquisition device (2) determines the corresponding height value (10) of the corresponding image point (7) according to the predetermined height determination method. An environmental image (5) is input into a first artificial neural network (13) in a predetermined height determination method. The first artificial neural network has been trained to assign corresponding image points (7) of the corresponding environmental image (5) to corresponding height values (10) in the world coordinate system of the vehicle environment (6) relative to a predetermined horizontal plane (12). The evaluation unit (9) assigns a corresponding height value (10) to the corresponding image point (7) of the image.
11. A method for training a first artificial neural network (13), the first artificial neural network being a first artificial neural network belonging to the acquisition device according to any one of claims 1 to 8, wherein, At least one environmental image (5) of the vehicle environment (6) is acquired by the camera unit (3) of the acquisition device (2). The corresponding environmental image (5) includes image points (7), which are located at the corresponding image coordinates of the environmental image (5). The environmental detection unit (4) obtains the corresponding height value (10) of the world point (8) of the vehicle environment (6) relative to the predetermined horizontal plane (12) in the world coordinate system of the vehicle environment (6). Assign the corresponding world point (8) on which the corresponding image point (7) is based to the corresponding image point (7). The corresponding image points (7) of the environment image (5) are assigned the corresponding height values (10) of the world points (8). In the training method, the image points (7) of the environment image (5) along with the corresponding determined height values (10) are input into the first artificial neural network (13).
12. The method according to claim 11, characterized in that, The environment detection unit (4) acquires the corresponding depth value (11) of the world point (8) of the vehicle environment (6) in the world coordinate system of the vehicle environment (6). This depth value describes the distance of the image point (7) to the camera unit (3) in the world coordinate system. Assign the corresponding world point (8) on which the corresponding image point (7) is based to the corresponding image point (7). The corresponding depth values (11) of the corresponding world points (8) are assigned to the corresponding image points (7) of the environment image (5). In the training method, the image points (7) of the environment image (5) along with the corresponding determined depth values (11) are input into the second artificial neural network (14).
13. The method according to claim 12, characterized in that, The first artificial neural network (13) and the second artificial neural network (14) are trained together in a multi-task learning training method.
Citation Information
Patent Citations
Agricultural unmanned vehicle obstacle sensing method and device based on monocular camera
CN112184700A
Object height estimation from monocular images
US20200380316A1
Intent-based dynamic change of region of interest of vehicle perception system
US20210097304A1
Obstacle sensing device, obstacle sensing method, and obstacle sensing program
WO2013035612A1