METHOD AND SYSTEM FOR ESTIMATE IN-DEPTH INFORMATION

DE502022007668D1Active Publication Date: 2026-04-30AUMOVIO AUTONOMOUS MOBILITY GERMANY GMBH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
AUMOVIO AUTONOMOUS MOBILITY GERMANY GMBH
Filing Date
2022-03-24
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Current 3D environmental sensing using stereo camera systems is hindered by uneven illumination caused by parallax, leading to difficulties in determining depth information in shadowed areas.

Method used

Utilizing a convolutional neural network to analyze geometric information from unevenly illuminated image areas, combining triangulation with geometric evaluation to estimate depth information, and adjusting depth data based on uneven illumination patterns.

Benefits of technology

Enables accurate and robust three-dimensional environment perception even in areas where triangulation is not possible, improving separation of foreground and background objects and enhancing depth determination.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method and a system for determining depth information for image information provided by imaging sensors of a vehicle, using an artificial neural network.

[0002] It is generally known that imaging sensors are used to capture the vehicle's surroundings in three dimensions. Stereo cameras are among the technologies used for 3D environmental perception. To calculate distance information, the image data provided by the two cameras is correlated, and the distance of a pixel to the vehicle is determined using triangulation.

[0003] The cameras for the stereo camera system are integrated, for example, into the front of the vehicle. The installation location is usually the windshield area or the radiator grille. To generate sufficient brightness for image analysis at night, the vehicle's headlights are typically used.

[0004] A problem with current 3D environmental sensing is that unevenly illuminated image areas in the image data acquired by the stereo camera system make it difficult to determine depth information, as the stereo camera system cannot obtain distance information in these unevenly illuminated areas. This is particularly true when the different installation positions between the headlights and the cameras result in shadows caused by parallax.

[0005] Publication US 2020 / 013176 A1 discloses a method for estimating a depth map based on images from a pair of cameras using a neural network.

[0006] Publication CN 112184731 A discloses depth estimation using a stereo camera and a neural network that processes the image information from the stereo camera.

[0007] Based on this, the object of the invention is to provide a method for determining depth information for image information, which enables an improved determination of depth information.

[0008] The problem is solved by a method having the features of independent claim 1. Preferred embodiments are the subject of the dependent claims. A system for determining depth information from image information is the subject of dependent claim 15.

[0009] According to a first aspect, the invention relates to a method for determining depth information from image information using an artificial neural network in a vehicle. The neural network is preferably a convolutional neural network (CNN).

[0010] The scope of protection of the present invention is defined by the attached claims.

[0011] The method comprises the following steps: First, at least one emitter and at least one first and one second receiving sensor are provided. The emitter can be configured to emit electromagnetic radiation in the visible spectrum. Alternatively, the emitter can emit electromagnetic radiation in the infrared spectrum, in the frequency range of approximately 24 GHz or approximately 77 GHz (emitter is a radar emitter), or laser radiation (emitter is a lidar emitter). The first and second receiving sensors are spaced apart from each other. The receiving sensors are adapted to the emitter type; that is, they are configured to receive reflected components of the electromagnetic radiation emitted by the at least one emitter.In particular, the receiving sensors can be designed to receive electromagnetic radiation in the visible or infrared spectral range, in the frequency range of approximately 24 GHz or approximately 77 GHz (radar receivers) or laser radiation (LIDAR receivers).

[0012] Subsequently, the emitter emits electromagnetic radiation, and the first and second receiving sensors receive reflected portions of this radiation. Based on these reflected portions, the first receiving sensor generates initial image information, and the second receiving sensor generates further image information.

[0013] The first and second image data sets are then compared to determine at least one unevenly illuminated image area in the first and second image data sets. This uneven illumination is caused by parallax due to the spaced-apart arrangement of the receiving sensors. If the first and second receiving sensors are not located at the projection center of an emitter, particularly a spotlight, the unevenly illuminated image area can also result from the parallax between the respective receiving sensor and its associated emitter. In other words, at least one image area is identified as an "unevenly illuminated image area" that is brighter or darker in the first image data set than in the second image data set.

[0014] Subsequently, geometric information from at least one unevenly illuminated area of ​​the image is evaluated, and depth information is estimated by the artificial neural network based on the results of this evaluation. In other words, the size or extent of the unevenly illuminated area is assessed, as this allows the neural network to draw conclusions about the three-dimensional shape of an object (e.g., a specific area of ​​the object is closer to the vehicle than another area) or the distance between two objects located in the vicinity of the vehicle.

[0015] The technical advantage of the proposed method lies in the fact that, even in unevenly illuminated areas where triangulation is not possible, the neural network can use the geometric information of this unevenly illuminated image area to infer the distance to one or more objects depicted within and / or around this unevenly illuminated area. This allows for a more accurate and robust three-dimensional environment perception, even against interference.

[0016] According to one embodiment, the unevenly illuminated image area arises in the transition zone between a first and a second object, which are at different distances from the first and second receiving sensors. The estimated depth information is therefore depth difference information, which contains information about the distance difference between the first and second objects and the vehicle. This allows for improved separation of foreground and background objects. A foreground object is an object positioned closer to the vehicle than a background object.

[0017] Furthermore, it is possible that the unevenly illuminated area of ​​the image relates to a single object, with the uneven illumination resulting from the three-dimensional structure of that object. This makes it possible to improve the determination of three-dimensional surface shapes of objects.

[0018] According to one embodiment, the emitter is at least a headlight that emits visible light in the wavelength range between 380 nm and 800 nm, and the first and second receiving sensors are each a camera. This allows the vehicle's existing front lighting and cameras operating in the visible spectral range to be used as detection sensors.

[0019] Preferably, the first and second receiving sensors form a stereo camera system. The image information provided by the receiving sensors is correlated, and based on the installation positions of the receiving sensors, the distance of the respective pixels in the image information from the vehicle is determined. This allows depth information to be obtained for the image areas captured by both receiving sensors.

[0020] According to one embodiment, at least two emitters in the form of the vehicle's headlights are provided, and a receiving sensor is assigned to each headlight such that the line of sight between an object to be detected and the headlight is essentially parallel to the line of sight between an object to be detected and the receiving sensor assigned to that headlight. "Essentially parallel" here means, in particular, an angle of less than 10°. Specifically, the receiving sensor can be located very close to the projection center of its assigned headlight, for example, at a distance of less than 20 cm.This means that the illumination area of ​​the headlight is essentially the same as the detection area of ​​the receiver sensor, resulting in a largely parallax-free installation situation, which leads to homogeneous illumination of the detection area of ​​the receiver sensor without lighting shadows from the headlight assigned to it.

[0021] According to one embodiment, the first and second receiving sensors are integrated into the vehicle's headlights. This ensures that the headlight's illumination range is essentially equal to the receiving sensor's detection range. This results in a completely or almost completely parallax-free installation.

[0022] According to one embodiment, the artificial neural network performs depth estimation based on the horizontally measured width of the unevenly illuminated image area. Preferably, the neural network is trained to use the dependence of the width of the unevenly illuminated image area on the three-dimensional shape of the surrounding area represented by this image area to estimate depth information. In particular, the horizontal width of the unevenly illuminated image area is suitable for determining depth differences relative to the unevenly illuminated image area. The depth difference can refer to a single contoured object or to several objects, where one object (also referred to as the foreground object) is located in front of another object (also referred to as the background object).

[0023] It goes without saying that, in addition to the horizontally measured width of the unevenly illuminated image area, further geometric information and / or dimensions of the unevenly illuminated image area can be determined to estimate depth information. This can include, in particular, a height measured vertically or a dimension measured obliquely (perpendicular to the horizontal).

[0024] According to one embodiment, the artificial neural network determines depth information in image areas captured by the first and second receiving sensors based on triangulation between pixels in the first and second image information and the first and second receiving sensors. Preferably, the artificial neural network determines the depth information by triangulation and also estimates the depth information based on the geometric information of the unevenly illuminated image area. In other words, depth determination by triangulation and the evaluation of geometric information from an unevenly illuminated image area are performed by one and the same neural network. Due to the use of several different mechanisms for determining depth information, improved and more robust three-dimensional environment determination can be achieved.

[0025] According to one embodiment, the neural network compares depth information obtained through triangulation with estimated depth information obtained by evaluating the geometric information of at least one unevenly illuminated image area, and generates adjusted depth information based on this comparison. This advantageously corrects triangulation inaccuracies, resulting in more reliable depth information overall.

[0026] According to one embodiment, the artificial neural network adjusts the depth information obtained through triangulation based on the evaluation of the geometric information of at least one unevenly illuminated image area. In other words, the depth information obtained through triangulation is modified based on the estimated depth information. This results in a more robust three-dimensional environment determination.

[0027] According to one embodiment, IR radiation, radar signals or laser radiation are emitted by the at least one emitter.

[0028] Accordingly, at least some of the receiving sensors can be infrared cameras, radar receivers, or laser receivers. In particular, the receiving sensors are selected according to the at least one emitter to which they are assigned. For example, receiving sensors are designed to receive infrared (IR) radiation when they are assigned to an IR emitter. Specifically, to detect the area to the side of or behind the vehicle, emitters and receiving sensors that do not emit light in the visible wavelength range can be used, as this would disturb other road users. This makes it possible to achieve at least partial all-around detection of the vehicle's surroundings.

[0029] According to one embodiment, to estimate depth information from image data depicting areas to the sides and / or behind the vehicle, more than one emitter and more than two receiver sensors are used to determine image data. Multiple sensor groups are provided, each comprising at least one emitter and at least two receiver sensors, and the image data from the respective sensor groups is combined to form a complete image. This enables at least partial all-around coverage of the vehicle's surroundings.

[0030] According to one embodiment, the sensor groups utilize at least partially electromagnetic radiation in different frequency bands. For example, a stereo camera system can be used in the front of the vehicle, employing an emitter that emits light in the visible spectral range, whereas emitters that utilize IR radiation or radar radiation can be used in the side areas of the vehicle.

[0031] According to a further aspect, the invention relates to a system for determining depth information from image information in a vehicle, comprising a computing unit that performs computational operations of an artificial neural network, at least one emitter configured to emit electromagnetic radiation, and at least one first and one second receiving sensor arranged at a distance from each other. The first and second receiving sensors are configured to receive reflected components of the electromagnetic radiation emitted by the emitter. The first receiving sensor is configured to generate first image information, and the second receiving sensor is configured to generate second image information based on the received reflected components. The artificial neural network is configured to: to compare the first and second image information to determine at least one unevenly illuminated image area in the first and second image information, wherein the unevenly illuminated image area arises from parallax due to the spaced arrangement of the receiving sensors; to evaluate the geometric information of the at least one unevenly illuminated image area and to estimate depth information based on the result of the evaluation of the geometric information of the at least one unevenly illuminated image area.

[0032] If the first and second receiving sensors are not located at the projection center of an emitter, especially a spotlight, the unevenly illuminated image area can also be caused by the parallax between the respective receiving sensor and the emitter assigned to it.

[0033] For the purposes of this disclosure, "image information" refers to any information that enables a multidimensional representation of the vehicle's surroundings. This includes, in particular, information provided by imaging sensors, such as cameras, radar sensors, or lidar sensors.

[0034] For the purposes of this disclosure, "emitter" refers to transmitting units designed to emit electromagnetic radiation. Examples include spotlights, infrared emitters, radar transmitters, and lidar transmitters.

[0035] The terms "approximately", "essentially" or "about" mean, within the meaning of the invention, deviations from the respective exact value by + / - 10%, preferably by + / - 5% and / or deviations in the form of changes that are insignificant for the function.

[0036] Further developments, advantages, and possible applications of the invention will also become apparent from the following description of exemplary embodiments and from the figures. All features described and / or illustrated are, individually or in any combination, fundamentally the subject matter of the invention, irrespective of their compilation in the claims or their cross-reference. The content of the claims is also incorporated into the description.

[0037] The invention will be explained in more detail below with reference to exemplary embodiments shown in the figures. The figures show: Fig. 1 is an exemplary schematic representation of a vehicle with a stereo camera system designed to detect objects in front of the vehicle; Fig. 2 is an exemplary schematic representation of first image information captured by a first detection sensor, on which the two objects and unequally illuminated areas in the transition zone between the objects are recognizable; Fig. 3 is an exemplary schematic representation of second image information captured by a second detection sensor, on which the two objects and unequally illuminated areas in the transition zone between the objects are recognizable; Fig. 4 is an exemplary schematic representation of a vehicle with several sensor groups designed to detect objects in the area surrounding the vehicle; and Fig.5. An example flowchart to illustrate the steps of a procedure for determining depth information from image information using an artificial neural network.

[0038] Figure 1Figure 1 shows an example of a vehicle 1 equipped with a stereo camera system. The stereo camera system comprises a first receiving sensor 4 and a second receiving sensor 5, which are, for example, image recording devices, in particular cameras. Furthermore, the vehicle 1 has a first emitter 3 and a second emitter 3', which are, for example, formed by the headlights of the vehicle 1. The emitters 3 and 3' are therefore designed to emit light visible to humans, in particular in the wavelength range between 380 nm and 800 nm. The receiving sensors 4 and 5 are accordingly designed to receive light in this wavelength range and provide image information. In particular, the first receiving sensor 4 provides first image information B1, and the second receiving sensor 5 provides second image information B2.

[0039] To evaluate the image information B1, B2 provided by the receiving sensors, the vehicle 1 has a computer unit 8, which is designed to evaluate the image information B1, B2. In particular, the computer unit 8 is designed to generate depth information from image information B1, B2 of the at least two receiving sensors 4, 5 in order to enable three-dimensional detection of the environment around the vehicle 1.

[0040] For the evaluation of image information B1, B2, an artificial neural network 2 is implemented in the computing unit 8. The artificial neural network 2 is designed and trained in such a way that it firstly calculates depth information for image information B1, B2 by means of triangulation and then verifies and modifies this calculated depth information using a depth information estimation. This estimation determines unevenly illuminated image areas by comparing image information B1, B2, evaluates their geometry or dimensions, and determines estimated depth information based on this. The depth information calculated by triangulation can then be adjusted based on this estimate.

[0041] In Fig. 1A first object O1 and a second object O2 are shown, located in front of vehicle 1 and illuminated by the vehicle 1's headlights. The receiving sensors 4, 5 can detect the portions of the light emitted by the headlights that are reflected by objects O1 and O2.

[0042] Objects O1 and O2 are at different distances from vehicle 1. Furthermore, from the perspective of vehicle 1 and with respect to the line of sight between objects O1 and O2 and receiving sensors 4 and 5, the second object O2 is located in front of the first object O1. For example, the front face of the second object O2, facing vehicle 1, is positioned a distance Δd in front of the front face of the first object O1, which also faces vehicle 1.

[0043] Due to the spaced arrangement of the emitters 3, 3' (here the headlights of vehicle 1) and the receiver sensors 4, 5, brightness differences arise in the first and second image information B1, B2 due to parallax, i.e. the first receiver sensor 4 provides image information B1 with brightness differences in different areas than in the second image information B2, which is generated by the second receiver sensor 5.

[0044] Figs. 2 and 3 This effect is illustrated in an exemplary and schematic way. Figure 2Figure 1 shows an example of the first image information B1 provided by the first receiving sensor 4, which is located on the left side of the vehicle 1 in the forward direction FR of the vehicle 1. Two unequally illuminated image areas D1, D2 are visible. These arise because the scene depicted by these image areas D1, D2 is illuminated by only one emitter 3, 3' each, and the first receiving sensor 4 sees the objects O1, O2 from the front, looking obliquely from the left. The unequally illuminated image area D2 therefore has a greater width b (measured horizontally) than the unequally illuminated image area D1.

[0045] The Figure 3The second image information B2, provided by the second receiving sensor 5, is shown as an example. This sensor is located on the right side of the vehicle 1 in the forward direction FR of the vehicle 1. The second image information B2 also shows two unequally illuminated image areas D1 and D2. These unequally illuminated areas arise because the scene depicted by these image areas D1 and D2 is illuminated by only one emitter 3 and 3' respectively, and the second receiving sensor 4 views the objects O1 and O2 from the front, obliquely from the right. Consequently, the unequally illuminated image area D1 has a greater width b' (measured horizontally) than the unequally illuminated image area D2.

[0046] It should be noted that, due to the spacing of the receiving sensors 4, 5 from each other, one emitter 3 is sufficient to generate unequally illuminated image areas D1, D2 in the first and second image information B1, B2. However, it is advantageous if each receiving sensor 4, 5 is assigned an emitter 3, 3', and these emitters 3, 3' are each located near their assigned receiving sensor 4, 5, where "near" means, in particular, distances of less than 20 cm. Preferably, the receiving sensor 4, 5 is integrated into the emitter 3, 3', for example, as a camera integrated into the headlight.

[0047] The neural network 2 is trained to compare the image information B1, B2, to determine unevenly illuminated image areas D1, D2, and to estimate depth information by evaluating geometric differences that exist between the unevenly illuminated image areas D1, D2 in the first and second image information B1, B2.

[0048] As previously explained, the neural network 2 is configured to determine the distance of the vehicle 1 to areas of the captured scene that are visible through the first and second receiving sensors 4, 5, and thus visible in both image information B1, B2, by means of triangulation. For example, the image information B1, B2 is combined into a single image, and depth information is calculated for the pixels of the combined image that correspond to an area represented in both image information B1, B2.

[0049] The disadvantage here is that for areas of a background object, in which Figs. 2 and 3 the object O1, which is not visible in both image information B1, B2 due to parallax (in the Figs. 2 and 3 (the unevenly illuminated areas D1 and D2) mean that no depth information can be calculated.

[0050] However, by means of an estimation process of the neural network 2, it is possible to estimate depth information by comparing the geometric dimensions of the differently illuminated areas D1, D2 in the image data B1, B2. In particular, the horizontally measured width of the differently illuminated areas D1, D2 can be used to estimate the depth information. For example, the neural network 2 can derive from the comparison of the geometric dimensions of the differently illuminated areas D1, D2 the distance Δd between objects O1, O2, i.e., in the illustrated embodiment, how far object O2 is positioned in front of object O1. This yields estimated depth information, based on which a correction of the depth information calculated by triangulation is possible. This generates modified depth information that is used for the three-dimensional representation of the vehicle's surroundings.

[0051] For example, if triangulation at a specific pixel point calculates a distance Δd between objects O1 and O2 of 2m, but the depth estimate based on the unequally illuminated areas only yields a distance between objects O1 and O2 of 1.8m, the depth information obtained through triangulation can be modified based on the estimated depth information, so that the modified depth information indicates, for example, a distance Δd between objects O1 and O2 of 1.9m.

[0052] It is understood that, based on the comparison of the unequally illuminated areas D1, D2, it is also possible to determine which object O1, O2 these areas can be assigned to, and thus a depth estimation is also possible in areas that cannot be detected by both receiving sensors 4, 5.

[0053] For training neural network 2, training data in the form of image information pairs simulating a vehicle environment can be used. The image information in these pairs represents the same scene from different directions, as perceived by the spaced-apart sensors 4, 5, 6, 6' from their respective positions. The image information in these pairs also includes unevenly illuminated areas, generated by at least one, preferably two, emitters 3, 3'. Depth information for these unevenly illuminated areas is also present in the training data. This allows neural network 2 to be trained and its weighting factors to be adjusted so that the depth information estimated from the geometric information of the unevenly illuminated areas closely approximates the actual depth information.

[0054] Fig. 4 Figure 1 shows a vehicle 1 equipped with several sensor groups S1-S4 for capturing environmental information about the vehicle. For example, sensor group S1 is used to capture the environment in front of vehicle 1, sensor group S2 to capture the environment to the right of vehicle 1, sensor group S3 to capture the environment behind vehicle 1, and sensor group S4 to capture the environment to the left of vehicle 1.

[0055] The sensor groups S1 - S4 each have at least one emitter 6, 6', preferably at least two emitters 6, 6', and at least two detection sensors 7, 7'.

[0056] The sensors of the respective sensor groups S1-S4 each generate three-dimensional partial environmental information within their detection range, as described above. Preferably, the detection ranges of sensor groups S1-S4 overlap, and thus so do the partial environmental information they provide. This partial environmental information can advantageously be combined to form overall environmental information, which is, for example, a 360° 360° view or a partial 360° view (e.g., greater than 90° but less than 360°).

[0057] Since lateral or rear illumination with visible light, similar to headlights, is not possible, sensor groups S2 to S4 can emit electromagnetic radiation in the non-visible wavelength range, such as IR radiation, radar radiation, or laser radiation. Thus, emitters 6, 6' can be, for example, infrared light emitters, radar emitters, or LiDAR emitters. The receiving sensors 7, 7' are each adapted to the radiation of the corresponding emitters 6, 6', i.e., IR receivers, radar receivers, or LiDAR receivers.

[0058] Fig. 5 shows a diagram illustrating the steps of a procedure for determining depth information from image information using an artificial neural network 2 in a vehicle 1.

[0059] Initially, at least one emitter and at least one first and one second receiving sensor are provided (S10). The first and second receiving sensors are arranged at a distance from each other.

[0060] Subsequently, electromagnetic radiation is emitted by the emitter (S11). This can be, for example, light in the visible spectral range, in the infrared spectral range, laser light, or radar radiation.

[0061] Subsequently, reflected components of the electromagnetic radiation emitted by the emitter are received by the first and second receiving sensors, and first image information is generated by the first receiving sensor and second image information by the second receiving sensor based on the received reflected components (S12).

[0062] The first and second image data are then compared to determine at least one image area with unequal illumination in the first and second image data (S13). This unequal illumination arises from the parallax effect caused by the spaced arrangement of the receiving sensors.

[0063] Subsequently, the geometric information of at least one unevenly illuminated image area is evaluated and depth information is estimated by the artificial neural network based on the result of the evaluation of the geometric information of at least one unevenly illuminated image area (S14).

[0064] The invention has been described above using exemplary embodiments. It is understood that numerous modifications and adaptations are possible without thereby departing from the scope of protection defined by the patent claims. Reference symbol list

[0065] 1 Vehicle 2 Neural network 3 First emitter 3' Second emitter 4 First receiver 5 Second receiver 6, 6' Emitter 7, 7' Receiver 8 Computer unit b, b'width B1first image information B2second image information D1, D2unequally illuminated area Δddistance / path O1first object O2second object S1 - S4 Sensor groups

Claims

1. Method for determining depth information relating to image information by means of an artificial neural network (2) in a vehicle (1), comprising the following steps: - providing at least one emitter (3, 3') and at least a first and a second receiving sensor (4, 5), the first and second receiving sensors (4, 5) being spaced apart from one another (S10); - emitting electromagnetic radiation by the emitter (3, 3') (S11); - receiving reflected proportions of the electromagnetic radiation emitted by the emitter (3, 3') by the first and second receiving sensors (4, 5) and generating first image information (B1) by the first receiving sensor (4) and second image information (B2) by the second receiving sensor (5) on the basis of the received reflected proportions (S12); - comparing the first and second image information (B1, B2) for determining at least one image area (D1, D2) included in the first and second image information which comprises brightness differences and which occurs by the parallax due to the spaced-apart arrangement of the receiving sensors (4, 5) (S13); - evaluating geometric information of the at least one image area (D1, D2) which comprises brightness differences and estimating depth information by the artificial neural network (2) on the basis of the result of the evaluation of the geometric information of the at least one image area which comprises brightness differences (S14).

2. Method according to claim 1, characterized in that the image area (D1, D2) which comprises brightness differences occurs in the transition area between a first object (01) and a second object (O2) which have a different distance from the first and second receiving sensors (4, 5) and in that the estimated depth information is depth difference information containing information relating to the distance difference between the first and second objects (O1, O2) and the vehicle (1).

3. Method according to claim 1 or 2, characterized in that the emitter (3, 3') is at least one headlight emitting visible light in the wavelength range between 380nm and 800nm and the first and second receiving sensors (4, 5) are each a camera.

4. Method according to any one of the preceding claims, characterized in that the first and second receiving sensors (4, 5) form a stereo camera system.

5. Method according to any one of the preceding claims, characterized in that the at least two emitters (3, 3') are the front headlights of the vehicle (1), and in each case one receiving sensor (4, 5) is assigned to a front headlight (3, 3') in such a way that the straight line of sight between an object to be detected (O1, O2) and the front headlight runs substantially parallel to the straight line of sight between an object to be detected (O1, O2) and the receiving sensor (4, 5) assigned to this front headlight.

6. Method according to any one of the preceding claims, characterized in that the first and second receiving sensors (4, 5) are integrated in the front headlights of the vehicle (1).

7. Method according to any one of the preceding claims, characterized in that the artificial neural network (2) estimates the depth on the basis of the width (b), measured in the horizontal direction, of the image area (D1, D2) which comprises brightness differences.

8. Method according to any one of the preceding claims, characterized in that the artificial neural network (2) determines depth information in image areas detected by the first and second receiving sensors (4, 5) on the basis of a triangulation between pixels in the first and second image information (B1, B2) and the first and second receiving sensors (4, 5).

9. Method according to claim 8, characterized in that the neural network (2) compares depth information determined by triangulation and estimated depth information obtained by evaluating the geometric information of the at least one image area (D1, D2) which comprises brightness differences and generates modified depth information on the basis of the comparison.

10. Method according to claim 8 or 9, characterized in that the artificial neural network (2) modifies depth information obtained by triangulation on the basis of the evaluation of the geometric information of the at least one image area (D1, D2) which comprises brightness differences.

11. Method according to any one of the preceding claims, characterized in that at least one emitter (6, 6') emits IR radiation, radar signals or laser radiation.

12. Method according to claim 11, characterized in that at least part of the receiving sensors (7, 7') are infrared cameras, radar receivers or receivers for laser radiation.

13. Method according to any one of the preceding claims, characterized in that, for estimating depth information relating to image information representing areas laterally adjacent to the vehicle (1) and / or behind the vehicle (1), more than one emitter (3, 3', 6, 6') and more than two receiving sensors (4, 5, 7, 7') are used to determine image information, a plurality of sensor groups (S1, S2, S3, S4) being provided which each have at least one emitter and at least two receiving sensors, and the image information of the respective sensor groups (S1, S2, S3, S4) being combined to form overall image information.

14. Method according to claim 13, characterized in that the sensor groups (S1, S2, S3, S4) at least partially use electromagnetic radiation in different frequency bands.

15. System for determining depth information relating to image information in a vehicle (1), comprising a computer unit (8) which executes arithmetic operations of an artificial neural network (2), at least one emitter (3, 3') which is configured to emit electromagnetic radiation, and at least one first and one second receiving sensor (4, 5) which are arranged at a distance from one another, the first and second receiving sensors (4, 5) being configured to receive reflected proportions of the electromagnetic radiation emitted by the emitter (3, 3'), and the first receiving sensor (4) being configured to generate first image information (B1) and the second receiving sensor (5) being configured to generate second image information (B2) on the basis of the received reflected proportions, the artificial neural network (2) being configured to: - compare the first and second image information (B1, B2) for determining at least one image area (D1, D2) included in the first and second image information which comprises brightness differences, the image area (D1, D2) which comprises brightness differences occurring by the parallax due to the spaced-apart arrangement of the receiving sensors (4, 5); - evaluating the geometric information of the at least one image area (D1, D2) which comprises brightness differences and estimating depth information on the basis of the result of the evaluation of the geometric information of the at least one image area (D1, D2) which comprises brightness differences.