Method for determining obstacles in an image captured by a camera of the surroundings of a vehicle

The method improves obstacle detection in vehicle surroundings by categorizing image areas and minimizing an error function to exclude 'unknown' regions, addressing high-cost training data needs and enhancing detection precision.

DE102024211522A1Pending Publication Date: 2026-06-03ROBERT BOSCH GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Conventional deep neural networks for obstacle detection in vehicle surroundings require large amounts of reliable training data, leading to high costs and effort, and struggle with unstructured environments where objects like vegetation cannot be clearly classified as passable or impassable.

Method used

A method using a modified minimization procedure in a deep learning system to determine the lower boundary of obstacles by categorizing image areas as 'passable', 'impassable', or 'unknown', and minimizing an error function that excludes 'unknown' areas, allowing for improved accuracy in obstacle detection.

Benefits of technology

Enhances the ability to accurately identify traversable and obstructive areas in vehicle surroundings, reducing computational effort and training data requirements while improving detection precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for determining obstacles in an image (B) of the environment (U) of an ego-vehicle (2) taken by a camera, the procedure is carried out by a neural network and includes the following measures: a) Dividing the image into different image areas (B1, B2, B3), where each image area is assigned to one of the three categories “passable”, “impassable” or “unknown”, b) Predicting a lower boundary of an impassable obstacle located lowest in the image by minimizing an error function (L1), where: - if the predicted lower limit (YPRB) is located above an actual lower limit (YTRB) of the lowest image area (B2) of the category "impractible" present in the image (B), this is taken into account in the error function (L1) according to measure c1), - if the predicted lower limit (YPRB) is located below an actual upper limit (YFT) of the lowest image area (B1) of category (B1) ‘passable’ in the image (B), this is taken into account in the error function (L1) according to measure c2), - if the predicted lower limit (YPRB) is located in an image area (B3) of the category “unknown”, this is not taken into account in the error function (L1) according to measure c3).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method for determining obstacles in an image taken by a camera of the surroundings of a vehicle.

[0002] Detecting obstacles or objects closest to a vehicle is essential for computer-assisted driving. Driver assistance functions available in vehicles today, such as automatic braking, parking, and adaptive cruise control, require reliable detection of obstacles in the surroundings to execute such maneuvers safely.

[0003] In recent years, deep neural networks (DNNs) have demonstrated remarkable generalization capabilities in tasks such as object detection in images depicting a vehicle's surroundings and in semantic segmentation. However, all these models require precise knowledge of the different types of objects, which is often not possible in the environment of an autonomous vehicle, as the environment is frequently unstructured and changes rapidly over time.

[0004] Encouraged by the success of deep neural networks in some of the most complex problems in the field of computer vision, it has now become possible to apply such networks to the task of generic object recognition.

[0005] In the professional world, a representation of the ego-vehicle's environment in the form of so-called "stixels" has become established for segmenting the obstacle nearest to the ego-vehicle within the camera's image space. In this process, the image of the ego-vehicle's environment, generated by the camera, is discretized into a fixed number of such stixels.

[0006] To predict the upper and lower edges of the stixels, which are important for obstacle detection, as accurately as possible for a given image, various models for deep neural networks have been proposed.

[0007] However, one disadvantage of such conventional methods is that the models used to train the neural network require an extremely large amount of reliable training data, which involves a great deal of effort and therefore high costs.

[0008] It is therefore an objective of the present invention to provide an improved embodiment of a method for obstacle detection, as described above, in which the aforementioned disadvantage is at least partially eliminated.

[0009] The objective is achieved through the subject matter of the independent claims. Preferred embodiments are the subject matter of the dependent claims.

[0010] The basic idea of ​​the invention is therefore to take into account, when determining a lower boundary of obstacles in an image using a minimization method, that an object detected in the image using object recognition, or the image area assigned to this object, cannot always be clearly assigned to the category "passable" or "impassable". This applies, for example, to vegetation, especially plants, near a road, which do not actually pose an obstacle for the ego vehicle, but may also be so large that they could damage the vehicle in a collision.

[0011] In the method presented here, the executing neural network is able to learn geometric structures that cannot be clearly assigned to either the "passable" or "impassable" categories using a modified minimization procedure. This allows for the determination of a lower boundary of obstacles in an image with improved accuracy. Consequently, it is possible to determine with greater accuracy which areas of the image can actually be traversed by the ego vehicle.

[0012] The method according to the invention is used to determine obstacles in an image of the surroundings of a motor vehicle, captured by a camera. The method is carried out by a deep learning system comprising at least one, preferably deep, neural network.

[0013] According to measure a) of the procedure, the image is divided into different image areas. In this process, each image area is assigned to one of three categories: "passable," "impassable," or "unknown." An image area assigned to the "passable" category is classified as flat or level, specifically as a road on which the ego-vehicle can therefore drive. An image area assigned to the "impassable" category is therefore considered an obstacle. This could be, for example, a lane marking, another vehicle, a traffic sign, or something similar. All image areas that cannot be assigned to either the "passable" or "impassable" categories are classified as "unknown," since it is not known whether the ego-vehicle can drive in these image areas or not.

[0014] The division of the image into image areas according to measure a) can be carried out by assigning the image areas to predetermined object classes, for which it is known whether the ego-vehicle can drive on them or not. For example, it is known that the ego-vehicle cannot drive on the object class "other motor vehicle", i.e., it is an obstacle. For this purpose, object recognition in the image must be performed for evaluation using the deep learning system or neural network.

[0015] In measure b), the prediction of a lower boundary of the image area categorized as "impervable," i.e., an obstacle located furthest down in the image and thus closest to the ego-vehicle or the camera. It is understood that not only a single image area categorized as "impervable," but also two or more such image areas can form this lower boundary.

[0016] Predicting or estimating this lower bound is performed using a minimization procedure that minimizes an error function. When calculating the error function, only contributions from image regions assigned to either the passable or impassable categories are considered. Image regions assigned to the unknown category, on the other hand, are not included in the error function.

[0017] The following three criteria are used for the specific determination and calculation of the error function L1: 1. If the predicted lower boundary YPRB is above an actual lower boundary YTRB of the lowest image area of ​​the category "not passable" that is present in the image, this is taken into account in the error function L1 according to measure c1) - which is explained in more detail below. The terms "above" and "below" or "lowest", etc., refer to an arrangement of the image showing the environment of the ego-vehicle in a drawing plane, in such a way that objects located lower in the image are positioned at a shorter distance from the ego-vehicle or the camera than image areas located higher in the image. 2. If the predicted lower limit YPRP is below an actual upper limit YFT of the lowest traversable category of the image, this is taken into account in the error function L1 according to measure c2) - which will be explained in more detail below. 3. If the predicted lower boundary YPRB lies within an image area categorized as "unknown," this is not taken into account in the error function L1 according to measure c3). Image areas whose traversability is unknown are therefore not considered when minimizing errors. Measure c3 thus ensures that potential obstacles are also considered when determining the lower boundary.

[0018] As already explained, the method according to the invention is carried out by a deep learning system with at least one, preferably deep, neural network. Thus, the deep learning system or the neural network can determine with increasing accuracy, through appropriate learning, where the lower boundary of obstacles lies in an image and where an actual upper boundary of the navigable area lies in the image.

[0019] One variant has proven advantageous in which the error function L1 consists of a lower part L1_U and an upper part L1_O. It is particularly advantageous because it involves very little computational effort if the error function L1 is formed by the sum of the lower part L1_U and the upper part L1_O.

[0020] According to an advantageous embodiment of the method according to the invention, the deep learning system or neural network can learn how to determine the lower bound with particular precision by setting the lower part L1_U of the error function L1 as the magnitude of the difference between the predicted lower bound YPRB and the actual lower bound YTRP in the aforementioned step c1) of the method. Accordingly, in this embodiment, in the aforementioned step c2) of the method, the lower part L1_U of the error function L1 is defined as the magnitude of the difference between the predicted lower bound YPRB and the actual upper bound YFT. In this embodiment, it is essential that a zero value L1_U = 0 is defined for the lower part L1_U of the error function L1 in the aforementioned step c3).

[0021] In a further preferred embodiment, according to measure c4) with respect to one or more lowest image regions of the category “not passable” that are present in the image, the upper part L1_O of the error function L1 can be defined as the amount of the difference between the predicted upper limit YTRT and the actual upper limit YTPT.

[0022] Particularly in the case of an image consisting only of "passable" and "impassable" image areas, the error function L1 with respect to the predicted upper and lower limits can be minimized in a known manner, in the sense of "complete monitoring," compared to the actual upper and lower limits. This enables the deep learning system to learn obstacle detection.

[0023] The invention also relates to a deep learning system comprising at least one neural network, which is configured / programmed to execute the aforementioned method according to the invention. The advantages of the method according to the invention, as explained above, are therefore transferred to the deep learning system according to the invention.

[0024] The invention also relates to a motor vehicle with at least one camera for monitoring an area around the vehicle. Furthermore, the motor vehicle includes a control unit that interacts with the camera and is configured / programmed to execute the aforementioned method according to the invention. The advantages of the method according to the invention, as explained above, are therefore transferred to the motor vehicle according to the invention.

[0025] Furthermore, the invention relates to a computer program product configured to execute the method according to the invention, in particular by means of the deep learning system. The computer program product contains instructions which, when executed by the vehicle's control unit and / or by the deep learning system, cause the latter to execute the method. The advantages of the method according to the invention, as explained above, are therefore transferred to the computer program product according to the invention.

[0026] The computer program product is preferably stored / stored on a memory that includes at least one non-volatile memory.

[0027] The invention also includes a computer-readable data carrier for executing the method. The data carrier contains instructions which, when executed, cause the vehicle's control unit and / or the deep learning system to execute the method according to the invention, as explained above. The advantages of the method according to the invention, as explained above, are therefore transferred to the data carrier according to the invention.

[0028] Further important features and advantages of the invention can be seen from the partial claims, the drawings and the associated description of figures based on the drawings.

[0029] It is understood that the features mentioned above, and those explained below, can be used not only in the respective given combination, but also in other combinations or on their own, without deviating from the scope of protection of the present invention.

[0030] Preferred embodiments of the invention are shown in the drawings and are explained in more detail in the following description, wherein the same reference numerals refer to the same or similar or functionally identical components.

[0031] The following shows, in any case schematically: Fig. 1 an image of a motor vehicle's route according to the invention, Fig. 2 a flowchart illustrating the method according to the invention, Fig. 3a a representation accordingly Fig. 1, in which obstacles present in the image were detected using a conventional, i.e., non-inventive, method, Fig. 3b the in Fig. 3a shown pre-air area, in which, in contrast to Fig. 3a, Obstacles present in the image were determined using the method according to the invention.

[0032] Fig. Figure 1 shows an image B of a forecourt V of a motor vehicle 1 according to the invention, which is in Fig. 1 is only roughly shown in the form of a rectangular frame - this is also referred to below as Ego Vehicle 2 - when it is driving in a lane F.

[0033] Image B was taken by a camera 3 located in the vehicle, which is used to monitor an environment U or a foreground V of the motor vehicle 1 or the ego vehicle 2.

[0034] In Fig. For a better description of image B, without loss of generality, a coordinate system K is defined with origin USP and X-direction X extending horizontally to the right away from origin USP, and with a Y-direction extending vertically downwards from origin USP. The terms "above" and "below" or "lowest" refer to an arrangement of image B, which represents the environment U of the ego-vehicle 2, in a drawing plane, such that objects located lower in image B are positioned at a shorter distance from ego-vehicle 2 or camera 3 than image areas located higher in image B.

[0035] How Fig. Figure 1 shows that the motor vehicle 1 or the ego-vehicle 2 comprises a control unit 4 (also not shown) that interacts with the camera 3. The camera 3 can be connected to the control unit 4 for data transfer (see arrow P) so that the images B captured by the camera are transferred to the control unit 4 for processing and, in particular, evaluation during the execution of the method according to the invention.

[0036] A deep learning system 5 is provided in the control unit 4 – for example, in the form of a computer program product that can be executed by the control unit 4 and that complies with the invention – which includes a deep neural network 6, which is not described in detail, wherein the deep learning system 5 or the deep neural network 6 is configured and programmed to execute the aforementioned method according to the invention. The computer program product thus contains instructions which, when the computer program product is executed by the control unit 4 of the motor vehicle, cause the deep learning system 5 or the neural network 6 to execute the correct procedure.

[0037] As from Fig. As can be seen in 1, image B is divided into different image areas B1 to B3, which in turn are assigned to three different categories.

[0038] Image area B1 corresponds to a road F on which the Ego vehicle 2 can drive, and which is therefore assigned to the category "passable".

[0039] The Stixels STX, which in Fig. The elements visible in image B are part of an image area B2, which is assigned the category "impervable". A Stixel STX is defined as a fixed-width rectangle that covers the surface of an obstacle 7 in image B. The Stixels STX thus mark and delineate obstacles 7 in image B with which the Ego vehicle 2 should not collide. For example, such an obstacle 7 is another vehicle 8 traveling in lane F. Furthermore, a third image area B3 can also be seen in image B, formed by vegetation 9 at an edge 10 of lane F. The vegetation 9 cannot be clearly assigned to the category of traversable or impassable. Consequently, the third image area B3 with the vegetation 9 is assigned the category "unknown".

[0040] The following describes, by way of example, the method according to the invention, which is carried out by the deep learning system 5, based on the following: Fig. 1 and the flowchart of Fig. 2 explained.

[0041] As already mentioned above, based on Fig. As explained in section 1, according to measure a) of the method according to the invention, image B is divided into different image areas B1, B2, B3. In the exemplary scenario, as already explained above, each of the three image areas B1, B2, B3 is assigned to the categories "passable", "impedable" or "unknown".

[0042] Image area B1, which has already been explained and assigned to the category "passable," is classified as a road F on which the Ego-Vehicle 2 can drive. Image area B2, which has already been explained and which contains, among other things, another motor vehicle 9, i.e., an obstacle 7, is assigned to the category "impassable." Image area B3, which has already been explained and contains vegetation 9, is assigned to the category "unknown" because it cannot be clearly determined whether it is passable for the Ego-Vehicle 2 or not.

[0043] The division of image B into image areas B1, B2, and B3 can be achieved by assigning these areas to predefined object classes. For this purpose, the deep learning system 5 or the neural network 6 can be used to perform object recognition in image B, allowing the recognized objects to be assigned to predefined object classes, where it is known whether they belong to the category "passable," "impassable," or "unknown." Such an object class could, for example, be the roadway (passable), another motor vehicle (impassable), or the aforementioned vegetation (unknown).

[0044] In a further step b), the deep learning system 5 or the neural network 6 predicts a lower boundary YPRB of the image area B2, i.e., the "impermissible" category located furthest down in image B with respect to the Y-direction and thus having the smallest distance to the ego vehicle 2 or the camera 4. The prediction or estimation of this lower boundary YPRB is performed using a minimization procedure in which an error function L1 is minimized.

[0045] The error function L1 to be calculated can consist of a lower part L1_U and an upper part L1_O. In the example scenario, the error function L1 is formed by the sum of the lower part L1_U and the upper part L1_O, i.e., L1 = L1_U + L1_O.

[0046] When error function L1 is defined, only the two image areas B1 and B2 whose objects are assigned to the category "passable" or the category "impassable" are considered. In contrast, image area B3, which is assigned to the category "unknown," is not considered by error function L1.

[0047] In particular, the following three criteria 1), 2) and 3) are used in the example to calculate the lower part L1_U of the error function L1: 1. If the predicted lower limit YPRB is above the actual lower limit YTRB of the lowest image area of ​​category B2 "not passable" in the image, this is taken into account in the error function L1 according to measure c1) (see flowchart in Fig. 2) For this purpose, in measure c1), the lower part L1_U of the error function L1 can be defined as the absolute value of the difference between the predicted lower limit YPRB of the image area B2 and the actual lower limit YTRB of the image area B2. Thus, L1_U=|YPRB−YTRB|, if YPRB <YTRB. 2. If the predicted lower limit YPRB is below an actual upper limit YTF of the lowest image area B1 of the category "passable" present in image B, this is taken into account in the error function L1 according to measure c2) (see flowchart in Fig. 2) For this purpose, in measure c2), the lower part L1_U of the error function L1 can be defined as the absolute value of the difference between the predicted lower limit YPRB of the image area B2 and the actual upper limit YTF of the image area B1. Thus, L1_U=|YPRB−YTF|, if YPRB>YTF. 3. If the predicted lower boundary YPRB lies within an image region of the category "unknown," this is ignored in the error function L1 according to measure c3). The image region B3, whose traversability is therefore unknown, is ignored when the error function L1 is determined and minimized. For this purpose, a zero value L1_U = 0 can be set for the lower part L1_U of the error function L1 in measure c3). Thus, the following applies: L1_U=0 if YPRB>YTRB.

[0048] In a further measure c4), with respect to image B, the lowest image area B2 of the category “not passable” can be defined as the upper part L1_O of the error function L1 as the amount of the difference between the predicted upper limit YPRT and the actual upper limit YTRT of image area B2.

[0049] In summary, the following results from the above: L1=L1_U+L1_O. L1_U=|YPRB−YTRB|, if YPRB <YTRB. L1_U=|YPRB−YTF|, if YPRB>YTF L1_U=0, mn if YPRB>YTRB L1_O=|YPRT−YTRT|

[0050] To illustrate the improvements achieved by the method according to the invention when the lower boundary YPRB of the image area B2 (unpassable) is predicted, the two show Fig. 3a and Fig. 3b, in any case in an analogous manner to Fig. 1, an image B of a surrounding area s U or of a foreground V. In Fig. Figure 3a shows the course of the estimated and sought variable YPRB when a conventional procedure is used, and in Fig. Figure 3b shows the behavior of the same variable YPRB when the procedure presented here is used.

[0051] It can be clearly seen that in Fig. 3b the curve YFT, i.e., the actual upper limit of the passable area, i.e., the roadway F, the actual conditions better than in Fig. 3a reflects this.

[0052] This is particularly evident from the fact that in Fig. 3a the right area B3 with vegetation V the area B2 (not passable) in Fig. 3b is assigned. In the left area of ​​the image, at least part of the vegetation V in image B is shown. Fig. 3b was recognized as impassable and therefore classified as a second image area B2.

[0053] Using the method according to the invention, impassable areas B2 can be better identified as obstacles.

Claims

Method for determining obstacles in an image (B) of the environment (U) of an ego-vehicle (2) captured by a camera (3), wherein the method is performed by a neural network (6) and comprises the following measures: a) dividing the image into different image regions (B1, B2, B3), each image region being assigned to one of the three categories "passable", "impassable", or "unknown", b) predicting a lower boundary of an impassable obstacle located at the lowest point in the image by minimizing an error function (L1), wherein: - if the predicted lower boundary (YPRB) is located above an actual lower boundary (YTRB) of the lowest image region (B2) of the category "impassable" present in the image (B), this is taken into account in the error function (L1) according to measure c1).- If the predicted lower limit (YPRB) is located below an actual upper limit (YFT) of the lowest image area (B1) of category (B1) "passable" in image (B), this is taken into account in the error function (L1) according to measure c2); - If the predicted lower limit (YPRB) is located in an image area (B3) of category "unknown", this is not taken into account in the error function (L1) according to measure c3). The method according to claim 1, characterized in that the method is carried out by a deep learning system (5) comprising at least one neural network (6). Method according to claim 1 or 2, characterized in that the image (B) is subdivided into the image areas (B1, B2, B3) according to measure a) by assigning these image areas (B1, B2, B3) to predetermined object classes. Method according to one of claims 1 to 3, characterized in that the fault function (L1) consists of a lower part (L1_U) and an upper part (L1_O). Method according to claim 4, characterized in that the error function (L1) is formed by the sum of the lower part (L1_U) and the upper part (L1_O). A method according to one of the preceding claims, characterized in that - in measure c1) the lower part (L1_U) of the error function (L1) is set as the magnitude of the difference between the predicted lower limit (YPRB) and the actual lower limit (YFT); and that - in measure c2) the lower part (L1_U) of the error function (L1) is set as the magnitude of the difference between the predicted lower limit (YPRB) and the actual upper limit (YFT); and that - in measure c3) a zero value (L1_U=0) is set for the lower part (L1_U) of the error function (L1). Method according to one of the preceding claims, characterized in that in a measure c4) for the lowest image area (B2) of the category “not passable” which is present in the image (B), the upper part (L1_O) of the error function (L1) is defined as the amount of the difference between the predicted upper limit (YPRT) and the actual upper limit (YTRT). Method according to one of the preceding claims, characterized in that in the case of an image (B) consisting only of image areas (B1, B2) of the categories “passable” and “unpassable”, the error function (L1) relating to the predicted upper and lower limits (YPRB, YPRT) is minimized compared to the actual upper and lower limits (YTRT, YFT). Deep learning system (5) comprising at least one, in particular deep, neural network (6) which is set up / programmed to perform the method according to one of the preceding claims. Motor vehicle (10),- comprising at least one camera (1), in particular for monitoring the environment of the motor vehicle (10),- with a control unit (13) which interacts with the camera and is set up / programmed to perform the method according to one of claims 1 to 8. Computer program product containing instructions which, when the computer program product is executed by a computer system and / or by the deep learning system according to claim 8, cause it to execute the method according to any one of claims 1 to 8. Data carrier containing instructions which, when executed by the motor vehicle (1) and / or by the deep learning system according to claim 9, cause it to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Determining drivable free-space for autonomous vehicles

    US20190286153A1