Method and apparatus for determining a hazard source on a lane

The image data is processed through monocular cameras and convolutional neural networks to determine the hazard sources in the lane, solving the problem of difficult to identify rare hazard sources in the prior art, and achieving fast and accurate hazard sources identification, which is suitable for autonomous driving and driver assistance systems.

CN113711231BActive Publication Date: 2025-07-22KERIDA EUROPE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080021323.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-15
Filing Date
2020-02-06
Publication Date
2025-07-22
Estimated Expiration
2040-02-06

AI Technical Summary

Technical Problem

Existing neural network-based methods are difficult to effectively classify and detect rare hazard sources on lanes, such as lost loads, exhaust pipes, personnel, potholes, ice or water membranes, and optical flow-based methods are computationally large and only effective at close range.

Method used

Using a monocular camera and a convolutional neural network, the first image area corresponds to the lane through image data processing, and the second image area is determined using the neural network intermediate layer output information to identify the dangerous source. Only a single image is required, which simplifies the implementation of the method.

Benefits of technology

It realizes the rapid and simple identification of hazardous sources in the lane in the detection area in front or rear of the vehicle, improves the accuracy and efficiency of identification, and is suitable for autonomous driving and driver assistance systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113711231B_ABST
    Figure CN113711231B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for determining a hazard source (14) on a lane (16) in a detection area (22) in front of or behind a vehicle (10) by means of a camera (18) of the vehicle (10), wherein an image of the detection area (22) is detected by means of the camera (18) to generate image data corresponding to the image, and a first image area (28) of the image is determined using a neural network (32), the first image area corresponding to the lane (16) in the detection area (22), and a second image area (30) of the image is determined using the neural network (32) based on the first image area (28), the second image area corresponding to the hazard source (14) on the lane (16) in the detection area (22).
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present invention relates to a method for determining a hazard source on a lane in a detection area in front of or behind a vehicle, in particular a road vehicle, by means of a camera of the vehicle. The present invention also relates to a device for determining a hazard source on a lane in a detection area. Background Art

[0002] In the context of autonomous driving and in driver assistance systems, the image data of images captured by a camera fixed to a vehicle should be processed in order to detect a hazard source on a lane in a detection area in front of or behind the vehicle. A possible hazard source is an object on the lane, such as a lost load / cargo or an exhaust pipe or a person. Other hazard sources on the lane are, for example, potholes, ice layers or water films, which may cause skidding.

[0003] Hitherto known neural network-based methods, such as monocular depth estimation (MonoDepth) or semantic segmentation, cannot classify such hazard sources due to the lack of training data and thus cannot detect such hazard sources either. In typical traffic situations, these hazard sources occur too rarely to generate suitable data items for training the neural network.

[0004] A method based on the optical flow of an image sequence of a detection area can only be used when the distance between the vehicle and the hazard source is very small. In addition, the computational complexity of this method is very high.

[0005] From the literature "Transfer Representation-Learning for Anomaly Detection" in the proceedings of the 33rd International Conference on Machine Learning by Andrews et al., two methods for detecting anomalies in data items are known. In both methods, the output of an intermediate layer of a neural network is used to detect anomalies by means of a support vector machine. In the first method, the neural network used is trained using data items different from the data items to be classified. In the second method, a neural network is used, which is trained using the non-anomalous part of the data items to be classified.

[0006] From the literature by Badrinarayanan et al., "SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation" (arXiv:1511.00561v3), a method for semantic classification of image points using a neural network is known. The neural network used includes: an encoder network that semantically segments an image into different image regions; and a decoder network that associates each image point with a corresponding region in the image region using the parameters of the encoder network.

[0007] From the literature by Chalapathy et al., "Anomaly Detection using One-Class Neural Networks" (arXiv:1802.06360v1), a method for detecting anomalies in complex data items is known. The neural network used includes an encoder network with fixed weights for classifying data items and a feed-forward network connected behind the encoder network, which is specifically designed to detect anomalies in the classified data items. Summary of the Invention

[0008] The object of the present invention is to provide a method and a device by which a hazard source on a lane in a detection area in front of or behind a vehicle can be simply determined by means of a camera of the vehicle.

[0009] The camera is preferably fixedly connected to the vehicle. The lane can in particular be a street. The vehicle is in particular a road vehicle. The hazard source is in particular an object on the lane, such as a lost load or an exhaust pipe or a person. The hazard source can also be a pothole, an ice plate or a water film, which may cause skidding.

[0010] In order to determine a hazard source on a lane in a detection area, in the method according to the present invention, a first image area is first used to limit the search for the hazard source to a determined sub-area of the detection area. Here, the neural network determines a display / representation of the image, in particular in the form of a matrix of feature vectors, using the image data. This display is used to classify the image areas of the image. Thus, the first image area is determined using the said display. The first image area corresponds here to the lane in the detection area. In addition, a second image area of the image is determined, for example, by means of the output information of the intermediate layer of the neural network in the first image area. This is carried out in particular using the same neural network as the one used to determine the first area. The second image area corresponds to the hazard source on the lane in the detection area. Alternatively, another neural network can be used to determine the second image area.

[0011] In the method according to the present invention, only a single image of the detection area is required to detect a hazard source. In particular, this image does not have to be a stereoscopic image. Thus, in order to carry out the method, only a monocular camera is required, i.e., a camera having only a single lens and capable of taking only non-stereoscopic images. Thus, the method according to the present invention can be carried out simply.

[0012] In a preferred embodiment, the second image area is determined by means of the first image area using the output information of the intermediate layer of the neural network. The intermediate layer is preferably a layer of the encoder network of the neural network. The hazard source is an anomaly, i.e., an undesired or atypical pattern, in the first image area which mainly includes only the lane in the detection area. It is known from the literature "Transfer Representation-Learning for Anomaly Detection" in the proceedings of the 33rd International Conference on Machine Learning by Andrews et al. that the output of the intermediate layer of a neural network can be used to detect anomalies in data items. This property of the neural network is used in the preferred embodiment to determine both the first area and to detect the hazard source in the image data corresponding to the first image area using the same neural network. Thus, the method can be carried out particularly simply since only a single neural network is required.

[0013] Preferably, a first feature vector is determined for each image point of the image using neural network with the aid of image data, and a first image region is determined with the aid of the first feature vector. The image points of the image are classified using the first feature vector. The classification is used to perform semantic segmentation of the image, i.e., to divide the image into regions that are coherent in terms of content, and thus the first image region is determined. In the present application, the feature vector is understood not only as the output information of the neural network itself, but also as the output information of the intermediate layer of the neural network.

[0014] In a preferred embodiment, a second feature vector is determined using neural network and a second image region is determined with the aid of the second feature vector. In particular, the output information of the intermediate layer of the neural network is used to determine the second feature vector. Alternatively or additionally, a second feature vector is determined for each predetermined sub-region of the first image region using neural network. The image points assigned to the first image region are classified using the second feature vector. This classification is used to determine the second image region and thus to determine the hazard source.

[0015] In another preferred embodiment, a second feature vector is determined for the first image region using neural network. In particular, the output information of the intermediate layer of the neural network is used to determine the second feature vector. The average value of the second feature vectors corresponding to / associated with the image points of the first image region is determined. For each second feature vector in the second feature vectors, the difference between the average value and the corresponding second feature vector is determined. If the value of the difference is higher than, lower than, or equal to a predetermined threshold, a predetermined sub-region of the image corresponding to the corresponding second feature vector is assigned as the second image region. The above difference is preferably weighted. In particular, the difference from the average value that is higher than, lower than, or equal to the predetermined threshold means an abnormal situation, i.e., a hazard source, and can be used to classify the image points of the first image region. Thus, the second image region corresponding to the hazard source can be determined with the aid of this classification.

[0016] In particular, the second image region is determined with the aid of the first image region using the output information of the encoder network of the neural network. The encoder network determines a representation of the first image region, in particular a representation in the form of a matrix of feature vectors. This representation is used for the classification of predetermined sub-regions of the image. The size of the predetermined sub-region depends in particular on the ratio between the resolution of the image and the resolution of the representation. Thus, the second image region can be determined particularly simply using this representation. Preferably, only a part of the encoder network, for example, the intermediate layer of the encoder network, is used to determine the second image region.

[0017] Preferably, image points of the image that are assigned to the first region are determined with the aid of image data in the case of a decoder network that uses a neural network. The decoder network determines, in particular in the case of a display determined by an encoder network, the image points of the image that are assigned to the first region. The decoder network in turn increases the resolution of the display determined by the encoder network, so that image points of the image can be assigned to the first image region.

[0018] Particularly advantageously, the neural network is a convolutional neural network. A convolutional neural network - which is also referred to as a Convolutional Neural Network (CNN) - generally comprises a plurality of convolutional layers and pooling layers. Image points of the image arranged in matrix form are generally used as input in the convolutional neural network. These matrices are each convolved with a convolutional matrix or filter kernel in the convolutional layer. In the pooling layer, the resolution of one of the outputs of the convolutional layer is reduced in each case in order to reduce the memory requirement and save computing time. In particular in image recognition, convolutional neural networks have a very low error rate.

[0019] Preferably, the neural network is trained with images of real, in particular typical, traffic situations taken by a camera of another vehicle or a camera of the vehicle itself.

[0020] In an advantageous refinement, an envelope enclosing the hazard source is determined with the aid of image data. The position of the envelope is preferably determined in three-dimensional coordinates, in particular in coordinates fixed to the vehicle or position-fixed coordinates. The envelope is in particular a rectangle enclosing the hazard source. The envelope can be used, for example, to drive around the hazard source or to initiate a braking process of the vehicle.

[0021] In an advantageous refinement, the hazard source is classified with the aid of image data in the case of an image recognition method, in particular another neural network. Different hazard sources can thereby be distinguished and thus the accuracy of the method increased.

[0022] The method described above can in particular be used by an autonomously driving vehicle which, for example, initiates an avoidance process, a braking process or an emergency braking process based on the detected hazard source. In addition, the aforementioned method can be used by a driver assistance system which, for example, notices the detected danger or automatically initiates a braking process based on the detected hazard source.

[0023] The invention also relates to a device for determining a hazard source on a lane in a detection region in front of or behind a vehicle, in particular a road vehicle. The device has the same advantages as the claimed method and can be improved in the same way. Description of the Drawings

[0024] Other features and advantages will be obtained from the following description that elaborates on multiple embodiments in conjunction with the accompanying drawings.

[0025] The figures show:

[0026] Figure 1 A vehicle having a device for determining a hazard source on a lane by means of a camera;

[0027] Figure 2 A flowchart showing a process for determining a hazard source on a lane by means of a camera;

[0028] Figure 3 A flowchart showing a detailed process for determining a first area on a lane;

[0029] Figure 4 A flowchart showing a detailed process for determining a second area on a lane;

[0030] Figure 5 A schematic illustration of an image of a detection area, in which the first area is highlighted;

[0031] Figure 6 Another schematic illustration of an image of a detection area, in which the second area is highlighted; and

[0032] Figure 7 A schematic diagram of a neural network. Detailed embodiments

[0033] Figure 1 A vehicle 10 having a device 12 for determining a hazard source 14 on a lane 16 by means of a camera 18 is shown. The vehicle 10 is designed as a passenger car in the illustrated embodiment.

[0034] Figure 1 The coordinate plane 20 of a position-fixed coordinate system is also shown. The first coordinate axis X is parallel to the lane 14 and extends in the driving direction of the vehicle 10. The second coordinate axis Y also extends parallel to the lane 14 and is perpendicular to the first coordinate axis X. The second coordinate axis Y extends transversely to the driving direction of the vehicle 10. The third coordinate axis Z points upward and is perpendicular to the plane defined by the first coordinate axis X and the second coordinate axis Y.

[0035] The device 12 includes a camera 18, which is fixedly connected to the vehicle 10. The camera 18 is oriented in the driving direction of the vehicle 10 or in the direction of the first coordinate axis X such that the camera can detect a detection area 22 on the lane 16 in front of the vehicle 10. In addition, the device 12 includes an image processing and evaluation unit 24, which is connected to the camera 18 via a cable 26. Alternatively, the evaluation unit 24 can also be located outside the vehicle 10, for example, in an external server. The image processing and evaluation unit 24 is designed to determine, using neural network 32 ( Figure 7 ), a first image area of the image based on image data corresponding to the image of the detection area 22 captured by the camera 18, where the first image area corresponds to the lane 16 in the detection area 22, and to determine, using neural network 32, a second image area of the image based on image data corresponding to the first image area, where the second image area corresponds to a hazard source 14 on the lane 16 in the detection area 22. The neural network 32 includes an encoder network 34 and a decoder network 36, as described in more detail below in conjunction with Figure 7 .

[0036] Figure 2 The flowchart shows an embodiment of a method for determining a hazard source 14 on the lane 16 in the detection area 22 of the camera 18.

[0037] In step S10, the process starts. Subsequently, in step S12, an image of the detection area 22 is detected by means of the camera 18. Subsequently, in step S14, image data associated with the image is generated. Subsequently, in step S16, a first image area 28 of the image is determined using neural network 32 based on the image data. For this purpose, all layers of the neural network 32 are used. The first image area 28 corresponds to the lane 16 in the detection area 22 and is shown in Figure 5 . The details of step S16 are described in more detail below in conjunction with Figure 3 . In step S18, a second image area 30 of the image is determined using neural network 32 based on the image data corresponding to the first image area 28. For this purpose, the output information of the intermediate layers of the neural network 32, preferably the output information of the encoder network 34 of the neural network 32, is used. The second image area 30 corresponds to the hazard source 14 on the lane 16 in the detection area 22 and is shown in Figure 6 . The other details of step S18 are described in more detail below in conjunction with Figure 4 . Subsequently, the process ends in step S20.

[0038] Figure 3 Shows in accordance with Figure 2Flowchart of the detailed process for determining the first region on lane 16 when using neural network 32 in step S16.

[0039] In step S100, the process is started. Subsequently, in step S102, the image data generated in step S14 is processed using neural network 32 to determine the first feature vector. For this purpose, all layers of neural network 32 are used. The first feature vector is arranged in matrix form and is a display of the information contained in the image of detection region 22 that is particularly suitable for further processing. In step S104, semantic segmentation of the image is performed using the first feature vector. Here, image regions that are coherent in terms of content are determined. Subsequently, in step S106, the first image region 28 corresponding to lane 16 in detection region 22 is determined based on the image semantic segmentation completed in step 104. Finally, in step S108, the process ends.

[0040] Figure 4 Shows according to Figure 2 Flowchart of the detailed process of step S18.

[0041] In step S200, the process is started. Subsequently, in step S202, second feature vectors are respectively determined for predetermined sub-regions of the first image region 28 using the output information of the intermediate layer of neural network 32. This intermediate layer of neural network 32 belongs to the encoder network 34 of neural network 32. Thus, the encoder network 34 determines a matrix-form display of the second feature vectors of the first image region 28 using the image data, which display is particularly suitable for further processing. Generally, this display has a lower resolution than the image itself, that is, there is not exactly a second feature vector for each image point of the first image region 28. The size of the predetermined sub-regions depends in particular on the ratio between the resolution of the image and the resolution of this display, that is, the matrix of the second feature vectors. For example, each of the predetermined sub-regions can correspond to one second feature vector of the display, and a plurality of image points are in turn assigned to this sub-region, and the plurality of image points are arranged, for example, in a rectangular sub-region of the image.

[0042] In step S204, the average value of the first feature vectors of all the image points assigned to the first image region is formed. Subsequently, in step S206, the difference between the average value and each second feature vector among the second feature vectors is determined. When the value of this difference is higher than, lower than, or equal to a predetermined threshold, the predetermined sub-region corresponding to the respective second feature vector is assigned as the second image region. After determining the difference between the average value and the second feature vectors for all the second feature vectors, the second image region 30 is thereby determined. Subsequently, in step S208, the process ends.

[0043] Figure 5 Schematic diagram showing the image of the detection area 22. In the illustration according to Figure 5 , the first image area 28 of the lane 16 is highlighted. In addition, in Figure 5 , the first coordinate axis X and the second coordinate axis Y of the image points of the image are shown. In the exemplary illustration according to Figure 5 , the first area reaches up to the value 200 of the second coordinate axis Y. Trees are arranged on the side of the lane 16. In addition, other vehicles are arranged on the side of the lane 16 and partially on the lane. If a vehicle is arranged on the lane 16 or partially on the lane 16, the lane 16 area of the lane 16 covered by the vehicle in the image is excluded from the first image area 28. The hazard source 14 is located on the lane 16, and the hazard source is exemplarily shown as an injured person in Figure 5 and Figure 6 .

[0044] Figure 6 Another schematic diagram showing the image of the detection area 22. The illustration according to Figure 6 is basically corresponding to the illustration according to Figure 5 , however, in Figure 6 , the second area 30 of the lane 16 is highlighted.

[0045] Figure 7 Schematic diagram showing the neural network 32, which is used to determine the first image area 28 and the second image area 30 in step S102 according to Figure 3 and step S202 according to Figure 4 . The neural network 32 includes an encoder network 34 and a decoder network 36, and the encoder network and the decoder network are each composed of multiple layers.

[0046] The encoder network 34 includes multiple convolutional layers 36 and pooling layers 40. For the sake of clarity, only one layer of these layers is shown respectively in Figure 7 . The existence of the convolutional layer makes the neural network 32 a so-called convolutional neural network. The encoder network 34 takes image data as input and determines the display of the image in the form of a feature vector matrix from these image data.

[0047] The decoder network 36 also includes multiple layers, and these layers are not shown separately for the sake of clarity in Figure 7 . The output of the network 32 is the display of the image data in the form of a matrix of feature vectors, which allows semantic segmentation of these image data. This display particularly allows determination of the first image area 28.

[0048] According toFigures 1 to 7 The method according to the invention and the device according to the invention are described by way of example with reference to embodiments. In particular, in the illustrated embodiment, the detection area 22 is located in front of the vehicle 10. It is obvious that the illustrated method embodiments can also be correspondingly used for areas behind the vehicle 10.

[0049] List of reference numerals:

[0050] 10 vehicle

[0051] 12 device

[0052] 14 hazard source

[0053] 16 lane

[0054] 18 camera

[0055] 20 coordinate plane

[0056] 22 detection area

[0057] 24 image processing and evaluation unit

[0058] 26 cable

[0059] 28, 30 image area

[0060] 32 neural network

[0061] 34 encoder network

[0062] 36 decoder network

[0063] 38 convolutional layer

[0064] 40 pooling layer

Claims

1. A method for determining a hazard source (14) on a lane (16) in a detection area (22) in front of or behind a vehicle (10) by means of a camera (18) of the vehicle (10), wherein, The image of the detection area (22) is detected by means of a camera (18). - Image data corresponding to the image is generated. - A first image area (28) of the image is determined by means of the image data using a neural network (32), and this first image area corresponds to a lane (16) in the detection area (22). - A second image area (30) of the image is determined by means of the first image area (28) using a neural network (32), and this second image area corresponds to a hazard source (14) on the lane (16) in the detection area (22). The second image area (30) is determined by means of the first image area (28) using the output information of the intermediate layer of the neural network (32). By means of the image data, a first feature vector is determined for each image point of the image using a neural network (32), and the first image area (28) is determined by means of the first feature vector. A second feature vector is determined for the first image area (28) using a neural network (32), wherein the average value of the second feature vectors corresponding to the image points of the first image area is determined. For each second feature vector among the second feature vectors, the difference between the average value of all the second feature vectors belonging to the first image area (28) and the corresponding second feature vector is determined, and when the value of the difference is higher than, lower than, or equal to a predetermined threshold, a predetermined sub-area of the image corresponding to the corresponding second feature vector is assigned as the second image area (30).

2. The method according to claim 1, characterized in that The neural network (32) at least includes an encoder network and a decoder network.

3. The method according to claim 1 or 2, characterized in that, The second image area (30) is determined by means of the first image area (28) using the output information of the encoder network (34) of the neural network (32).

4. The method according to claim 1 or 2, characterized in that, The image points of the image assigned as the first image area (28) are determined by means of the image data using a neural network (32).

5. The method according to claim 1 or 2, characterized in that, The neural network (32) is a convolutional neural network.

6. The method according to claim 1 or 2, characterized in that, The neural network (32) is trained using images of real traffic conditions captured by the camera (18) of the vehicle (10) or the camera of another vehicle.

7. The method according to claim 1 or 2, characterized in that, The neural network (32) is trained using images of real and typical traffic conditions captured by the camera (18) of the vehicle (10) or the camera of another vehicle.

8. The method according to claim 1 or 2, characterized in that, An envelope enclosing the hazard source (14) is determined by means of the image data.

9. The method according to claim 1 or 2, characterized in that, The hazard source (14) is classified by means of the image data using an image recognition method.

10. The method according to claim 1 or 2, characterized in that, The hazard source (14) is classified by means of the image data using an image recognition method and using another neural network.

11. The method according to claim 1, wherein The vehicle is a road vehicle.

12. A device for determining a hazard source (14) on a lane (16) in a detection area (22) in front of or behind a vehicle (10) by means of a camera (18) of the vehicle (10), the device having a camera (18) and an image processing and evaluation unit, the camera being designed to detect an image of the detection area (22), the image processing and evaluation unit being designed to generate image data corresponding to the image, determine a first image area (28) of the image by using a neural network (32), the first image area corresponding to the lane (16) in the detection area (22), and determine a second image area of the image by using the first image area (28) and a neural network (32), the second image area corresponding to the hazard source (14) on the lane (16) in the detection area (22). Determine the second image area (30) by using the output information of the intermediate layer of the neural network (32) with the first image area (28). Respectively determine a first feature vector for each image point of the image by using the neural network (32) with the image data, and determine the first image area (28) by using the first feature vector. Determine a second feature vector for a first image region (28) in the case of using a neural network (32), wherein, Determine the average value of the second feature vectors corresponding to the image points of the first image area. For each second feature vector among the second feature vectors, determine the difference between the average value of all the second feature vectors belonging to the first image area (28) and the corresponding second feature vector, and when the value of the difference is higher than, lower than or equal to a predetermined threshold, assign a predetermined sub-area of the image corresponding to the corresponding second feature vector to the second image area (30).

Citation Information

Patent Citations

  • Risk prediction method

    CN107180220A

  • Method and Apparatus for Determining a Road Condition

    US20150371095A1