Method, device, computer program and machine-readable storage medium for detecting an object

By using regional candidate networks and quality models in object detection, selecting and fusing object assumptions, the problems of poor processing of information loss and redundant object assumptions in the prior art are solved, and more efficient and accurate object detection is achieved.

CN112287961BActive Publication Date: 2025-07-01ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010709336.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-22
Filing Date
2020-07-22
Publication Date
2025-07-01
Estimated Expiration
2040-07-22

AI Technical Summary

Technical Problem

The prior art has the problem of information loss in object detection, especially in the process of abandoning the assumption of redundant object, and fails to effectively utilize the output structure and model knowledge of RPN.

Method used

The sensor signal is processed through the regional candidate network, object hypothesis is generated, and the best object hypothesis is selected using the quality model, identifying and fusing redundant object hypothesis to improve the accuracy of object detection.

Benefits of technology

This method reduces the number of object assumptions by considering the objective function and model knowledge of RPN, while retaining effective processing of redundant object assumptions, improving the quality and accuracy of object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112287961B_ABST
    Figure CN112287961B_ABST
Patent Text Reader

Abstract

A method for detecting an object in the surrounding environment of a vehicle based on sensor signals, wherein the sensor signals represent the surrounding environment of the vehicle, the method comprising the steps of: processing the sensor signals by means of a region candidate network to obtain at least one object hypothesis for each anchor point, wherein the object hypothesis includes an object probability and a bounding box; selecting the best object hypothesis by means of a quality model, wherein the quality model depends on the anchor point and the bounding box of the object hypothesis; identifying object hypotheses redundant with respect to the selected object hypothesis, wherein the redundant object hypotheses are identified according to the anchor points of the redundant object hypotheses by means of an objective function assigned to the region candidate network; and fusing the selected object hypothesis with the identified redundant object hypotheses for object detection. The invention also relates to a correspondingly configured device, a correspondingly configured computer program, and a machine-readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting an object, a correspondingly configured device, a correspondingly configured computer program, and a machine-readable storage medium. Background Art

[0002] Convolutional Neural Networks ("CNN") are a form of artificial neural network. Typically, a CNN is constructed as a stack of layers consisting of one or more successive so-called convolutional layers ("Convolutional Layer") followed by pooling layers ("Pooling Layer"). The sequence consisting of convolutional layers and pooling layers can be repeated.

[0003] The input of a CNN consists of matrices representing data. Typical application areas of CNNs are especially the processing of image data and video data.

[0004] On the convolutional layer, a discrete convolution is typically performed with the aid of a filter kernel. Typically, in the case of using a discrete convolution, the dimensions of the input do not change.

[0005] On the pooling layer, an operation of combining is applied to the input matrix. Here, typically, the dimensions of the matrix are reduced.

[0006] A Region Proposal Network (RPN) is a form of CNN that estimates whether an object is located in the provided representation of the surrounding environment (region) for a determined position (anchor point) of the provided representation of the surrounding environment. For this purpose, the RPN calculates for each anchor point a so-called objectness score (Objectness-Score) and a so-called bounding box ("Bounding Box"). The objectness score represents the probability (confidence) that an object exists for this anchor point. The bounding box represents the spatial extent scale of the object.

[0007] Due to the high spatial resolution of the anchor points, the RPN typically estimates a relatively large number of redundant object hypotheses for the objects located in the processed surrounding environment. The subsequent processing task is then to determine a certain amount of object detections starting from a certain amount of redundant object hypotheses, where the certain amount of object detections contains as few estimates as possible for each actual object.

[0008] It is known from Alexander Neubeck and Luc Van Goo, "Efficient nonmaximum suppression.", 18th International Conference on Pattern Recognition (ICPR'06), 3:850 - 855, 2006 that by means of so-called Non-maximum Suppression (NMS), object hypotheses with the highest object existence likelihood scores are selected from all redundant object hypotheses and the remaining object hypotheses are discarded.

[0009] In Gregory P. Meyer, Ankit Laddha, Eric Kee, Carlos Vallespi-Gonzalez, and Carl K. Wellington, "Lasernet: An efficient probabilistic 3d object detector for autonomous driving", 2019, a method is proposed instead of NMS in which overlapping object hypotheses are iteratively combined (clustered) and then fused.

[0010] A serious deficiency of traditional post-processing according to the prior art is information loss, which is caused by the discarding of redundant object hypotheses.

[0011] Furthermore, in the methods for post-processing according to the prior art, the identification of redundant object hypotheses is phenomenologically implemented without considering the known output structure of the RPN as model knowledge.

[0012] The quality of object hypotheses is also obtained purely phenomenologically by learned hyperparameters. Summary of the Invention

[0013] Against this background, the present invention implements a method for detecting objects in the vehicle's surroundings in relation to sensor signals. The sensor signals come from sensors for sensing the vehicle's surroundings. These sensor signals represent the vehicle's surroundings.

[0014] The method comprises the following steps:

[0015] Processing the sensor signals by means of a region candidate network to obtain at least one object hypothesis for each anchor point, wherein the object hypothesis comprises an object probability and a bounding box;

[0016] Selecting the best object hypothesis by means of a quality model, wherein the quality model depends on the anchor point and the bounding box of the object hypothesis;

[0017] Assume an identification-redundant object hypothesis relative to the selected object hypothesis, wherein the redundant object hypothesis is identified with respect to the anchor point of the redundant object hypothesis by means of an objective function assigned to the region candidate network;

[0018] Fuse the selected object hypothesis with the identified redundant object hypothesis for object detection.

[0019] The object detection task can be improved by the method of the present invention. This is achieved by considering the objective function assigned to the region candidate network and the available model knowledge.

[0020] Currently, sensors for sensing the surrounding environment of a vehicle can be understood as environmental sensors. Environmental sensors are based on the sensing of different physical effects. The most well-known environmental sensors include video sensors, radar sensors, laser sensors, and acoustic sensors, especially ultrasonic sensors. Combinations of the above sensors are conceivable. In addition, other sensor technologies suitable for generating signals representative of the vehicle's surrounding environment are conceivable.

[0021] Currently, the object probability can be understood as the object presence likelihood score output by the RPN, that is, the presence of the (searched) object at the corresponding anchor point in the processed surrounding environment representation signal.

[0022] Currently, the bounding box can be understood as an enclosing boundary that is placed around the hypothesized (searched) object at the corresponding anchor point.

[0023] Currently, the objective function assigned to the region candidate network or the RPN objective function can be understood as a function assigned to or associated with the used RPN, which indicates whether the given position is near the bounding box starting from the given bounding box and the given position of the processed surrounding environment. Here, "near" can currently be understood as a given distance around the bounding box.

[0024] The present invention is based on the following recognition: The relative position of the anchor point of the object hypothesis with respect to the bounding box of the object hypothesis is a measure of the quality of the object hypothesis. In other words, the present invention is based on a quality model related to the anchor point.

[0025] A suitable implementation of the quality model related to the anchor point is the covariance matrix of the parameters of the bounding box of the object hypothesis.

[0026] To calculate this quality model, for example, the object hypotheses of the RPN of the validation data set can be analyzed and evaluated. Since the relative anchor point position is known for each hypothesis, the corresponding expected quality can thus be obtained statistically (in the form of an average value and a covariance matrix).

[0027] Furthermore, the present invention is based on the following recognition: for a given object hypothesis, based on the RPN objective function, it is known which additional anchor points also support the object hypothesis.

[0028] After the fusion step, the number of object hypotheses can be reduced.

[0029] Although the number of object hypotheses is reduced, it is still conceivable that the remaining object hypotheses have intersecting bounding boxes. This can indicate that the same object in the surrounding environment is supported by two separate object hypotheses. By suitable methods, such as those known from Alexander Neubeck and Luc Van Gool: Efficient nonmaximum suppression, 18th International Conference on Pattern Recognition (ICPR’06), 3:850–855, 2006 or Gregory P. Meyer, Ankit Laddha, Eric Kee, Carlos Vallespi-Gonzalez, and Carl K. Wellington. Lasernet: An efficient probabilistic 3d object detector for autonomous driving, 2019, this remaining redundancy can be resolved.

[0030] According to an embodiment of the method according to the invention, in the step of identification by means of the RPN objective function, such object hypotheses are identified as redundant: the object hypotheses are located within a predefined distance with respect to the selected object hypothesis or with respect to the bounding box of the selected object hypothesis according to their anchor points.

[0031] According to this embodiment, when the object hypothesis is located within a predefined distance with respect to the selected object hypothesis according to its anchor point, the object hypothesis is identified as redundant. Thus, redundant object hypotheses are identified according to the criterion of proximity. In other words, the selection of the object hypotheses identified as redundant is carried out independently of their object existence likelihood scores. Thus, according to this embodiment, so-called false positive object detections can be recognized and the false positive object detections can be blocked accordingly or their number can be reduced.

[0032] False positive object detections occur when, despite the fact that there is actually no object in the corresponding region of the surrounding environment, a high object existence likelihood score is obtained for an anchor point due to an artifact in the input data or due to the insufficiency of the RPN used.

[0033] Typically, false positive object detections are individual defective locations. Correspondingly, the object hypotheses at this anchor point are not supported by such object hypotheses: the object hypotheses are located nearby according to the objective function of the RPN and are thus redundant object hypotheses.

[0034] Then, in the fusion step, the selected object hypotheses are identified as defective and accordingly do not result in object detection.

[0035] This can be achieved, for example, by the following method: in the fusion step, the corresponding object probabilities (object existence likelihood scores) of the object hypotheses to be fused are fused. Thereby, the correct low object probabilities of the nearby object hypotheses can be used to reduce the single defective high object probability, so that the fused object hypotheses ultimately do not support object detection.

[0036] According to an embodiment of the method of the present invention, in the fusion step, the selected object hypotheses and the identified object hypotheses are fused according to their corresponding quality models.

[0037] For the fusion of object hypotheses, a suitable fusion mechanism is appropriate. This especially includes the weighted least squares method. In addition, the median method, the mean shift method or the random sample consensus method (Ransac) can be considered.

[0038] Regardless of the selected fusion mechanism, this embodiment has the following advantages: for fusion, the quality models of the corresponding object hypotheses are utilized.

[0039] According to an embodiment of the method of the present invention, the quality model of the object hypothesis depends on the relative position of the anchor point of the object hypothesis with respect to the bounding box of the object hypothesis.

[0040] This embodiment is based on the following recognition: the quality of the object hypothesis depends to a large extent on the relative position of the anchor point with respect to the belonging bounding box. Different from the phenomenological method for determining the quality of the object hypothesis (the phenomenological method can be created, for example, by means of unsupervised learning), the proposed quality model comes from explicit modeling. This method brings existing model knowledge into object detection and thus leads to improved detection results.

[0041] An empirical method has shown that the closer the corresponding anchor point of the object hypothesis is to the center of the bounding box, the higher the quality of the object hypothesis.

[0042] According to an embodiment of the method of the present invention, the quality model additionally depends on the region candidate network.

[0043] This embodiment is based on the following recognition: The quality of the object hypothesis is closely associated with the Region Proposal Network (RPN) in which the object hypothesis has been established. By additionally considering the RPN when establishing the quality model, in addition to the relative position of the anchor with respect to the belonging bounding box, an improved quality model can be created. The quality model for the object hypothesis is best if the corresponding bounding box encloses the detected object as tightly as possible.

[0044] According to one embodiment, the quality model additionally depends on other influences on the quality of the object hypothesis.

[0045] One of these influences may be the geometric relationship between the sensor and the object to be detected. This geometric relationship can be, for example, the distance between the sensor and the object to be detected.

[0046] For example, if it is known that object hypotheses at a large distance have predictably poor quality, this modeling approach is useful.

[0047] According to one embodiment of the method of the present invention, in the fusion step, the fusion of the selected object hypothesis and the identified object hypothesis is continued as the new selected object hypothesis, and then, the method is continued with an identification step using the new selected object hypothesis.

[0048] According to this embodiment, the fusion of the object hypothesis initially selected as the best hypothesis and the identified object hypothesis is performed iteratively.

[0049] According to an alternative embodiment of the method of the present invention, the selection step is postponed. This embodiment starts with an identification step. In the identification step, in this embodiment, redundant object hypotheses are not identified based on the selected (best) object hypothesis, but rather based on the entire set of object hypotheses. Then, after the fusion step, a selection step is performed based on the fused object hypotheses, which no longer have redundancy.

[0050] Another aspect of the present invention is a device configured to perform all steps of the method according to the present invention.

[0051] Another aspect of the present invention is a computer program configured to perform all steps of the method according to the present invention.

[0052] Another aspect of the present invention is a machine-readable storage medium on which a computer program according to the present invention is stored. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings.

[0054] The drawings show:

[0055] Figure 1 Schematic diagram of the regional candidate network according to the present invention;

[0056] Figure 2 Flowchart of an embodiment of the method according to the present invention. Detailed implementation manners

[0057] Figure 1 Schematically shows the working manner of the regional candidate network (RPN) 10.

[0058] The input data 11 is processed in the RPN 10 and generates a network output 12 with a fixed structure. The network output 12 includes anchor points 121a, 121b represented as dots 121a or crosses 121b. The anchor points 121a represented as dots represent the following anchor positions: at these anchor positions, (dynamic) objects are estimated according to the objective function of the RPN. The anchor points 121b represented as crosses represent the following anchor positions: at these anchor positions, no (dynamic) objects are estimated according to the objective function of the RPN.

[0059] Object hypotheses 13 are obtained for the anchor points 121a, 121b. Such object hypotheses include the probability of object estimation, i.e., the so-called object existence likelihood score, and include the so-called bounding box 131.

[0060] The bounding box 131 includes a width w and a height h, and relative offsets Δw in the vertical direction and Δh in the horizontal direction with respect to the position of the anchor point 121a.

[0061] Figure 2 Shows a flowchart of an embodiment of the method 200 of the present invention.

[0062] In step 21, the input data 11 representing the surrounding environment, typically sensor data, is provided to the regional candidate network (RPN) 10. After being processed by the RPN 10, there are one or more object hypotheses 13.

[0063] In step 22, the object hypotheses 13 are subsequently processed for object detection.

[0064] In step 221, the best object hypothesis 13 is selected according to a quality model. Here, the relative positions of the anchor points 121a, 121b of the object hypothesis 13 with respect to the bounding box 131 of the object hypothesis 13 can be considered as the quality model. Appropriately, the vertical offset Δw and the horizontal offset Δh are considered.

[0065] In step 222, starting from the selected object hypothesis 13, object hypotheses 13 that support the object to be detected are identified according to the RPN objective function.

[0066] In step 223, the selected object hypothesis 13 is fused with the identified object hypothesis 13.

[0067] The result of the subsequent processing may include the objects detected in the provided surrounding environment. If no object appears in the provided surrounding environment, the result of the method should reflect this and accordingly no object is detected in the provided surrounding environment.

Claims

1. A method for detecting an object in the surroundings of a vehicle based on sensor signals of sensors for sensing the surroundings of the vehicle, wherein, The sensor signal represents the vehicle's surroundings, and the method has the following steps: - Processing the sensor signal with a region candidate network to obtain at least one object hypothesis for each anchor point, where the object hypothesis includes an object probability and a bounding box; - Selecting the best object hypothesis with a quality model, where the quality model depends on the anchor point of the object hypothesis and the bounding box; - Identifying object hypotheses redundant with respect to the selected object hypothesis, where the redundant object hypotheses are identified based on the anchor points of the redundant object hypotheses with the aid of an objective function assigned to the region candidate network; - Fusing the selected object hypothesis with the identified redundant object hypotheses for object detection.

2. The method according to claim 1, wherein, In the step of performing the identification with the aid of the objective function, an object hypothesis is identified as redundant if the object hypothesis is within a pre-given distance from the selected object hypothesis and / or from the bounding box of the selected object hypothesis based on the objective function according to its anchor point.

3. The method according to claim 1 or 2, wherein In the step of the fusion, the selected object hypothesis and the identified object hypothesis are fused according to their corresponding quality models.

4. The method according to claim 1 or 2, wherein The quality model of the object hypothesis depends on the relative position of the anchor point of the object hypothesis with respect to the bounding box of the object hypothesis.

5. The method according to claim 4, wherein The quality model additionally depends on the region candidate network.

6. The method according to claim 4, wherein The quality model additionally depends on other influences on the quality of the object hypothesis.

7. The method according to any one of claims 1, 2, 5, and 6, wherein In the step of the fusion, the fusion of the selected object hypothesis and the identified object hypothesis is continued as a newly selected object hypothesis, and subsequently, the method is continued with the newly selected object hypothesis in the step of the identification.

8. The method according to claim 6, wherein, The quality model depends on the geometric relationship between the sensor and the object to be detected.

9. An apparatus for performing all steps of the method according to any one of claims 1 to 8.

10. A computer program product having stored thereon a computer program configured to perform all steps of the method according to any one of claims 1 to 8.

11. A machine-readable storage medium having stored thereon a computer program configured to perform all steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Automobile driving scene target detection method based on deep convolutional neural network

    CN107169421A

  • Method for positioning video targets on basis of region candidate frame tracking

    CN108280844A