Method and device for perceiving the environment of a vehicle that is at least partially automated.

By evaluating sensor data at a lower resolution with defined focus regions and using neural networks, the method achieves a balance between computing power and accuracy in environmental perception for automated vehicles.

DE102020211023B4Active Publication Date: 2026-01-15VOLKSWAGEN AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102020211023
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-09-01
Publication Date
2026-01-15
Estimated Expiration
2040-09-01

AI Technical Summary

Technical Problem

Automated vehicles face challenges in achieving robust and accurate environmental perception with limited computing power, necessitating a method that balances computing requirements with accuracy for real-time operation, especially in emergency situations.

Method used

A method and device that evaluate sensor data at a lower initial resolution overall, with defined focus regions at a higher resolution, using neural networks to reduce computing power while maintaining high accuracy in critical areas.

Benefits of technology

This approach allows for reduced computing power consumption while ensuring high accuracy in relevant focus regions, enhancing the vehicle's environmental perception capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods for perceiving the environment of a vehicle that is at least partially automated (50), wherein the environment of the vehicle (50) is detected by means of at least one sensor (51), wherein sensor data (10) detected by the at least one sensor (51) are evaluated by means of an evaluation device (2) using at least one evaluation method in a first resolution, and wherein the acquired sensor data (10) in at least one defined focus region (8) are evaluated by the evaluation device (2) using at least one evaluation method in a second resolution, wherein the first resolution is smaller than the second resolution, and wherein the evaluation results (11,12) are combined and output, wherein, in the case of several defined focus regions (8), each focus region (8) is assigned a priority (17), and wherein, when evaluating the respective associated sensor data (10), the focus regions (8) are processed in the order of the assigned priorities (17).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method and a device for perceiving the environment of a vehicle that is at least partially automated.

[0002] Robust and accurate environmental perception is essential for automated driving. Artificial neural networks form the backbone of this environmental perception. Cameras are the predominant sensor used to detect the vehicle's surroundings, as they provide a high information density and are relatively inexpensive. Computing power in automated vehicles is typically very limited. Therefore, environmental perception methods must operate within this limited capacity. At the same time, the method must be capable of real-time operation and accurate enough to handle emergency situations.

[0003] From DE 10 2016 213 494 A1, a camera device for capturing the surrounding area of ​​a vehicle is known, comprising an optronic system including a high-resolution image acquisition sensor and a wide-angle optic, wherein the optronic system is configured to output a sequence of images of the surrounding area with a periodic change between high-resolution and reduced-resolution images.

[0004] The invention is based on the objective of improving a method and a device for perceiving the environment of an at least partially automated vehicle, particularly with regard to the required computing power and accuracy.

[0005] The problem is solved according to the invention by a method with the features of claim 1 and a device with the features of claim 8. Advantageous embodiments of the invention are set forth in the dependent claims.

[0006] In particular, a method for perceiving the environment of a vehicle that is at least partially automated is provided, wherein the vehicle's environment is detected by means of at least one sensor, wherein sensor data detected by the at least one sensor are evaluated by means of an evaluation device using at least one evaluation method in a first resolution, and wherein the detected sensor data in at least one defined focus region are evaluated by means of the evaluation device using the at least one evaluation method in a second resolution, wherein the first resolution is smaller than the second resolution, and wherein the evaluation results are combined and output.

[0007] Furthermore, in particular a device for perceiving the environment of a vehicle that is at least partially automated is provided, comprising an evaluation device, wherein the evaluation device is configured to evaluate sensor data acquired by at least one sensor using at least one evaluation method in a first resolution, and to evaluate the acquired sensor data using the at least one evaluation method in at least one defined focus region in a second resolution, wherein the first resolution is smaller than the second resolution, and to combine and output the evaluation results.

[0008] The method and the device enable a compromise between the computing power required for evaluating the acquired sensor data and the accuracy of the evaluation. This is achieved by evaluating the sensor data at a first resolution that is lower than the original resolution. Additionally, at least one focus region is defined within the acquired sensor data; that is, a portion of the sensor data is reduced compared to the original sensor data. For example, a section of a two-dimensional camera image can form such a focus region. Within this at least one focus region, the sensor data is evaluated at a second resolution. The first resolution is lower than the second resolution. Specifically, the second resolution corresponds to the original, i.e., the full, resolution of the acquired sensor data.Since the captured sensor data is evaluated overall at a lower resolution than the original resolution, and only the at least one focus region is evaluated at a higher, in particular full, resolution, the computing power required for evaluation can be reduced overall, and yet high accuracy can still be achieved in the at least one focus region.

[0009] A sensor is, in particular, a camera that detects the vehicle's surroundings. However, a sensor can also be a lidar sensor, a radar sensor, or an ultrasonic sensor.

[0010] The acquired sensor data can be one-dimensional or multi-dimensional. In particular, the acquired sensor data is two-dimensional. The acquired sensor data primarily consists of captured camera images. However, the acquired sensor data can also include point clouds from a lidar or radar sensor, or ultrasonic data.

[0011] An evaluation method is, in particular, an evaluation method for environmental perception and / or environmental interpretation in which at least one perceptual function is performed. Such a perceptual function can be, for example, one of the following: object recognition, determination of bounding boxes, semantic segmentation, instance segmentation, or panoptic segmentation, etc.

[0012] The evaluation process is provided and carried out primarily using a machine learning method. Specifically, it may be intended to use (trained) neural networks. For example, a first trained neural network may be provided by the evaluation system, which processes the acquired sensor data at the lower first resolution and executes one or more of the aforementioned evaluation methods. A second trained neural network, also provided by the evaluation system, evaluates the acquired sensor data in the at least one focus region at the higher, specifically the original, second resolution.

[0013] In principle, however, the evaluation procedure can also be provided and carried out in other ways and / or with other means; that is, classic methods of sensor data evaluation can also be used.

[0014] It is specifically intended that the acquired sensor data be downscaled to the first resolution. This is done using methods known per se. In the case of camera images as acquired sensor data, for example, the number of image elements (pixels) is reduced, so that the image is scaled down to a lower total number of image elements. For the at least one focus region, the acquired sensor data is accordingly cropped to a specific area, so that the evaluation process, for example, the second neural network, only receives the sensor data within the area of ​​the at least one focus region as input data. It may be provided that the second neural network is instantiated again for each defined focus region or is provided and executed separately.

[0015] Parts of the device, in particular the evaluation unit and its components, can be designed individually or collectively as a combination of hardware and software, for example as program code that runs on a microcontroller or microprocessor. However, it is also possible for parts to be designed individually or collectively as an application-specific integrated circuit (ASIC).

[0016] A vehicle is, in particular, a motor vehicle. However, a vehicle can also be any other land, rail, water, air, or spacecraft.

[0017] It is intended that, for multiple defined focus regions, each focus region is assigned a priority. When evaluating the associated sensor data, the focus regions are processed in the order of their assigned priorities. The priority can be assigned, for example, based on an object class associated with an object within the context of environmental perception. Thus, the vulnerability of other road users can be prioritized, with pedestrians and cyclists receiving a higher priority compared to other vehicles. It can also be stipulated that children receive a higher priority than adults, as children's behavior is generally less predictable than that of adults. This makes it possible to concentrate limited computing power on the highest-priority focus regions.

[0018] In one embodiment, at least one focus region is defined by a rule-based system depending on the vehicle's position, its planned trajectory, and / or its surrounding context. This allows the focus region to be specifically directed to particularly relevant areas in the vehicle's environment. For example, it may be possible to mark or store particularly relevant areas in the environment on a map, depending on the vehicle's position. For instance, the map may indicate that a school is located near the vehicle's current position. Based on this information, a focus region can be defined, for example, to encompass the roadside and sidewalk in front of the school.Further examples of areas that can be examined more closely using a focus region during evaluation include: crosswalks, intersections of side streets or driveways, oncoming traffic, etc. If a planned trajectory is known, a focus region can also be defined based on this trajectory. For example, the planned trajectory might involve a left turn and crossing into oncoming traffic. Based on the planned trajectory, a focus region can then be placed on the opposite lane to detect approaching vehicles with improved resolution. An environmental context can be derived from both the acquired sensor data and an environmental map to, for example, better detect the areas within a shared space and the surrounding areas using defined focus regions.The rule-based system contains, in particular, links that connect a vehicle's position and / or a planned trajectory, or a type or class of planned trajectory (left turn, right turn, overtaking, etc.), and / or an environmental context with at least one focus region. This focus region can be defined, for example, by positional information relative to the vehicle and / or in relation to the captured sensor data. For instance, a focus region can be defined by specifying image element areas in captured camera images. The rule-based system might, for example, define a focus region as covering an area around the center of a preceding road. Depending on the specific environment, further focus regions can then be defined according to predefined rules (e.g., school ahead, crosswalk ahead, bike path crossing, etc.).The rule-based system can be provided, for example, by means of a logic module set up for this purpose in the device.

[0019] In one embodiment, temporal changes in the acquired sensor data are detected, and at least one focus region is defined based on at least one detected temporal change in the acquired sensor data. This allows a focus region to be defined or newly created when something in the environment changes over time. A temporal change here means, in particular, that in temporally successive sensor data acquisitions, for example, successive camera images ("frames"), these differ at least in certain areas. For example, a traffic light depicted in successive camera images might change from a green signal to a red signal.Since changes represent a potential danger to the vehicle (or other road users causing the change) and are therefore of greater interest in environmental perception, focus regions can be dynamically defined depending on the environment or a change of state in the environment.

[0020] In one embodiment, it is provided that movements in the vicinity of the vehicle are detected in the acquired sensor data, wherein at least one focus region is defined depending on at least one detected movement in the acquired sensor data.

[0021] This allows, for example, sensor data depicting moving objects in the environment to be evaluated with higher resolution. The underlying principle here is that moving objects in the environment, unlike static objects, are potentially more dangerous and therefore require more attention and greater accuracy in their evaluation. It may be intended that at least one evaluation method provides information on the movement of objects in the environment, for example, using appropriate object motion detection techniques.

[0022] In one embodiment, at least one focus region is defined using a machine learning method. For example, a neural network can be trained to estimate areas where focus regions should be defined, based on the acquired sensor data and / or the vehicle's position and / or the surrounding environment (e.g., pedestrian zone, highway, crosswalk, underground parking garage, etc.). This allows focus regions to be estimated and defined even for unknown environments.

[0023] In one embodiment, the system takes into account the uncertainty of the evaluation results estimated by the evaluation method when combining them. This allows the evaluation results to be weighted by their estimated uncertainty. For example, if camera images are evaluated as sensor data and an image element-by-image assignment to objects is performed (e.g., for multiple object classes), an uncertainty can be specified for each image element, which can then be used to evaluate the assignment. Such image element-by-image assignment and the associated uncertainty specification can be performed, for example, using a neural network.The evaluation results then provide an assignment with associated uncertainties for the entire camera image at a lower initial resolution and an assignment with associated uncertainties for the (at least one) focus region at a second, higher resolution. When combining these assignments, they can be weighted according to their respective uncertainties and summarized, in particular by summing them. When using neural networks, assigning image elements to objects (or object classes) in peripheral areas of the focus region is subject to greater uncertainty than in more central areas, because less information is available at the edges (since information beyond the clipped edges is missing).When combining the results, a greater weight can be given to the evaluation results of the first lower resolution based on the respective assigned uncertainties, so that overall the result in the peripheral areas of the focus region is improved.

[0024] In one embodiment, it is provided that when a focus region is abandoned, an evaluation result is retained in the second resolution for at least a predetermined period. This allows the evaluation results generated for the higher second resolution to continue to be used, at least for the predetermined period, and thus improve accuracy in environmental perception despite the abandonment of this focus region. Abandoning a focus region here means, in particular, that the sensor data in the abandoned focus region are no longer evaluated in the second, higher resolution, but only in the first, lower resolution. The predetermined period can, in particular, range from a few hundred milliseconds to several seconds.

[0025] Further features for the design of the device result from the description of embodiments of the method. The advantages of the device are the same in each case as in the embodiments of the method.

[0026] Furthermore, a vehicle is created, comprising at least one device according to one of the described embodiments.

[0027] The invention is explained in more detail below with reference to preferred embodiments and the figures. These show: Fig. 1 a schematic representation of an embodiment of the device for perceiving the environment of a vehicle that is at least partially automated; Fig. 2a a schematic representation to illustrate the procedure (captured sensor data); Fig. 2b a schematic representation to illustrate the procedure (evaluation result with reduced first resolution); Fig. 3a a schematic representation to illustrate a procedure according to the method described in this disclosure; Fig. 3b a schematic representation of a section of sensor data belonging to a defined focus region; Fig. 3c a schematic representation of resolution-reduced sensor data; Fig. 4 A schematic representation to illustrate an overall result compiled from the evaluation results according to the procedure.

[0028] In Fig. Figure 1 shows a schematic representation of an embodiment of the device 1 for perceiving the environment of a vehicle 50 that is at least partially automated. The device 1 performs the method described in this disclosure for perceiving the environment of a vehicle 50 that is at least partially automated.

[0029] The device 1 comprises an evaluation unit 2. The evaluation unit 2 includes a resolution reduction module 3, a clipping module 4, a first neural network 5, a second neural network 6, and a merging module 7. Parts of the device 1, in particular the evaluation unit 2 and its components, can be configured individually or collectively as a combination of hardware and software, for example, as program code executed on a microcontroller or microprocessor.

[0030] The device 1 receives sensor data 10 acquired from at least one sensor 51 of the vehicle 50. The sensor 51 is, in particular, a camera, and the sensor data 10 are, in particular, camera images captured from the surroundings.

[0031] The resolution of the acquired sensor data 10, in particular the acquired camera images, is reduced to a first resolution by the resolution reduction module 3, and reduced-resolution sensor data 10- is provided. In parallel, the cropping module 4, depending on a defined focus region 8, extracts a section 9 from the acquired sensor data 10, in particular from the camera images. A second resolution of the section 9 corresponds to a resolution of the original sensor data 10. The first resolution is therefore lower than the second resolution.

[0032] The sensor data 10 in the first resolution are fed to the first neural network 5. The section 9 in the second resolution is fed to the second neural network 6. The first neural network 5 and the second neural network 6 perform the same evaluation procedure, in particular a perception function, for example, semantic segmentation, in which each image element in the captured camera images is assigned an object class; that is, for each image element, it is estimated which object the image element represents. Evaluation results 11, 12 are output by the neural networks 5, 6 and fed to the merging module 7.

[0033] Although the procedure is described using neural networks 5, 6, other machine learning methods or classical sensor data evaluation can also be used to evaluate the sensor data 10 in the first resolution and the section 9 from the sensor data 10 corresponding to the focus region 8 in the second resolution. It should be noted, however, that the evaluation method, i.e., the perceptual function for environmental perception or environmental interpretation (e.g., object recognition, semantic segmentation, semantic instance segmentation, bounding boxes, etc.), is the same for both approaches.

[0034] The merging module 7 combines the evaluation results 11 and 12 into a single overall result 13. In the case of camera images, for example, evaluation result 11 is upscaled back to its original resolution (which corresponds to the second resolution), and evaluation result 12 for section 9 is inserted into the upscaled evaluation result 11, image element by image, at the correct position. It may also be possible to calculate an average of the evaluation results 11 and 12.

[0035] When combining the evaluation results 11 and 12, it may be possible to take into account an uncertainty estimated by the evaluation method. The neural networks 11 and 12 estimate the evaluation results 11 and 12, for example, in the form of object classes assigned to individual image elements, and simultaneously provide an additional uncertainty for each image element-wise estimate. A weighting factor can be defined, or already defined, based on each estimated uncertainty when combining the results. In particular, at the edges of the section 9, where estimates from the second neural network 6 have a greater uncertainty due to a lack of additional information from areas beyond the edge, the overall result 13 can be improved in this way, since the evaluation result 11 from the first neural network 5 can be given greater weight in these areas.

[0036] It can be provided that at least one focus region 8 is determined by means of a rule-based system 18 depending on the position 53 of the vehicle 50 and / or a planned trajectory 54 of the vehicle 50 and / or an environmental context. To provide the rule-based system 18, the device 1 can, for example, have a logic module 19 in which the rule-based system 18 is stored and applied. For example, depending on the position 53 of the vehicle 50, a street and / or environmental map can be consulted to determine which focus regions 8 are located in the environment. Depending on the planned trajectory 54, a focus region 8 can, for example, be placed on oncoming traffic (e.g., when turning and crossing an oncoming lane). An environmental context could, for example, be a play street, a residential area with families with small children, or a wildlife crossing, etc.This concerns focus regions 8, which are then, for example, defined as areas near vehicles parked at the edge of the play street, oncoming traffic, or on the roadside bordering the forest. Position 53 and the planned trajectory 54 are supplied to the logic module 19, for example, by a navigation system and / or a vehicle control unit (not shown) of vehicle 50.

[0037] It can be provided that temporal changes in the acquired sensor data 10 are detected, wherein at least one focus region 8 is defined depending on at least one detected temporal change in the acquired sensor data 10. For this purpose, the device 1 can, for example, have a change detection module 14 with which areas in the acquired sensor data 10 are detected in which a change (e.g. a change in a traffic light phase or a flashing of a turn signal of another vehicle, etc.) is detected.

[0038] It can be provided that movements in the vicinity of the vehicle 50 are detected in the acquired sensor data 10, wherein at least one focus region 8 is defined depending on at least one detected movement in the acquired sensor data 10. For this purpose, the device 1 can, for example, have a motion detection module 15 that detects movements in the acquired sensor data 10, in particular acquired camera images, and defines an associated area, in particular image element area in at least one camera image, as the focus region 8.

[0039] It may be provided that at least one focus region 8 is determined by means of a machine learning method. For example, another neural network 16 can be trained to estimate and determine focus regions 8 depending on acquired sensor data 10, in particular acquired camera images, and / or other information (an environmental context and / or a position 53 of the vehicle 50, etc.).

[0040] It is intended that, for multiple defined focus regions 8, each focus region 8 is assigned a priority 17, and that when evaluating the respective sensor data 10, the focus regions 8 are processed in the order of their assigned priorities 17. The priorities 17 can be selected, for example, depending on the vulnerability of another road user encompassed by the focus region 8. For this purpose, a prioritization module (not shown) of the device 1 can be provided.

[0041] It may be provided that, when a focus region 8 is abandoned, an evaluation result 12 is retained in the second resolution for at least a predetermined period of time. This is done, for example, in the merging module 7.

[0042] The device 1 makes it possible to find a compromise in the execution of the evaluation procedures with regard to the required computing power and accuracy. Although computing power can be reduced because the acquired sensor data 10 are processed entirely at a reduced resolution, relevant areas can be evaluated at a higher resolution, in particular at the original resolution, by defining focus regions 8.

[0043] The overall result 13 can, for example, be fed to a maneuver planner 52 of the vehicle 50 and taken into account there in the maneuver planning for automated driving.

[0044] In the Fig. Figure 2a shows a schematic representation to illustrate the procedure. It depicts captured sensor data 10 in the form of a camera image. Fig. 2b shows a semantic segmentation as evaluation result 11, in which each image element of the camera image is assigned an object class (indicated by different hatching patterns). For this purpose, the data in the Fig. The sensor data 10 shown in Figure 2a, i.e., the camera image, is used at a lower resolution than the original resolution; that is, the camera image has been scaled down. A child 30 standing at the end of the line of vehicles visible on the right in the camera image is not detected by the semantic segmentation due to the reduced resolution and therefore does not appear in the evaluation result 11.

[0045] In the Fig. Figure 3a shows the procedure according to the method described in this disclosure. In the acquired sensor data 10, that is, in the camera image, (at least) one focus region 8 is defined. A simple case is shown in which the focus region 8 encompasses the middle of the road ahead. A section 9 corresponding to the focus region 8 ( Fig. 3b) is evaluated at a second resolution, which corresponds to the original resolution of the sensor data 10. The complete sensor data 10, i.e., the complete camera image, is scaled down to a lower first resolution, to resolution-reduced sensor data 10- ( Fig. 3c).

[0046] After cutting out the focus region 8 and downscaling the sensor data 10, the focus region 8 and the resolution-reduced sensor data 10 are each evaluated by the evaluation procedure, which again includes a semantic segmentation as an example, in which each image element is assigned an object class.

[0047] The resulting evaluation results are then combined to form an overall result 13. Fig. Figure 4 shows a schematic representation of the combined overall result 13.

[0048] The overall result 13 clearly shows that the image elements belonging to the child 30 were correctly classified and assigned the object type "person" or "child" because a higher resolution could be selected in this area than in the remaining areas by defining a focus region 8. The semantic segmentation thus detected or recognized the child 30 in the captured sensor data 10.

[0049] Therefore, by means of the method and device 1 described in this disclosure, a compromise can be found between the necessary computing power and the accuracy in evaluating sensor data 10.

[0050] For the sake of clarity, only one focus region 8 was defined and shown in the examples described above. In principle, further focus regions can be defined. The procedure is analogous in each case. Reference symbol list 1 Device 2 Evaluation unit 3. Trigger reduction module 4 Cutting module 5 first neural network 6 second neural network 7 Merging Module 8 Focus region 9 Excerpt 10 recorded sensor data 10 reduced-resolution sensor data 11 Evaluation result 12 Evaluation result 13 Overall result 14 Change Detection Module 15 Motion Detection Module 16 other neural networks 17 Priority 18 rule-based system 19 Logic module 30 children 50 vehicles 51 Sensor 52 maneuver planners 53rd position 54 planned trajectory

Claims

[1] Method for perceiving the environment of a vehicle that is at least partially automated (50), wherein the environment of the vehicle (50) is detected by means of at least one sensor (51), wherein sensor data (10) detected by the at least one sensor (51) are evaluated by means of an evaluation device (2) using at least one evaluation method in a first resolution, and wherein the acquired sensor data (10) in at least one defined focus region (8) are evaluated by the evaluation device (2) using at least one evaluation method in a second resolution, wherein the first resolution is smaller than the second resolution, and wherein the evaluation results (11,12) are combined and output, wherein, in the case of several defined focus regions (8), each focus region (8) is assigned a priority (17), and wherein, when evaluating the respective associated sensor data (10), the focus regions (8) are processed in the order of the assigned priorities (17). [2] Method according to claim 1, characterized by , that at least one focus region (8) is determined by means of a rule-based system (18) depending on a position (53) of the vehicle (50) and / or a planned trajectory (54) of the vehicle (50) and / or an environmental context. [3] Method according to claim 1 or 2, characterized by , that temporal changes in the acquired sensor data (10) are detected, wherein at least one focus region (8) is determined depending on at least one detected temporal change in the acquired sensor data (10). [4] Method according to any of the preceding claims, characterized by , that movements in the vicinity of the vehicle (50) are detected in the acquired sensor data (10), wherein at least one focus region (8) is determined depending on at least one detected movement in the acquired sensor data (10). [5] Method according to any of the preceding claims, characterized by , that at least one focus region (8) is determined using a machine learning method. [6] Method according to any of the preceding claims, characterized by , that when combining the evaluation results (11,12) an uncertainty of the evaluation results (11,12) estimated by the evaluation procedure is taken into account. [7] Method according to any of the preceding claims, characterized by , that when a focus region (8) is abandoned, an evaluation result (12) is retained in the second resolution at least for a specified period of time. [8] Device (1) for perceiving the environment of a vehicle (50) that is at least partially automated, comprising: an evaluation facility (2), wherein the evaluation device (2) is configured to evaluate sensor data (10) acquired by at least one sensor (51) by means of at least one evaluation method in a first resolution, and to evaluate the acquired sensor data (10) by means of the at least one evaluation method in at least one defined focus region (8) in a second resolution, wherein the first resolution is smaller than the second resolution, and to combine and output the evaluation results (11,12), wherein the evaluation device (2) is further configured to assign a priority (17) to each of the focus regions (8) when several focus regions (8) are defined, and to process the focus regions (8) in the order of the assigned priorities (17) when evaluating the respective associated sensor data (10). [9] Vehicle (50) comprising at least one device (1) according to claim 8.

Citation Information

Patent Citations

  • Camera device and method for detecting an environment area of a vehicle of one's own

    DE102016213494A1