Radar-based occupancy grid map
By generating a radar-based three-dimensional occupancy grid map and feature grid map, and using CNN for refinement, the improved space for object detection in vehicle assistance and control functions is solved, achieving higher precision vehicle environment detection and more reliable auxiliary system performance.
Patent Information
- Application Number
- CN202411536671.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-10-31
- Publication Date
- 2025-06-06
AI Technical Summary
Existing vehicle assistance and control functions rely on sensor detection of the vehicle's external environment, especially in object detection.
By generating a radar-based three-dimensional occupancy grid map (3D OGM) and multiple feature grid maps (FGMs), and using a convolutional neural network (CNN) for refinement processing, the refined OGM is generated for use by vehicle auxiliary systems.
It improves the detection accuracy of the vehicle environment and the performance of the auxiliary system, enhances the navigation and safety functions of the vehicle, reduces the false alarm rate and improves the reliability of the driver assistance system.
Smart Images

Figure CN120107925A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to improvements in operating a vehicle, and particularly to methods and systems for generating radar-based occupancy maps. Background Art
[0002] Assistance and control functions in a vehicle rely on sensor detection of the vehicle's external environment. External sensors can include different types of sensor systems, such as cameras, radar, and lidar systems. The sensor data obtained by these sensor systems is then processed by an onboard computing platform known as an electronic control unit.
[0003] The vehicle control function is navigation. Modern vehicles include driver assistance systems such as front / rear approach warning and parking assistance. These functions rely on object detection in the vehicle's external environment.
[0004] This control and accessibility needs to be improved. Summary of the invention
[0005] In this context, a method, system and computer program product are presented.
[0006] More specifically, a computer-implemented method for driver assistance in a vehicle includes generating a three-dimensional occupancy grid based on radar point sensor data of the vehicle's environment. Figure 3 Based on the radar point sensor data, multiple feature grid maps FGMs are generated, where the corresponding feature dimension of each FGM corresponds to the feature of the radar point sensor data. Based on the 3D OGM and multiple FGMs, a refined OGM is generated. The refined OGM is provided for use by the vehicle's auxiliary system.
[0007] In some embodiments, the refined grid map includes at least one of a refined 3D OGM and a feature map, wherein dimensions of the feature map indicate one or more semantically classified traffic infrastructure features of the vehicle environment.
[0008] In some embodiments, the multiple FGMs include one or more of a radar cross section FGM, a radial velocity FGM, and a distance FGM, wherein a dimension of the radar cross section FGM indicates a radar cross section of a detected stationary environmental element (such as a traffic infrastructure element), a dimension of the radial velocity FGM indicates a radial velocity of the detected stationary environmental element (such as a traffic infrastructure element), and a dimension of the distance FGM indicates a distance to the detected stationary environmental element (such as a traffic infrastructure element).
[0009] In some embodiments, a refined grid map is generated using a convolutional neural network (CNN), and includes inputting the 3D OGM and multiple FGMs into the CNN.
[0010] In some embodiments, the method further includes applying, by the CNN, a two-dimensional convolution to the x and y spatial dimensions of the 3D OGM and the multiple FGMs, and treating the z dimension of the 3D OGM and the feature dimensions of the multiple FGMs as channels.
[0011] Some embodiments include repeating the result of the two-dimensional convolution along the z dimension.
[0012] In some embodiments, the method further comprises: applying two-dimensional convolution to the x and y dimensions of the 3D OGM respectively for any layer of the z dimension of the 3D OGM; and applying one-dimensional convolution to the z dimension of the 3D OGM respectively for any cell of the x and y dimensions.
[0013] In some embodiments, the method further comprises: concatenating the results of the convolutions; maximally reducing the z dimension of the concatenated results; sequentially downsampling the x and y dimensions; and sequentially upsampling the x and y dimensions.
[0014] In some embodiments, the method further includes repeating the upsampling result along the z dimension; and concatenating the repeated upsampling result with the repeated result of the convolution along the channel.
[0015] In some embodiments, the method further comprises reducing the channel dimension to 1 to output a refined 3D OGM.
[0016] In some embodiments, the method further comprises reducing the z dimension to output the feature map.
[0017] In some embodiments, reducing the z dimension to output a feature map includes: determining two cumulative maxima along the z dimension; and concatenating the two maxima with the reduced channel dimension result.
[0018] In some embodiments, the method further includes adaptively realigning the 3D OGM and the plurality of FGMs based on a current orientation of the vehicle.
[0019] In some embodiments, adaptively realigning the 3D OGM and multiple FGMs based on a current orientation of the vehicle includes: in response to determining that the current orientation of the vehicle deviates from a deviation between reference points of the 3D OGM and multiple FGMs exceeding a given threshold, realigning the 3D OGM and multiple FGMs with the current orientation of the vehicle by integer translation of the 3D OGM and multiple FGMs.
[0020] Another aspect relates to an electronic control unit (ECU) for a vehicle, the ECU being arranged to implement a method as described herein.
[0021] Another aspect relates to a vehicle equipped with a radar sensor system for collecting radar point sensor data to be provided to an ECU and the ECU is communicatively coupled to the radar sensor system.
[0022] A final aspect relates to a computer program product comprising instructions which, when executed on a computer, cause the computer to perform a method as described herein.
[0023] These and other objects, embodiments and advantages will become readily apparent to those skilled in the art from the following detailed description of the embodiments with reference to the accompanying drawings, and the present disclosure is not limited to any particular embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The foregoing and other objects, features, and advantages of the present subject matter will become apparent from the following description of exemplary embodiments with reference to the accompanying drawings, wherein like reference numerals are used to represent like elements, and wherein:
[0025] Figure 1 A high-level flow chart of the method is shown.
[0026] Figure 2 is a more detailed flowchart of the present method implemented by a convolutional neural network.
[0027] Figure 3 Further details of the implementation of 3D OGM input and processing are given.
[0028] Figure 4 Provides a more detailed understanding of the implementation of FGM input and processing.
[0029] Figure 5 Details of the pyramid treatment are explained in detail.
[0030] Figure 6 Details of downsampling to generate a two-dimensional transportation infrastructure map are shown.
[0031] Figure 7 and Figure 8 The improvements provided by this feature are presented through Intersection Over Union and Precision / Recall curves.
[0032] Fig. 9A , Fig. 9B , Fig. 9C and Fig.9D A visual example of the improvement provided by this feature is shown.
[0033] Fig.10 An example of map orientation used in a moving vehicle is shown.
[0034] Fig.11is a diagram of a computing system that implements the functionality described herein. DETAILED DESCRIPTION
[0035] Deep convolutional neural networks (CNNs) are used in the field of image analysis and perception, for example for object detection and semantic segmentation. The input to a neural network is usually a multidimensional tensor representing an image, such as a 2D image / grid with multiple channels (e.g., three channels such as RGB) or 3D voxels defined using a specific spatial coordinate system. Each image grid cell contains either a point or some pre-computed features. In the former case, the network is responsible for automatically learning the appropriate features.
[0036] The architecture of a CNN for image processing to detect features can be based on segmenting a 3D dimensional representation into different 2D projections. The segmentation can utilize one or more downsampling functions. For example, focal loss is a loss function used in classification tasks, such as as part of an object detection process. Corner pooling is used to implement an anchor-free object detector. Feature pyramids can be used for object detectors and semantic segmentation tasks.
[0037] Occupancy grid maps (OGMs) and inverse sensor models are used to model and understand the environment, for example in the fields of robotics and navigation. OGMs provide a structured way to represent and estimate the occupancy or presence of objects or obstacles in a given geographic area. For example, OGMs can be used for tasks like obstacle avoidance, path planning, and simultaneous localization and mapping (SLAM). Robots and autonomous vehicles can use these maps to understand their surroundings, plan safe paths, and avoid collisions with obstacles.
[0038] The OGM divides the environment into a grid of equal-sized cells. In 2D space, these cells are typically arranged in rows and columns, resulting in a structured grid that covers the entire area of interest. In a 3D OGM, a cell refers to an equally-sized cubic volume that discretes the mapped area. The size of the cell can vary depending on the specific use case and the level of detail desired. The size of the cell also forms the resolution of the OGM, which affects the accuracy and level of detail in the map's graphical representation. Higher-resolution grids provide more detail, but require more computing resources to process, maintain, and update.
[0039] Each grid cell in the OGM is associated with at least one probability representing the likelihood of occupancy. Typically, the probability is binary, where a single cell is considered to be occupied (1 or true) or unoccupied (0 or false). More precise probability values may also be used to indicate partial occupancy (values between 0 and 1).
[0040] OGM can be generated by fusing sensor data, such as laser rangefinders, sonar, or cameras. These sensors provide information about the environment, and the map is updated based on the sensor readings. When sensor measurements indicate the presence of an obstacle within a cell, the occupancy probability of that cell is increased. OGM can also be integrated with additional features specific to radars.
[0041] The OGM can be continuously updated as updated sensor data becomes available. Bayesian techniques such as Bayes' rule or recursive Bayesian filters (such as occupancy grid map algorithms) can be used to update the probability values in the grid cells based on current sensor measurements and previous data.
[0042] The Feature Grid Map (FGM) referred to below can be considered a variant of the OGM. As further detailed below, the FGM is divided into cells of equal size like the OGM. However, unlike the standard OGM, the FGM has two classic Cartesian dimensions (horizontal, x and y) and a third dimension that quantifies the presence of a given feature (particularly a traffic infrastructure element) in the underlying sensor data.
[0043] Such a mechanism is used here for the refinement and fusion of multi-occupancy and feature-based grid maps (OGMs), with the goal of generating grid maps with improved quality to facilitate other downstream applications in the vehicle, such as a high-quality three-dimensional (3D) environment map of the vehicle (which is typically moving). Optionally, a 2.5D semantic map of static elements of the vehicle environment can also be generated.
[0044] In an implementation example, these outputs are created by a deep artificial neural network. The spatial 3D map output of the network provides a denser, cleaner and more effective representation of the vehicle environment than the original map data (OGM or multiple OGMs). This improvement in the quality of the map representation can be used as a basis for improving downstream tasks in the vehicle, such as guardrail detection, obstacle detection that can be driven under / over, key point detection for self-localization, etc. In turn, these improved downstream tasks can promote safety in the vehicle by being applied to assistance systems and / or for autonomous driving. Therefore, the refined raster map output has a variety of practical applications for vehicle operation, safety and assistance.
[0045] Thus, a computer-implemented method for driver assistance in a vehicle is provided and described herein, the method comprising: generating a three-dimensional occupancy grid based on radar point sensor data of an environment of the vehicle Figure 3 D OGM; generating a plurality of feature grid maps FGM based on the radar point sensor data, wherein a corresponding feature dimension of each of the FGMs corresponds to a feature of the radar point sensor data; generating a refined grid map based on the 3D OGM and the plurality of FGMs; and providing the refined grid map for use by an auxiliary system of the vehicle.
[0046] Depend on Figure 1 A general visualization of the method described herein is given. The process inputs (10) radar point sensor data of the environment of a vehicle. The vehicle can be any type of movable transportation vehicle, such as a car (including a taxi), a truck, a bus, a train, a ship, an airplane, etc. The vehicle is equipped with a radar system that collects the radar point sensor data. The radar system can be implemented by one of a plurality of radar sensors mounted on the vehicle and communicatively connected to a computing device (electronic control unit) to provide sensor data. For example, if the vehicle is a car, four radar sensors can be installed, one at each corner of the car to allow 360° coverage.
[0047] Each reflection or echo received by the radar system represents a data point. These data points include information such as the time it took for the radar signal to reflect off an object and return, and the direction (azimuth and elevation) from which the reflection came. These data points are processed to detect and track objects in the radar system's field of view.
[0048] Each detected radar point is characterized by a position in 3D space, such as Cartesian x, y, z coordinates (longitudinal x, lateral y, altitude z) or spherical coordinates (range, azimuth, elevation), as well as additional features such as radar cross section (RCS) and radial velocity (Doppler effect). Since the present method is mainly focused on determining the position of static elements of the vehicle's geographical environment, point detections with absolute radial velocities above a certain threshold (indicating a moving object) (e.g., 0.3 m / s) can be discarded from the input.
[0049] Radar point sensor data generated by a radar system of the vehicle is used to generate a plurality of grid maps. More specifically, the radar point sensor data is processed to generate (11) a 3D OGM and to generate (12) a plurality of FGMs. The three dimensions (x, y, z) of the OGM constitute a Cartesian coordinate system, where x may refer to the longitudinal direction, y may refer to the lateral direction, and z may refer to the height direction. The 3D OGM may be generated using an inverse sensor model (ISM) configured to present radar point detections in a grid of the OGM according to a Gaussian distribution.
[0050] In ISM, static uncertainty in the radial direction and static uncertainty in azimuth and elevation are assumed. In Cartesian space, the uncertainty in the direction orthogonal to the radial direction is linearly proportional to the radar point detection distance from the radar sensor. Using these three uncertainties, a 3D Gaussian distribution is defined with the detection x, y, z position as the average, which determines which elements of the grid each point detection contributes to and to what degree. The total strength of each point's contribution is also scaled using specific confidence metrics, such as signal-to-noise ratio, indicator values for azimuth and elevation lookup confidence, and a bistatic flag. This information is used to derive a scaling factor for the Gaussian bell wave for each radar point.
[0051] In addition, additional grid maps are generated from other features detected by the radar point. As mentioned above, these additional grid maps are referred to herein as feature grid maps (FGMs). The FGM has two Cartesian dimensions (horizontally, x and y) and a third dimension corresponding to the features. Therefore, the FGM can be understood as a stack of several two-dimensional OGMs, each of which is calculated using only the detection of features within a specific range. Alternatively, the FGM can be viewed as a 2D OGM, where each spatial unit contains a histogram about the features. The features encoded as the third dimension of the FGM are processed as the third coordinate during the processing of the neural network. For example, a point detection in an RCS FGM is encoded by the coordinates x, y, rcs, while a point detection in a range FGM is encoded by the coordinates x, y, range.
[0052] The FGM is generated in a similar way to the 3D OGM. For the FGM, the 3D Gaussian is reduced to two dimensions. The feature dimension can follow the nearest neighbor assignment scheme.
[0053] In some embodiments, three FGMs are generated based on the three radar data point features. More specifically, in some embodiments, the multiple FGMs include one or more of a radar cross section FGM having a dimension indicating the radar cross section of the detected stationary environmental element, a radial velocity FGM having a dimension indicating the radial velocity of the detected stationary environmental element, and a distance FGM having a dimension indicating the distance to the detected stationary environmental element. The stationary environmental elements may include traffic infrastructure elements of interest represented in the map, as well as other environmental elements such as trees, houses, fences, etc.
[0054] For ease of processing, the input 3D OGM and the input FGM typically use the same grid structure and grid size, ie, they have the same resolution. Further details of an exemplary FGM used herein are discussed below.
[0055] Continue to refer to Figure 1, the 3D OGM and multiple FGMs are used to generate (13) a refined grid map. The term refined grid map can include one or more map representations. In some embodiments, the refined grid map includes at least one of the refined 3D OGM and a traffic infrastructure feature map. The traffic infrastructure feature map is a two-dimensional map with two spatial dimensions (vertical, horizontal), and the third dimension of the feature map indicates one or more traffic infrastructure features (e.g., guardrails, bridges, signs) of the vehicle environment.
[0056] Thus, for example, the refined grid map may again become a 3D OGM with improved quality, i.e. a more correct indication of occupancy in the grid. The refined OGM may also include map representations indicating features of particular relevance to vehicle navigation and / or assistance, such as traffic infrastructure elements that can be driven under (e.g., tunnels, trees, bridges across roads, gates, highway signs) or traffic infrastructure elements that can be driven over (e.g., physical objects that are above the road surface but low enough to be driven over with sufficient caution, such as speed bumps, curbs, etc.). For example, a 3-meter-high crossing barrier may be marked as a traffic infrastructure element that can be driven under for a passenger car, but not for a truck or bus.
[0057] In general, the terms "drive-under" and "drive-over" are used herein to refer to the drivability conditions of a road or path. These terms relate to the ability of a vehicle to safely drive over a traffic infrastructure or road, i.e., to safely drive under traffic infrastructure elements that restrict the drive-over height of the traffic infrastructure elements, and to safely drive over traffic infrastructure elements.
[0058] Finally, the refined grid map is provided (14) to the vehicle's functions for use by the vehicle's assistance systems. The refined OGM may be used for any assistance, safeguarding, safety or control function of the vehicle. For example, a secondary assistance system of the vehicle may warn the driver of an unpassable traffic infrastructure element in the vehicle's path based on the refined OGM. An autonomously driven vehicle may stop before an unpassable traffic infrastructure element or navigate around such an unpassable traffic infrastructure element. The improved quality of the improved output grid map may help improve the quality of such secondary assistance systems. For example, false alarms for interruption assistance systems (which activate an interruption operation without an actual reason) may be reduced.
[0059] The refined OGM may be used to display traffic infrastructure information to the driver of the vehicle via the vehicle's graphical user interface to better assist the driver in driving and navigation tasks. For example, traffic infrastructure elements that may be drivable under or drivable over may be graphically highlighted in the 2D or 3D map representation shown, while traffic infrastructure elements that may not be clearly drivable under or drivable over for a given vehicle may be graphically labeled differently to highlight any potential obstacles or hazards to the vehicle.
[0060] In some embodiments, a refined grid map is generated using a convolutional neural network (CNN), and includes inputting the 3D OGM and multiple FGMs into the CNN. Figures 2 to 6 Details of this implementation using CNN are shown. Figure 2 presents a more general overview, while Figures 3 to 6 Added some additional scheme details.
[0061] Figure 2 An implementation example utilizes a convolutional neural network (CNN). Other implementations such as alternative machine learning models or neural network types are also possible. Both the 3D OGM and multiple FGMs are input to the neural network and initially processed by the CNN (activities 20, 21). As described above, one input is a 3D Cartesian OGM with dimensions x, y, and z. The OGM is a 3D map of the static environment of the vehicle as perceived by the radar sensor, and for radar systems, especially in the z dimension (vertical), this is prone to noise and inaccuracy.
[0062] Typically, the process (20, 21) includes convolution of the input image. For example, two dimensions of the input image are convolved, and the third dimension is treated as an image channel. The outputs of the process (20, 21) are aggregated by concatenating the outputs along the image channels (22).
[0063] The output of the cascade (22) is further processed by pyramid processing (23), using multiple layers within the CNN with different spatial resolutions or receptive fields. This mechanism helps capture features present in the input image at different scales and helps determine the structure of objects in the radar input OGM / FGM regardless of the size or location of the object. Figure 5 The details of pyramid processing (23) are further explained.
[0064] The output of the pyramid processing (23) and the output of the cascade (22) are then cascaded again along the image channels (24). The output of this second cascade (24) is input to the feed-forward network (25) (FFN). For each xyz voxel, the FFN (25) gradually reduces the channel dimension of the image data to 1, thereby generating an improved 3D OGM (26). Through further downsampling activities (27), a 2D traffic infrastructure feature map (28) can also be output.
[0065] Figure 3 and Figure 4 A more detailed view of an exemplary process of inputting a 3D OGM and multiple FGMs is provided. Figure 3 and Figure 4 Together, they constitute multiple input and processing paths and stages of a CNN to efficiently aggregate information across different dimensions and input grids. In the first stage of the convolution process formed by blocks 31, 33, 34, 42, 45, 48, all input grid maps are processed using convolutions applied to their spatial dimensions. Thus, the input grid maps are first convolved.
[0066] Some embodiments further apply CNN two-dimensional convolutions to the x and y spatial dimensions of the 3D OGM and the multiple FGMs; and treat the z dimension of the 3D OGM and the feature dimensions of the multiple FGMs as channels. Some embodiments include repeating the result of the two-dimensional convolution along the z dimension. Some embodiments also apply two-dimensional convolutions to the x and y dimensions of the 3D OGM for any layer of the z dimension of the 3D OGM respectively; and apply one-dimensional convolutions to the z dimension of the 3D OGM for any unit of the x and y dimensions respectively.
[0067] More specifically, the 3D OGM (30) is processed through three paths ( Figure 3 ). The first path (31, 32) is given by a 2D convolution over the spatial x, y dimensions, treating the z dimension (depth, the vertical dimension of the 3D OGM) as an image channel. Typically, convolution (31) (and other convolutions performed by CNNs) perform a convolution operation on the input data using a set of filters (also called kernels). The applicable size of the filters (e.g. 3×3 or 5×5) depends on the implementation. The convolution consists of element-wise multiplication of the filter with a portion of the input data in the z dimension, and then summing the results to generate a single output value. The convolution operation is performed over the spatial dimensions of the input OGM (vertical x and horizontal y). The result of the convolution is repeated (32) to regain the z dimension, i.e., the x, y layers of the 2D convolution are stacked along the z dimension (Z-multiple copies) to form the output graph of this processing path.
[0068] The same convolution process is used for the input FGM ( Figure 4). The RCS FGM (41) is convolved (42) over the spatial dimensions, the range FGM (44) is convolved (45) over the spatial dimensions (longitudinal direction x, transverse direction y)(longitudinal direction x, transverse direction y), and the RV FGM (47) is convolved (48) over the spatial dimensions (longitudinal direction x, transverse direction y)(longitudinal direction x, width direction y), while the corresponding feature dimensions are considered as image channels. All three convolutions (42, 45, 48) are repeated along the z dimension (43, 36, 49).
[0069] Again, refer to 3D OGM ( Figure 3 ), the 3D input OGM is processed by two additional processing paths. On the one hand, 2D convolution (33) is performed separately for all z layers in the xy plane (repeated on z), thus naturally maintaining the z dimension. On the other hand, 1D convolution along the z dimension is applied to all xy units separately (repeated on x, y), thus preserving the x and y dimensions.
[0070] The resulting convolutional graph data is along the channel (see Figure 2 ) and FGM( Figure 4 )’s convolutional stages are cascaded (22) to form a combined intermediate representation of the input graph data.
[0071] Some embodiments concatenate the (repeated) results of convolution (repeated convolution 31, convolution 33, convolution 34, repeated convolution 42, 45, 48), maximally reduce the z dimension of the concatenated results, sequentially downsample the x and y dimensions, and sequentially upsample the x and y dimensions. Some embodiments further repeat the upsampled results along the z dimension; and concatenate the repeated upsampled results with the repeated results of the concatenation of the convolutions along the channels.
[0072] Reference Figure 5 , the output of the cascade (22) is processed by the pyramid process (23). Here, the z dimension of the intermediate representation is reduced. For example, a maximum reduction operation (50) is applied to the z dimension, which reduces the z dimension to the maximum value in the z dimension. Thus, for each spatial position (i.e., each x, y position), the maximum value in the z dimension (among all channels) at that position is selected.
[0073] Furthermore, the graph data is first sequentially downsampled (51) and then sequentially upsampled (52) back to the original size. Downsampling (51) can again be achieved through multiple convolutional layers followed by pooling layers (e.g., max pooling or strides in convolution) to gradually reduce the spatial dimensions of the graph while increasing the number of channels. This allows for capturing high-level and abstract features in the graph data.
[0074] Upsampling (52) takes the lower resolution maps from downsampling (51) and upsamples them to match the spatial dimensions of the higher resolution feature maps. This can be achieved using techniques such as transposed convolution operations, nearest neighbor, bilinear interpolation, and / or maximum demodulated sum.
[0075] Feature pyramid processing can also include further processing steps, such as combining the graphs from downsampling and upsampling. At each level of the pyramid, the graph data can represent information at a specific spatial resolution. Thus, a combined graph can be created by fusing high-level abstract features from downsampling with finer-grained details from upsampling.
[0076] The output of the pyramid processing, i.e., multiple graphs at different spatial resolutions, is repeated (53) in the z dimension and then concatenated (24) with the existing graph data output by the cascade (22) to form another intermediate representation of the graph data.
[0077] Some embodiments further reduce the channel dimension to 1 to output a refined 3D OGM. Some embodiments further reduce the z dimension to output a traffic infrastructure feature map. In some embodiments, reducing the z dimension to output the feature map includes determining two cumulative maxima along the z dimension; and concatenating the two maxima and the reduced channel dimension result.
[0078] Figure 6 Further details of downsampling (27) to generate a 2D traffic infrastructure feature map (28). A 2D traffic infrastructure feature map is generated by adopting the intermediate graph representation of the feedforward network (25) and by applying a 1D convolutional downsampling network to reduce the z dimension. Optionally, an edge pooling operation is performed at each resolution level of the downsampling (27) to facilitate detection of object boundaries along the z dimension. Edge pooling calculates two cumulative maxima along the z dimension, one in each direction, namely the cumulative maximum of the z dimension (60) and the reverse cumulative maximum of the z dimension (61). To this end, the z dimension is scanned from bottom to top (60) and from top to bottom (61) to cumulatively determine the maximum value. These two results and the original input are then cascaded along the channel dimension (62) (max pooling).
[0079] CNNs can be trained using real data, for example in the form of a LiDAR-based OGM. An OGM computed from a LiDAR point cloud can be used as a real target for network-based refinement of a LiDAR-based OGM. Each point in the cloud has previously been assigned a semantic class label of “road”, “obstacle” or “driveable over” (e.g., curb, speed bump) by a pre-existing semantic classification algorithm. This allows the computation of a 4D semantic OGM with three spatial dimensions and a fourth dimension corresponding to the semantic class. This OGM can also be interpreted as three independent 3D OGMs computed using points separated by class.
[0080] Based on this 4D LiDAR-based semantic OGM, the following data can be calculated for CNN training purposes:
[0081] 3D Labels: A 3D OGM with the road removed (the vehicle’s radar system typically cannot identify roads or similar driving paths due to the lack of objects). Therefore, the 3D OGM used as a reference for the 3D OGM CNN output characterizes the “obstacles” and “drive-over” parts of the vehicle’s environment.
[0082] 2.5D Semantic Labeling: A 2D map with three (non-exclusive) exemplary categories, corresponding to "low obstacle", "high obstacle" and "drive-over". A "low obstacle" is defined as any obstacle located 2.5 meters or less above the road surface, while a "high obstacle" is defined as any obstacle located more than 2.5 meters above the road surface, such that parts of the map shown as "high obstacles" rather than "low obstacles" can be safely driven under by passenger vehicles (such as bridges, tunnels, highway signs). For other vehicle types (trucks, buses, ships, etc.), the categories "low obstacle" and "high obstacle" can be appropriately defined based on the vehicle height.
[0083] Visibility Mask: A 3D binary mask representing the elements of the vehicle environment that are visible to the radar sensor, computed using ray tracing.
[0084] CNNs can be trained using a modified focal loss operation that addresses the class imbalance problem in object detection tasks, where the number of background (non-object) elements far exceeds the number of object samples. This imbalance leads to poor learning and degraded model performance. In traditional focal loss variants, focal weights are calculated separately for each element of the grid, where poorly classified elements are weighted higher than well-classified elements. Since these weights do not take into account the neighborhood of the element, focal loss tends to focus on single problematic elements rather than problematic regions. At object boundaries, this can lead to oscillations during training and blurred edges during inference. To alleviate these problems, this paper proposes to apply max pooling to the focal weights. Therefore, an element with a high focal weight causes a small neighborhood around the element to be emphasized in the loss, rather than just the element itself. This can improve the visual quality of the generated graphs.
[0085] An example of the quality of the output image can be obtained from Figure 7 and Figure 8 The curve graph and Fig. 9A , Fig. 9B , Fig. 9C and Fig.9D Figure and image example export.
[0086] Figure 7 is the intersection-over-union (IoU) graph, and Figure 8 is the precision / recall curve. In both figures, the solid line represents the performance of the current method introduced in this paper, where the dashed line represents the input 3D OGM. Figure 7 It is shown that for the grid map refined by convolutional neural network (CNN), its IoU value (as a function of IoU threshold for target classification, i.e. predicted voxels within a given distance threshold from the true voxel) is significantly higher than that of the input 3D OGM (30). When the classification threshold is about 0.4, the consistency of the refined output grid map with the real data is at the best working point. Figure 8 The trade-off between precision and recall is shown. Typically, precision is defined as the ratio of true positives to all positives (true positives and false positives), while recall is defined as the ratio of true positives to the sum of true positives and false negatives. Figure 8 As can be seen from the graph, the refinement mechanism described in this paper outperforms the input occupancy grid map (OGM).
[0087] Fig. 9A , Fig. 9B , Fig. 9C and Fig.9D A visual example of an improvement of the mechanism described herein is provided. Fig. 9A An example input OGM (30) is shown in a 2D bird's-eye view, and Fig. 9B The corresponding real LiDAR-based image is shown. Fig. 9C is an example of a corresponding refined grid map as output according to the present method, and Fig.9D It also provides a complementary self-vehicle perspective. Fig. 9C The picture is Fig. 9A The original input map is significantly improved as the traffic infrastructure features become more visible.
[0088] Some embodiments also adaptively realign the 3D OGM and the plurality of FGMs according to a current orientation of the vehicle. In some embodiments, adaptively realigning the 3D OGM and the plurality of FGMs according to the current orientation of the vehicle includes realigning the 3D OGM and the plurality of FGMs with the current orientation of the vehicle by integer translation of the 3D OGM and the plurality of FGMs in response to determining that the current orientation of the vehicle deviates from a deviation between reference points of the 3D OGM and the plurality of FGMs exceeding a given threshold.
[0089] An exemplary visualization of this adaptive realignment scheme is given by Fig.10The refined grid map is intended to cover a certain area in front of the ego vehicle in order to provide information on driving directions. But as the ego vehicle moves, its position and orientation relative to the grid map changes. Eventually, this will cause the vehicle to leave the grid map area or turn to one of the edges of the map area, such that the grid map no longer actually covers a sufficient portion of the area in front of the vehicle.
[0090] It is unwise to compensate for such changes in vehicle orientation by repeatedly interpolating on a grid centered on the vehicle, as interpolation artifacts will accumulate rapidly. Instead, an adaptive realignment scheme is proposed for the square grid, which can keep the refined grid map aligned with the area in front of the ego vehicle using only occasional lossless integer transforms. The goal of the realignment scheme is to keep the ego vehicle near a certain target point. This target point is located on a ring around the edge of the grid and changes dynamically based on the current vehicle orientation. When the ego vehicle turns, the target point slides along the indicated outer ring. If it is determined that the offset of the ego vehicle from the target point exceeds a given offset threshold, the input grid map (3D OGM) and multiple FGMs undergo integer translations to bring the ego vehicle and the target point as close as possible. Therefore, sharp turns of the ego vehicle result in larger and / or more frequent translations to keep the vehicle roughly oriented towards the center of the grid, while lighter turns of the ego vehicle result in less frequent translations, and very slight turns below the offset threshold do not result in any translation.
[0091] Adaptive re-centering can be implemented at the level of the network input data, i.e., by moving the 3D OGM and multiple FGMs when the offset from the target point exceeds a given threshold. More specifically, the re-centering offset occurs between updates of the input 3D OGM based on current and past versions of the grid map data derived from previous radar point sensor data. In this way, the refined grid map output by the network is indirectly and dynamically adapted by adapting the network input map data to the current orientation of the ego vehicle in the manner described above.
[0092] In practice, the refined grid map remains aligned with the current orientation of the ego vehicle within a given offset threshold, so that the ego vehicle drives towards the center of the grid. This ensures in a computationally efficient way that most of the output refined grid map grids are always located in front of the ego vehicle, even if the orientation of the grid in the coordinate system has not changed.
[0093] Fig.111 is a diagram of the internal components of a computing system (100) that implements the functionality described herein. The computing system (100) may be located in a vehicle and includes at least one processor (101), a user interface (102), a network interface (103), and a main memory (106) that communicate with each other via a bus (105). Optionally, the computing system (100) may also include a static memory (107) and a disk drive unit (not shown), which also communicate with each other via the bus (105). A video display, an alphanumeric input device, and a cursor control device may be provided as examples of the user interface (102).
[0094] In addition, the computing system (100) may also include a network interface (103) to communicate with the vehicle's onboard sensor system and other computing systems such as an electronic control unit. The sensor system is used to obtain sensor data to provide to the computing system (100) for processing.
[0095] The computing system (100) may include a graphics processing unit (GPU) (104) that may be particularly arranged to perform at least some portions of the above-described CNN operations. For example, the GPU (104) may feature a single instruction multiple thread (SIMT) architecture that is particularly suitable for performing machine learning mechanisms such as the above-described CNN-related operations. Using the GPU (104) for such operations may enable real-time processing, enabling the generation and / or updating of refined grid maps at regular intervals (e.g., at 50 milliseconds or 100 milliseconds intervals, depending on the frequency of incoming radar point sensor data and updated input maps).
[0096] The main memory (106) may be a random access memory (RAM) and / or any other volatile memory. The main memory (106) may store program code for a module (108), a module (109), and a module (110), wherein the module (108) is used to generate an input 3D OGM based on sensor data, the module (109) is used to generate multiple input FGMs based on sensor data, and the module (110) is used to generate a refined grid map (refined 3D OGM and / or 2DM traffic infrastructure feature map) based on the input 3D OGM and the input multiple FGMs. Other modules required for other functions described herein may be stored in the memory (106). The memory (106) may also store additional program data (111) for providing the functions described herein. Portions of the program data (111) and modules (108, 109, 110) may also be stored in a separate, for example, cloud storage and at least partially executed remotely.
[0097] According to one aspect, a vehicle is provided. The methods described herein may be stored as program code (108, 109, or 110) and may be at least partially contained in the vehicle. Portions of the program code (108, 109, or 110) may also be stored and executed on a cloud server to reduce the computational workload on a computing system (100) of the vehicle.
[0098] According to one aspect, a computer program comprising instructions is provided. When a computer executes the program, these instructions cause the computer to perform the methods described herein. The program code embodied in any system described herein can be distributed individually or collectively as a program product in a variety of different forms. Specifically, the program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon, and the computer-readable program instructions are used to cause a processor to perform various aspects of the embodiments described herein.
[0099] Computer-readable storage media, which are non-transitory in nature, may include volatile and nonvolatile, and removable and non-removable tangible media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media may also include random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technology, portable compact disk read-only memory (CD-ROM) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be read by a computer.
[0100] The computer-readable storage medium itself should not be interpreted as a transient signal (e.g., a radio wave or other propagating electromagnetic wave, an electromagnetic wave propagating through a transmission medium such as a waveguide, or an electrical signal transmitted through a wire). Computer-readable program instructions can be downloaded from the computer-readable storage medium to a computer, another type of programmable data processing device, or another device, or downloaded to an external computer or external storage device via a network.
[0101] It should be understood that although specific embodiments and variations are described herein, further modifications and substitutions are obvious to those skilled in the relevant art. In particular, examples are provided by illustrating principles, and a variety of specific methods and arrangements for implementing these principles are provided.
[0102] In some embodiments, the functions and / or actions specified in the flow chart, sequence diagram and / or block diagram can be reordered, processed in series and / or processed in parallel without departing from the scope of the present disclosure. In addition, any flow chart, sequence diagram and / or block diagram may include more or less blocks than those shown in the embodiments of the present disclosure.
[0103] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present disclosure. It should also be understood that when used in this specification, the terms "include" and / or "comprise" specify the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. In addition, to the extent that the terms "include", "comprise", "have", "with", "consist of" or variations thereof are used in the detailed description or claims, these terms are intended to be included in a manner similar to the term "comprise".
[0104] Although the description of various embodiments has illustrated the method, and although these embodiments have been described in considerable detail, it is not the intention of the applicant to restrict or in any way limit the scope of the appended claims to such details. Additional advantages and modifications will be apparent to those skilled in the art. Therefore, the present disclosure in its broader aspects is not limited to the specific details, representative devices and methods, and illustrative examples shown and described. Therefore, the described embodiments should be understood as provided as examples for the purpose of teaching general features and principles, and should not be understood as limiting the scope as defined in the appended claims.
Claims
1. A computer-implemented method for driver assistance in a vehicle, the method comprising: - generating a three-dimensional occupancy grid map 3D OGM based on radar point sensor data of the vehicle's environment; - generating a plurality of feature grid maps FGM based on the radar point sensor data, wherein a corresponding feature dimension of each of the FGMs corresponds to a feature of the radar point sensor data; - generating a refined OGM based on the 3D OGM and the plurality of FGMs; and - providing the refined OGM for use by auxiliary systems of the vehicle.
2. The method according to claim 1, wherein: The refined OGM includes at least one of a refined 3D OGM and a feature map, wherein a dimension of the feature map indicates one or more traffic infrastructure elements of the vehicle's environment.
3. The method according to claim 1 or 2, wherein: The plurality of FGMs include one or more of the following: - a radar cross section FGM having dimensions indicative of the radar cross section of the detected stationary environmental element; - a radial velocity FGM having a dimension indicative of the radial velocity of the detected stationary environmental element; and - A range FGM having a dimension indicating the distance to the detected stationary environmental element.
4. The method according to any one of claims 1 to 3, wherein: The refined OGM is generated using a convolutional neural network (CNN), and includes: - Inputting the 3D OGM and the multiple FGMs into the CNN.
5. The method according to claim 4 further comprises, by the CNN, - applying a two-dimensional convolution to the x-spatial dimension and the y-spatial dimension of the 3D OGM and the plurality of FGMs; and - Treating the z dimension of the 3D OGM and the feature dimensions of the multiple FGMs as channels. The method of claim 5 , further comprising repeating the result of the two-dimensional convolution along the z dimension.
7. The method according to claim 5 or 6, further comprising: - applying a two-dimensional convolution to the x and y dimensions of the 3D OGM respectively for any layer of the z dimension of the 3D OGM; and - Applying a one-dimensional convolution to the z dimension of the 3D OGM for any cell in the x and y dimensions respectively.
8. The method according to claim 7, further comprising: - concatenating the results of said convolutions; - z dimension of the maximum reduced cascade result; - downsampling the x dimension and the y dimension in sequence; as well as - Upsampling the x dimension and then the y dimension.
9. The method according to claim 8, further comprising: - repeating the upsampling result along the z dimension; as well as - Concatenate the repeated upsampling results with the repeated results of the cascade of convolutions along the channel.
10. The method according to claim 9, further comprising: - Reduce the channel dimension to 1 to output the refined 3D OGM.
11. The method according to claim 10, further comprising: - reducing the z dimension to output the feature map.
12. The method according to claim 11, wherein: Reducing the z dimension to output the feature map includes: - determining two cumulative maxima along said z dimension; and - Concatenate the two maximum values with the reduced channel dimension result.
13. The method according to any one of claims 1 to 12, further comprising: - adaptively realigning the 3D OGM and the plurality of FGMs according to the current orientation of the vehicle.
14. The method according to claim 13, wherein: Adaptively realigning the 3D OGM and the plurality of FGMs according to a current orientation of the vehicle comprises: - In response to determining that the current orientation of the vehicle deviates from the offset between the reference points of the 3D OGM and the plurality of FGMs by more than a given threshold, realigning the 3D OGM and the plurality of FGMs with the current orientation of the vehicle by integer translation of the 3D OGM and the plurality of FGMs.
15. An electronic control unit arranged to implement the method according to any one of claims 1 to 14.
16. A vehicle, comprising: - Radar systems for collecting radar point sensor data; as well as - communicatively coupled to an electronic control unit of the radar system, the electronic control unit being arranged to implement a method according to any one of claims 1 to 14.
17. A computer program product comprising instructions which, when executed on a computer, cause the computer to perform the method according to any one of claims 1 to 14.