Method and processor circuit for detecting a location solely based on a radar scan and motor vehicle equipped with a radar device

DE102024116874B3Active Publication Date: 2025-11-06CARIAD SE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102024116874
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2025-11-06
Estimated Expiration
2044-06-14

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention provides a method for detecting a location, which includes the steps of obtaining scan data X from a radar device of a vehicle and encoding the obtained scan data X of a point cloud into D respective features F. e for N' sample points X' from the point cloud, where N' <N. Um Radarrauschen und ähnliche Unzulänglichkeiten zu kompensieren, umfasst das Verfahren die Schritte des Schätzens einer Punktbedeutung F v for each intermediate sample point, where the point significance indicates the degree to which the features at the sample point originate from an existing object at that location and not from radar echoes, reflections, or sensor noise. The method can be improved by assigning a corresponding radar cross-sectional area (RCS) value, σ, to each of the N scan points. i received from the radar device and the received RCS values ​​σ i can be combined to create an RCS feature vector f rcsto generate a feature that can be combined with the features generated for the sampling points.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method and a processor circuit for detecting a location based on features derived from coded scan data. The scan data is obtained from a vehicle's radar system. The scan data describes 3D (three-dimensional) coordinates of scan points acquired by the radar system during a radar scan of the location. The invention also provides a motor vehicle comprising a radar system for generating the scan data.

[0002] The invention relates to the detection of locations that can be detected using only a vehicle radar, thus eliminating the need for GNSS (Global Navigation Satellite System) or other additional sensor data during operation. The invention focuses on the problem of obtaining a compressed yet informative representation of the scene, i.e., a suitable descriptor vector. This is achieved using a vehicle radar, such as those already integrated into conventional vehicles. These vehicle radars are popular sensors because they are compact, affordable, and resistant to adverse weather conditions. Furthermore, they provide additional signals that can aid in understanding the scene, including the Doppler velocity of the radar targets and the radar cross-section (RCS).The Doppler velocity provides an estimate of the radial velocity of the measurement and proves to be a useful indicator for dynamic object detection. The detection cross-section indicates the detectability of an object and depends on its material and viewing angle.

[0003] Radar-based location detection uses radar scans to estimate whether a current location has been visited before. Several methods utilize rotating radars for this task, employing hand-crafted features or contrastive learning. They use the intensity image provided by the rotating radar and encode it into a descriptor representing each location. Hereinafter, the term "location" refers to the place, or geocoordinates, or position (coordinates and direction / orientation) from which the radar scan was acquired. The term "scene" refers to the appearance of that location as seen by the radar during acquisition. The location can be uniquely identified by location data, such as geocoordinates or a location ID, such as a unique number or a location description (e.g.,"Home carport"), which can be entered into a digital map or digital model of the surroundings.

[0004] Methods relying solely on vehicle radar generally utilize a variant of intensity scan context (ScanContext) as an additional component of their SLAM pipeline (Kim, Giseop and Choi, Sunwook and Kim, Ayoung, “Scan context++: Structural place recognition robust to rotation and lateral variations in urban environments”, IEEE Transactions on Robotics - TRO, Vol. 38, No. 3, pp. 1856-1874, 2021; and Wang, Han and Wang, Chen and Xie, Lihua, “Intensity scan context: Coding intensity and geometry relations for loop closure detection”, IEEE International Conference on Robotics and Automation - ICRA, 2020). However, their primary focus is on developing a complete 3D radar SLAM system, with minimal emphasis on the location detection component.In particular, only the AutoPlace algorithm (Kaiwen Cai, Bing Wang, Chris Xiaoxuan Luz, “AutoPlace: Robust Place Recognition with Single-chip Automotive Radar,” in Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2022) focuses on achieving highly accurate place recognition using vehicle radar. This method limits the geometric understanding between points and restricts the features to a planar space. Furthermore, it relies on the aggregation of point clouds, which requires accurate odometry and the availability of radar submaps for place recognition. The descriptor vector resulting from AutoPlace is also high-dimensional, making it less compact for storage in a database. The algorithm also projects the point cloud of scan points into 2D images, resulting in the loss of point-specific geometric information.

[0005] Furthermore, DE 10 2023 206 446 A1 discloses a method for processing a point cloud, wherein the received point cloud data contain points with at least one feature value each and point clouds are mapped onto a grid with a plurality of grid cells, wherein each grid cell is assigned a respective sub-area of ​​the spatial area and at least one feature value is determined according to a first mapping rule depending on the points lying in the sub-area.

[0006] DE 10 2023 106 974 A1 discloses a method for detecting objects with intrinsic motion, wherein detection data from a radar sensor comprise individual points of a point cloud.

[0007] The aim of the present invention is to provide a radar-based location or position detection system for a motor vehicle that can overcome the shortcomings of automotive radar devices (noisy and sparse scans).

[0008] This problem is solved by the subject matter of the independent claims. Advantageous further developments with practical and non-trivial embodiments of the invention are described below, in the dependent claims, and in the figures.

[0009] As one solution, the invention comprises a method for detecting a place or location, wherein the method comprises the following steps, which can be performed, for example, by a processor circuit in a vehicle at the location or by a stationary backend server of that vehicle: • Obtaining scan data X from a vehicle radar device, wherein the scan data X comprises 3D (3-dimensional) coordinates of N radar scan points of a point cloud, where N is the number or count of radar points, which can be in a range from, for example, 200 to 10,000, as is known from the vehicle radar devices described in the introduction. For example, the scan data X ∈ ℝ N×3 , i.e., real numbers or floating-point numbers. • Coding the received scan data X into D, respective features F e for N' sample points X' from the point cloud, where N' <N (d.h. Unterabtastung). Die Kodierung kann auf maschinellen Lernverfahren basieren, wie sie aus dem Stand der Technik verfügbar sind.

[0010] These steps can be carried out based on the known state of the art in processing radar scan points.

[0011] To overcome the shortcomings of vehicle radar devices, the procedure includes the following steps: • Estimating the meaning of a point F v For each sampling point, the point significance indicates the extent to which the features at the sampling point originate from an object present at the location, particularly a stationary object, and not from, for example, radar echoes, reflections, or sensor noise. This estimation of the point significance can be based on an analysis of the available coordinates and / or other quantities measured by the radar, such as the relative velocity and / or radar cross-sectional area (RCS) values. A statistical method for noise and / or correlation detection, an analytical method (e.g., for geometric structure detection), and / or a machine learning method can be used. • Generating extended features F e+v by combining the features F eand their point meaning F v for each sampling point, where the "combination" can be implemented, for example, as element-wise addition or multiplication, to name a few examples. F e+v ∈ ℝ N'×D or in F e+v For each sampling point D, there are feature values, i.e., local or localized feature values. • Compressing the extended features F e+v to generate a D-dimensional descriptor vector f szene , i.e. f szene ∈ ℝ 1×DOr, in fszene, we have D feature values ​​for the entire scene, i.e., global or globalized feature values ​​that refer to the whole scene. Feature compression into a vector can be performed as is known from the prior art. For example, feature compression into a vector can be achieved by combining the corresponding feature values ​​for each of the N' sample points into a single value, resulting in D values ​​that can form the elements of the descriptor vector. A descriptor vector offers the advantage that the features are expressed "globally," i.e., independently of their position or coordinates in the scene. • pairwise comparison of the descriptor vector f szene with stored descriptor vectors, where each stored descriptor vector is associated with location data, • Signaling the location data associated with the stored descriptor vector for which the comparison satisfies a predefined match criterion. The match criterion can be designed by a person skilled in the art to ensure a minimum similarity between the compared vectors, i.e., the closest vector and also a similarity greater than a predefined threshold.

[0012] The inventive approach processes the point cloud originating from the vehicle radar point by point, without projecting it onto an image. The size of the descriptor vector f szeneThe scene size can be much smaller compared to the state of the art (D < 1000). Point significance can be expressed as a probability value or as a value in the range of 0 to 1 (where, for example, 0 means "not valuable for location estimation," i.e., most likely noise, and 1 means "valuable for location estimation," i.e., most likely an actual radar reflection event). To derive point significance, the point evaluation module assesses predefined properties of those obtained scan points, or only those of the subset of scan points, that lie within a predefined neighborhood region around the respective sample point. Such properties can include the spatial density of the sample points within the neighborhood region.In this way, incorrect sample points resulting from sensor noise (sporadic sample points without adjacent scan points) can be detected and ignored, making it possible to use a comparatively noisy radar chip in a vehicle radar device.

[0013] The step of signaling the selected location data provides a recognition result for the location or place where a vehicle may currently be. This allows at least one control unit of the vehicle to react to the signal. For example, the location data can be made available to an automated driving function of the vehicle, such as a driver assistance system and / or an autonomous driving function. The automated driving function can calculate a driving trajectory for the vehicle and control the vehicle accordingly (steer and / or accelerate). The invention also includes a navigation system for a vehicle, wherein the navigation system is configured to recognize a location according to an embodiment of the method according to the invention. The navigation system can include the processor circuitry for carrying out the method steps.

[0014] The invention also includes further developments that bring additional technical advantages.

[0015] A further advantage arises if the procedure also includes the following • Obtain, for each of the N sampling points, a corresponding radar cross-sectional area value, RCS value σ i with value index i = 1, ..., N of the radar device, • Combining the obtained RCS values ​​σ i , to obtain a D-dimensional RCS feature vector f rcs to generate. In other words, a "global" scene description is provided based on RCS values. • Generating the descriptor vector f szene includes integrating the RCS feature vector f rcs into the descriptor vector f szene In other words, when generating the descriptor vector f szene The values ​​of the RCS feature vector f will be rcsalso included or taken into account. It should be noted that "global" is meant in the sense described above.

[0016] Considering a global RCS analysis offers the advantage that the materials of objects in the scene can also be evaluated at the location. The global evaluation offers the advantage that the analysis is permutation-invariant. Furthermore, providing a D-dimensional vector as a result offers the advantage that element-wise combination, e.g., addition, with another vector of D elements is possible, derived from the extended features F. e+v This can be derived, as explained in more detail below. A further advantage arises when the N RCS values ​​σ i They can first be assigned to the individual sample points. This means that individual RCS values ​​σ iThe D dimensions or D features associated with each sample point must be mapped to them. For this purpose, a development includes: • Map each of the obtained RCS values ​​σ i into a D-dimensional RCS feature intermediate vector fσ i = MLP r (σ i ) by a multilayer perceptron, MLP, (where r stands for RCS), and • Calculating the RCS feature vector f rcs as the average vector of the N intermediate RCS feature vectors f σi .

[0017] Therefore, the RCS values ​​σ i First, they are evaluated locally, i.e., for each individual sample point, before being combined globally for the entire scene.

[0018] The method is particularly advantageous when it can be implemented in a single processing framework or architecture, so that calibrating or training this framework aligns or coordinates all processing steps. For this purpose, the method can be implemented as an artificial neural network (ANN) operated by a vehicle's processor circuit and / or a vehicle's backend server. • coding into features F e is performed by an encoder module of the ANN, where the encoder module includes, for example, a foldable neural network, • estimating the point meaning F v is performed by a scoring module of the ANN, where the scoring module is a multilayer perceptron MLP. v includes (where v stands for value or worth) • combining to form the RCS feature vector f rcsis carried out by an RCS network, RCSN, of the ANN, where the RCSN is the MLP r comprises, where the encoder module and the point-rating module are connected by a first algebraic operator to provide the extended features F e+v based on an algebraic combination of the corresponding features of F e and F v to generate and integrate the RCS feature vector f rcs into the descriptor vector f szene is performed by a second algebraic operator.

[0019] Each module can thus be implemented as a subnetwork within the ANN. Connecting the modules via algebraic operators offers the advantage that a training method, particularly the well-known backpropagation learning, can incorporate all modules into a single training session. Examples of algebraic operators include element-wise multiplication and / or addition (sum) and / or averaging and / or a normalization operation (e.g., division by the sum of the squared values). The ANN can therefore be trained using scan data from a multitude of radar scans with annotated geocoordinates specifying the location or true location data. A 2D projection of the scan points is not required; the 3D coordinates of the scan points can be directly processed into features. The integration of RCS values ​​directly into the training and evaluation / recognition process is also possible.The RCS network module, RCSN, and the point meaning estimation modules have been proven to improve location detection results.

[0020] At the heart of this development is a deep learning model that utilizes the point coordinates (scan data) and the RCS information from the radar sensor / radar device to enhance radar-based position detection capabilities. The machine learning model (the ANN) is based on a point encoder that extracts features from the original point cloud, a network that encodes the RCS values, a point evaluation module that estimates the significance of a point for position detection, and a global descriptor extractor for spatially clustering the global descriptor vector. This approach has been shown to achieve high accuracy while keeping the scene descriptor comparatively small.

[0021] This framework or architecture requires that the extended features F e+v ∈ ℝ N'×D are compressed into a D-dimensional vector. It has been shown that a particularly advantageous method is given when the extended features F e+v into an extended feature vector f VLAD ∈ ℝ D This can be implemented using a NetVLAD module that combines the first algebraic operator with the second algebraic operator. A NetVLAD module is available, for example, from Arandjelovic et al. (Arandjelovic, Relja and Gronat, Petr and Torii, Akihiko and Pajdla, Tomas and Sivic, Josef, “NetVLAD: CNN architecture for weakly supervised place recognition”, IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2016) and / or from Uy et al. (MA Uy and GH Lee, Pointnetvlad, “Deep point cloud based retrieval for large-scale place recognition”, In Proc. of the IEEE / CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2018).

[0022] A further advantage arises when the second algebraic operator is the descriptor vector f. szene generated as fscene=fVLAD+frcs‖fVLAD+frcs‖2.

[0023] This normalizes the values ​​of the descriptor vector and makes the comparison with stored descriptor vectors from other devices more independent of radar equipment.

[0024] A further advantage arises when the encoder module includes a neural network (NN) that implements 3D point convolution with L layers, where each layer contains I intermediate features from its preceding layer I-1 based on a layer-specific point convolution kernel g. l of spherical shape with radius r lderives the respective intermediate sample points of this layer I, where for input layer I = 1 the preceding layer I-1 is the scan data X, and for output layer I = L its intermediate sample points are used as the sample point X' and its intermediate features as the D features. Training using training scan data with labels indicating the true location data can lead to values ​​of g. l lead to the geometric shapes that define the geometric shapes that are characteristic or useful for distinguishing between the scenes from different locations contained in the training scan data.

[0025] A further advantage arises if, to estimate the point significance, at least one of the following properties in the respective neighborhood of the respective sample point is evaluated: • A distance measure L2 of the compared vectors must yield the smallest value of all compared vectors (current descriptor vector and stored descriptor vector) and • The distance measurement must be smaller than a predefined threshold.

[0026] This has the advantage of identifying the most similar stored descriptor vector while simultaneously guaranteeing a similarity above the specified threshold, thus ensuring the identification of an unknown location or ensuring that only location data from known locations is used.

[0027] A further advantage arises if, for the estimation of the point significance, at least one of the following properties in a respective neighborhood of the respective sample point is evaluated: • a volumetric density of the scan points or sample points, • a count of the scan points or sample points, • a geometric arrangement of the scan points or sample points (e.g. arrangement in a line, corner).

[0028] The properties can be compared to a corresponding threshold (density or number) or a template (geometric shape or arrangement). Such properties of multiple adjacent scan or sample points are highly unlikely in the case of noise, interference echoes, or reflections. The use of the described MLP v and a backpropagation training ensures that suitable properties and thresholds are identified based on labeled training scan data.

[0029] A further advantage arises when the coordinates of the sample points correspond to the coordinates of a subset of sample points selected by raster scanning. Such a downsampling method can be implemented using a prior art computer library, e.g., a voxel grid filter or voxel grid downsampling.

[0030] A further advantage arises when the scan data X for each detection is obtained from a single radar scan, with the scan data X originating from a non-rotating vehicle radar chip of the radar device. This makes location detection usable in an automated driving function of a vehicle, as the reaction time can be reduced to the duration of a single scan cycle (no aggregation of scan points is required). The radar chip can consist of a patch antenna.

[0031] A further advantage arises when the compression of the extended features generates a descriptor vector of size D with fewer than 1000 elements, preferably fewer than 512 elements. For example, D=256 has proven sufficient for fast yet reliable location detection.

[0032] A further advantage arises if the procedure also includes: entering the descriptor vector f szenea new location is added to a database of stored descriptor vectors if the matching criterion for all stored descriptor vectors remains violated. This process thus enables the learning of new locations. The location data for a new location can be based on a unique ID, such as a UUID (Universally Unique Identifier), provided, for example, by a random number generator, and / or (during a test drive) geocoordinates from a GNSS. Radar-based location detection uses radar scans to estimate whether the current location has been visited before. New locations can be added to the database of stored descriptor vectors (along with the location data). In a garage or building, or more generally in an area with no or unreliable GNSS reception, the location data can provide relative coordinates that refer to a map or a position relative to a predefined reference point, such as...Define the entrance to this area. For example, upon entering the area, the vehicle can be provided with map data of the area and a database of stored descriptor vectors linked to location data based on coordinates on the map. The vehicle can then perform radar-based navigation within the area.

[0033] For use cases or application situations that may occur in the procedure and are not explicitly described here, it may be provided that an error message and / or a request for user feedback is issued and / or a default setting and / or a predefined initial state is set according to the procedure.

[0034] As a further solution, the invention provides a processor circuit comprising instructions which, when executed by the processor circuit, cause it to perform a method according to an embodiment of the method according to the invention. The processor circuit can be configured as an electronic control unit (ECU) for a vehicle, as a stationary backend server for that vehicle, or (in a distributed implementation) as being partly located in the vehicle and partly in the backend server. The vehicle and the backend server can be connected or coupled via a data connection that includes a radio link and / or an internet connection.

[0035] The processor circuit can consist of at least one microprocessor and / or at least one microcontroller and / or at least one FPGA (Field Programmable Gate Array) and / or at least one DSP (Digital Signal Processor). In particular, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an NPU (Neural Processing Unit) can be used as the microprocessor. Furthermore, the processor circuit can contain program code that is suitable for carrying out the embodiment of the method according to the invention when executed by the processor circuit. The program code can be stored in a data memory of the processor circuit. The processor device can, for example, be based on at least one printed circuit board and / or on at least one SoC (System on Chip).

[0036] As a further solution, the invention provides a motor vehicle comprising a radar device for generating scan data correlated with received radar waves and connected to an embodiment of the processor circuit according to the invention. The motor vehicle according to the invention is preferably designed as a motor vehicle, in particular as a passenger car or truck, or as a passenger bus or motorcycle.

[0037] As a further solution, the invention also includes a computer-readable storage medium with program code which, when executed by a processor circuit, causes the processor to execute an embodiment of the method according to the invention. The storage medium can be provided at least partially as non-volatile data storage (e.g., as flash memory and / or as an SSD - Solid State Drive) and / or at least partially as volatile data storage (e.g., as RAM - Random Access Memory). The storage medium can be located in the computer or computer network. However, the storage medium can also be operated, for example, as an app store server and / or cloud server on the internet. The program code can be provided as binary code and / or assembly code and / or as source code of a programming language (e.g., C) and / or as a program script (e.g., Python). Alternatively, the computer-readable storage medium can also be implemented by a signal containing computer-readable data, e.g., via a signal.B. a time-varying voltage signal and / or a radio signal.

[0038] The invention also includes combinations of the features of the described embodiments. Thus, the invention also includes embodiments that each comprise a combination of features from several of the described embodiments, provided that the embodiments have not been described as mutually exclusive.

[0039] The following describes exemplary embodiments of the invention. The following will be shown: Fig. 1 a schematic representation of radar-based location detection; Fig. 2 a schematic diagram of an implementation of an exemplary embodiment of a method according to the invention; Fig. 3 a schematic representation of a module for estimating point significance; Fig. 4. A sketch to illustrate the test results; and Fig. 5. A sketch illustrating further test results.

[0040] The embodiments described below are preferred embodiments of the invention. In these embodiments, the described components each represent individual features of the invention, which are to be considered independently of one another and which also develop the invention independently of one another. The disclosure is therefore intended to include combinations of features of the embodiments other than those shown. Furthermore, the described embodiments can also be supplemented by other features of the invention that have already been described.

[0041] In the illustrations, identical reference symbols denote elements with the same function.

[0042] Unlike previous methods, the idea presented here introduces an approach for location detection on a single radar scan. It proposes a way to utilize radar point coordinates to capture the geometric features of the environment in a compressed yet informative descriptor. The machine learning model ANN leverages the additional RCS information provided by the sensor and compensates for radar noise and sparsity by focusing on points relevant for location detection. Our approach is capable of working with 2D and 3D radar sensors in real-world scenarios, as demonstrated by experimental evidence.

[0043] Fig.Figure 1 illustrates an example of radar location detection. Scan points from a query scan are encoded by our network and compared to the map database using the L2 distance metric. Location detection is achieved by identifying the best-matching descriptor from the database.

[0044] Fig.Figure 2 shows an example of a point meaning estimation module. (Left) During training, query points that have a correspondence within a radius δ, represented as a circle, are considered valuable for location detection. Those without a correspondence can be highly inconsistent between scans, so their predicted meaning should be reduced. Although all points are checked, we only display the radius for a small subset to achieve clarity. (Right) The resulting probabilities of our point meaning estimation module. Areas with higher density and geometric line patterns are considered more important than random points far from the sensor.

[0045] The approach or framework: The presented approach aims at location detection using a single radar scan. This involves comparing scans stored in a database with current scans acquired while navigating an environment (see Fig. 1) Initially, the radar sensors capture the environment during the robot's first pass at a location. Each scan is processed using our neural network ( Fig. 2) combines local and global point information from the measured point cloud and converts it into a location descriptor. This descriptor is then stored in a map database. During navigation, the database is queried with the current measurements and the data is processed.

[0046] However, these sensors are affected by bad weather or poor lighting conditions. In this paper, we use automotive radar to address the problem of locating a vehicle on a map using single radar scans. The effectiveness of radar is independent of environmental conditions, and it provides additional information not available in LiDAR, such as Doppler velocity and radar cross-section. However, sparse and noisy radar measurements make location detection challenging. Recent research in automotive radar addresses the sensor's limitations by combining multiple radar scans and using high-dimensional scene representations.

[0047] In contrast, an artificial neural network (ANN) architecture is proposed here that focuses on each point of individual radar scans without relying on additional odometry input for scan aggregation. We extract local and global features point by point, resulting in a compact descriptor vector of the scene. The machine learning model (i.e., the ANN) enhances local feature extraction by estimating the significance of each point for location detection and improves the global descriptor by leveraging the radar cross-sectional information provided by the sensor. To demonstrate the method, publicly available databases of labeled radar scan data, "nuScenes" and "4DRadarDataset," which include 2D and 3D automotive radar sensors (H. Caesar and V. Bankiti, AH Lang and S. Vora, VE Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O.Beijbom, “nuScenes: A multimodal dataset for autonomous driving”, arXiv, 2019, arXiv:1903.11027, http: / / arxiv.org / pdf / 1903.11027 and Li, Xingyi and Zhang, Han and Chen, Weidong, “4d radar-based pose graph slam with ego-velocity pre-integration factor”, IEEE, RAL, volume 8, number 8, pages 5124-5131, 2023). The results show that the approach achieves state-of-the-art results for location detection in a single scan using vehicle radars.

[0048] Due to noise, it can be difficult to distinguish between direct measurements of objects and those resulting from disturbances or reflections. This difficulty is further compounded by the small number of points per scan, as it can be challenging to identify objects and structural elements in the environment.

[0049] The main contribution of this work is a novel solution for location detection using single scans from vehicle radars, without relying on additional odometry input. Our focus is on a deep learning model that addresses sparsity through point-by-point processing of the scan. It utilizes the point coordinates and the RCS information provided by the sensor to enhance radar-based location detection capabilities. Our model is based on a point encoder that extracts features from the original point clouds, a network that encodes the RCS values, a point evaluation module that estimates the significance of a point for location detection, and a global descriptor extractor for spatially clustering the global descriptor vectors. Our method achieves the best results for 2D and 3D location detection in vehicle radar with a comparatively small scene descriptor.

[0050] The idea presented here introduces a novel approach to mileage-free, single-scan location detection using only vehicle radars. It proposes a method for utilizing radar point coordinates to capture the geometric features of the environment in a compressed yet informative descriptor. Our model leverages the additional RCS information provided by the sensor and addresses radar noise and sparsity by focusing on points critical for location detection. Our approach is capable of operating with both 2D and 3D radar sensors in real-world scenarios, as demonstrated by experimental evidence.

[0051] The presented approach aims for mileage-free, single-scan radar location detection. This is achieved by comparing scans stored in a database with current scans during navigation, see [reference]. Fig.1. Initially, the radar sensors capture the environment during the robot's first pass at a location. Using our neural network ( Fig. 2) Combining local and global point information from the measured point cloud, we convert each scan into a location descriptor. We then store the descriptor in a map database. During navigation, we query the database with current measurements, which use the same encoder-descriptor network as in Fig.The two points were converted into descriptors. This allows us to identify locations with high similarity according to our scoring function. Our approach uses the Doppler velocity provided by the radar to filter out dynamic point outliers. We use point folding to capture local information, along with a scoring method to estimate each point's contribution to location detection. Additionally, we incorporate global point information within the scan using RCS data and a global NetVLAD descriptor.

[0052] Dynamic point pre-filtering. A crucial advantage of vehicle radars is the measurement of the Doppler velocity of point targets. This represents the relative radial velocity of the measured point with respect to the vehicle. This velocity information cannot be used as an additional channel for the position detection network, as a vehicle can approach the same location at different speeds. Nevertheless, it allows the differentiation between static and dynamic points (Zeller, Matthias and Behley, Jens and Heidingsfeld, Michael and Stachniss, Cyrill, “Gaussian radar transformer for semantic segmentation in noisy radar data”, IEEE, Robotics and Automation Letters - RAL, Vol. 8, No. 1, pp. 344–351, 2022). We follow the Autoplace algorithm by Cai et al.(Cai, Kaiwen and Wang, Bing and Lu, Chris Xiaoxuan, “Autoplace: Robust place recognition with single-chip automotive radar”, IEEE International Conference on Robotics and Automation - ICRA, 2022) and focus exclusively on the static points of the radar scan for location recognition. The basic idea is that static points in the environment should correspond to the speed of the ego vehicle. Points with a different speed likely correspond to moving objects and are therefore considered outliers. We preprocess all scans, resulting in filtered point clouds that contain only points in the scene that are likely to be static.

[0053] Scan encoder. The main goal of the scan encoder is to obtain a feature space representation F. e ∈ ℝ N'×D of the filtered static radar point cloud X ∈ ℝ N×3In the case of 2D radar systems, we assume that the z dimension is zero. The encoding should contain spatial data sufficiently informative for place recognition, which involves capturing contextual information from points at various scales. Autoplace achieves this by projecting its radar point cloud onto a 2D image and encoding it with a convolutional neural network. As a result, most of the vertical and geometric information is lost. We focus on capturing 3D contextual information for individual points directly from the radar scan using rigid kernel point convolutions, KPConv (Wiesmann, Louis and Nunes, Lucas and Behley, Jens and Stachniss, Cyrill, “KPPR: Exploiting momentum contrast for point cloud-based place recognition”, IEEE RAL, Vol. 8, No. 2, pp. 592–599, 2022) or another point convolution algorithm, such as the one known as PointConv. For one point xil from the point cloud XL=[x1l,x2l,…,xnl]T in the plane l the folding of the features F l-1 ∈ ℝ Nl-1×Dl-1 with the kernel g l is given as: (Fl−1∗gl)(xil)=∑xkl∈Nl(xl)g(xkl−xil)lfkl−1, where Nl(xil)={xkl|xkl−xil <rl} the neighboring points of xil within a radius r l ∈ ℝ. Example values ​​for N l and r l can be seen, for example, in the implementation example below.

[0054] Our encoder consists of a sequence of five convolutional layers and downsampling layers that capture local features F l We capture data at different levels with varying radii. We use raster downsampling at different scales and find that downsampling before convolution calculations aids in processing contextual information about the points. We also increase the radius of Nl(xil) in each new convolutional downsampling block for an extended receptive field. We are also testing adding an extra channel to the input point cloud that takes into account the RCS information of each point.

[0055] Example of the implementation of the point importance estimator: Due to the small number of outliers caused by noise and multipath reflections in radar point clouds, it is crucial that the position detection network is able to identify useful anchor points within the scan. Since we know that noise can appear and disappear randomly between individual scans, we propose a point importance estimator module to guide the network training and focus on points relevant for position detection.

[0056] Fig.Figure 3 shows a comparison of the recognition results at each location of the nuScenes for individual radar scans for (left) the ANN as presented in this disclosure, and (right) Autoplace without temporal coding.

[0057] By identifying valuable measurements, this module also helps to focus on the points we know with greater certainty to exist, e.g., high-density locations and patterns in the environment, as in Fig. 3. We achieve this by adding an additional function to our network that outputs the probability that a point is important for location detection.

[0058] Using a subsampling scan X'=[x1',x2',…,xn']T∈ℝN'×3 To estimate the significance of the corresponding coded local features, we code the probability that a point xj'∈ℝ3 is valuable, in a feature vector f vj ∈ ℝD using an MLP with three learning layers, ReLU activations and layer normalization, so that fvj=MLPv(xj').

[0059] The coded probabilities of all sampled points are then converted to F v ∈ ℝ N'×D stacked and added to the output of the embedding, resulting in the following: Fe+v=Fe+Fv, where F e+v ∈ ℝ N'×D Instead of multiplying the probabilities as weights, we add them to the normalized features from the encoder to improve the feature embeddings, which correspond to the points that are more valuable for location detection.

[0060] The final importance probability of a point lies in the range [0,1]. It indicates how valuable the point is for location determination. It is estimated as P(xj')=sigmoid(ϕ(fvj)), where ϕ is a linear projection of the feature vector of ℝ D on R. We use this evaluation later in (9) to calculate the loss during training between points that are present in only one of the scans being compared. A linear projection can be implemented as a module based on an artificial neural network, which can be part of the training.

[0061] Example of radar cross-sectional network (RCS) implementation: To maximize the potential of radar sensors, we propose an additional module that uses RCS information to enhance the final feature descriptor. RCS measures the radar's detectability of an object based on the object's properties and the measurement angle. It provides additional information for each point in X about the properties of a given location, making it a valuable attribute for location detection. However, our experiments show that adding RCS as an additional channel only marginally improves location detection performance.

[0062] Instead, we propose an RCS network module that displays the RCS feature representation of the entire radar scan. RCS ∈ ℝ DIt is learned and permutation-invariant. This improves the global descriptor vector and makes it more informative than a descriptor containing only point information.

[0063] To capture the RCS information σ ∈ ℝ N Based on the scan, we suggest using a two-layer MLP that achieves the RCS value σ. i each point into a feature coding f σi ∈ ℝ D so that fσi=MLPr(σi).

[0064] Subsequently, the features of all points are aggregated and normalized over the feature dimension so that the result is permutation-invariant. fRCS=∑i=1Nfσi∑i=1Nfσi2.

[0065] The result f RCSIt acts as a global descriptor containing the distribution of RCS values ​​for this specific scan. While Autoplace uses an additional stage that performs a new histogram ranking after network prediction, we integrate our RCS network module into the model, thus avoiding this additional post-processing step.

[0066] Global descriptor database. The descriptor extraction layer combines local features into a single global descriptor vector. We use a NetVLAD layer to aggregate the local features F. e+v from (3) to K learning cluster centers, which leads to f VLAD ∈ ℝ D This leads to the following: These learnable centers represent points calculated from groups of similar local descriptors. We combine the VLAD descriptor with the RCS descriptor to obtain the final global descriptor vector f. szene after fscene=fVLAD+frcs‖fVLAD+frcs‖2.

[0067] The resulting descriptor vector represents the location measured by a radar scan. It contains the local information from the point encoder and the point meaning estimator, as well as the global features extracted using the RCS network and NetVLAD.

[0068] During map capture, each scan is recorded as a different descriptor vector. fszenem∈ℝD in a KDTree (JL Bentley, “Multidimensional Binary Search Trees Used for Associative Searching”, Communications of the ACM, Volume 18, Number 9, Pages 509-517, Year 1975) map database M={fscene1m,fscene2m,...,fsceneMm}. The query scan is also known as a feature descriptor vector. fszeneq coded and compared with those that were in M The comparison was made using the L2 distance metric. We use the superscript "m" to indicate that the feature vector belongs to the map database.

[0069] Example of implementing metric learning for location detection: The goal of the training is for the network to learn a useful and concise representation of the environment. For a query descriptor fszeneq "The feature descriptor vector must be similar for locations that are the same, called positive patterns" fszenepos and dissimilar for places that are very different, which is referred to as a negative pattern. fszeneneg. Positive and negative samples are defined during training based on GNSS distance information. However, GNSS is no longer required for the operation of our location detection network during inference.

[0070] Positive samples are defined as those located within a radius R1 of the query measured by the GNSS site. Furthermore, positive samples acquired from a nearby site may appear quite different from the queries themselves due to occlusion or different viewpoints. Therefore, during training, the positive samples within R1 that exhibit the smallest L2 distance between the descriptors are selected. We also observe performance improvements when training with multiple positive samples for the same location, as the network learns that one and the same location can be measured in different ways and from different viewpoints.

[0071] Negative samples are those located farther from a larger radius R2 such that R2 > R1. However, since the dataset may contain highly diverse scans from different locations, randomly selecting negative samples can lead to poor discrimination and generalization capabilities. As suggested by Uy et al. (Uy, Mikaela Angelina and Lee, Gim Hee, “Pointnetvlad: Deep point cloud-based retrieval for large-scale place recognition”, IEEE / CVF Computer Vision and Pattern Recognition Conference - CVPR, 2018), we employ a hard negative mining strategy to find the most similar negative pattern to the query descriptor. fszeneq to be found in the feature space. This allows the network to learn to distinguish difficult scenes where a false database match is very similar to the query scan.

[0072] We use the triplet margin loss from Balntas et al. (Balntas, Vassileios and Riba, Edgar and Ponsa, Daniel and Mikolajczyk, Krystian, “Learning local feature descriptors with triplets and shallow convolutional neural networks”, The British Machine Vision Conference - BMVC, 2016), which minimizes the distance of the query to a positive sample. dpos=L2(fszeneq, fszenepos) and the distance in relation to the hardest H negative samples dnegh=L2(ƒszeneq, ƒszenenegh) maximized. The triplet margin loss for a constant margin α is given as: Lt=∑h=1Hmax(dpos−dnegh+α, 0).

[0073] To account for the estimated value of the points from the point meaning module, we also propose measuring the nearest point correspondences between the query and the associated positive scan. First, we align both scans by transforming them into global GNSS coordinates. Then, as described in Fig. As shown, a hyperparameter for the radius difference δ is introduced, which determines whether a specific point in the query scan has a corresponding point in the positive scan. We only consider points that exhibit consistency, i.e., are present in both scans, as valuable for position detection. Note that the GNSS signal is only required for calculating the loss during training, but is dispensable during operation.

[0074] If we interpret this as a binary classification problem, we assign a label to the points in the query scan that have a match in the positive scan. P^(xj;)=1 to, and the points without correspondence P^(xj;)=0 For the sake of simplicity Pj=P(xj;) The final binary cross-entropy loss is: Lν=−∑j=1N'Pjlog(P^j)+(1−Pj)log(1−P^j).

[0075] The final loss is a weighted sum of the triplet margin loss in (8) and the point loss in (9) L=Lt+γLν, where γ is an adjustable parameter.

[0076] An experimental evaluation: The aim of the experiment is mileage-free single-scan radar position detection using vehicle radars. The capabilities of our system become clear. The results of the following evaluation support the main statements that the method (i) achieves state-of-the-art single-scan position detection on vehicle radar while maintaining a compact scene representation, (ii) provides a novel method for using RCS information to describe the scene, which improves accuracy, and (iii) improves feature extraction by estimating the significance of points within the scan for position detection.

[0077] Experimental setup: The goal of our approach is to reliably determine the position on a given map based on a single radar scan. To evaluate our method, we conduct experiments in real-world driving scenarios using 2D and 3D radar datasets, nuScenes and the 4DRadar dataset, respectively. These datasets contain the point cloud output provided by the radar sensors, so the result of our point-based method is independent of the key point extraction algorithm required for point-based approaches in NavTech radar datasets (Dan Barnes and Matthew Gadd and Paul Murcutt and Paul Newman and Ingmar Posner, “The Oxford Radar RobotCar Dataset: A Radar Extension to the Oxford RobotCar Dataset”, ICRA, 2020, and G. Kim and YS Park and Y. Cho and J. Jeong and A. Kim, “Mulran: Multimodal range dataset for urban place recognition”, ICRA, 2020).First, we evaluate our work on single-scan radar location detection and compare it with state-of-the-art solutions. We provide quantitative and qualitative results for comparison. Subsequently, we conduct an ablation study of our system to analyze the contribution of each module to the final result and to determine how the new hyperparameters affect location detection. Table 1: Comparison with the state of the art in nuScenes. R@1 % [%] R@1 / 5 / 10 [%] Output dimming. KPPR 88.0 66.1 / 78.1 / 81.6 256 RadVLAD 33.5 1.74 / 6.09 / 10.5 32768 RadVLAD RCS 17.2 0.97 / 3.19 / 4.84 32768 FFT-RadVLAD 16.3 0.29 / 1.74 / 3.29 32768 FFT-RadVLADRCS 11.2 0.19 / 0.58 / 1.64 32768 ScanContext 24.7 15.3 / 20.8 / 22.2 1200 ScanContext RCS 6.38 3.58 / 4.84 / 5.22 1200 Auto+TE7 Run 85.9 78.9 / 83.1 / 84.3 4096 Our 7 sweep 87.5 78.7 / 82.7 / 83.3 256 Car 1 run 83.0 60.8 / 72.8 / 76.5 9216 Auto+TE1 run 86.4 73.4 / 81.2 / 83.0 4096 Our 1 sweep 88.2 76.2 / 82.3 / 83.9 256

[0078] Details of the example implementation: An example of the ANN network architecture is in Fig.Figure 2 is shown. We train the model for 80 epochs with a stack size of 4. Following KPPR (Wiesmann, Louis and Nunes, Lucas and Behley, Jens and Stachniss, Cyrill, “KPPR: Exploiting momentum contrast for point cloud-based place recognition”, IEEE RAL, Vol. 8, No. 2, pp. 592-599, 2022), we set the descriptor size D = 256. We use the Adam optimizer (DP Kingma and J. Ba, title Adam: “A Method for Stochastic Optimization”, ICLR, 2015) with a learning rate of 5 × 10 -6 and a learning rate decay of 0.95 every 20 steps. We set the parameter for the triplet margin loss α = 0.1. We set the KPConv radius in each layer to r = [3,6,12,24,48] m and a raster downsampling of size 0.25,0.5,1,2,4 m. The kernel size is set to 5 for nuScenes and 7 for the 4DRadar dataset. The number of cluster centers K for NetVLAD is set to 64.

[0079] NuScenes provides insights into how our approach performs on 2D automotive radars over extended periods. We follow a similar procedure to Cai et al. (Autoplace), training with scans of the "Boston Seaport" location during the first 105 days. We divide our test sets at a 4:1 ratio, using scans taken after the 105-day threshold. We consider a prediction correct if it falls within R1 = 9 m of the actual location. During training, negative samples are taken from outside an R2 = 18 m radius. We also train with 5 positive samples for each location to improve performance.

[0080] The 4DRadar dataset demonstrates the performance of our method for location detection in loop closures using 3D radar systems for motor vehicles, as the images were acquired within a short timeframe with minimal environmental variations. The dataset is divided into four sequences. We use "Campus 1" and "Campus 4" for training and "Campus 2" and "Campus 3" for evaluation. Furthermore, we consider a prediction correct if it lies within R1 = 2 m. During training, negative samples are taken from outside a radius of R2 = 4 m.

[0081] State-of-the-art comparison: The first experiment evaluates the accuracy and resulting descriptor size of our single-scan radar position detection system compared to other methods in 2D and 3D radar datasets. Similar to other work, we denote recognition as "R" and measure the metrics R@1 / 5 / 10 and R@1%, which indicate whether the query selected by the system is among the top 1%, 5%, 10%, or 1% of candidates in the database. We also show the dimensionality of the output descriptor, which describes the size of the descriptor stored in the database. A descriptor with higher dimensionality requires more storage space, while compressed descriptors require less space and can be stored more efficiently.

[0082] Fig.Figure 4 shows the results of a qualitative location prediction in nuScenes Boston Seaport. The first image shows the location where the query measurement was performed and the predictions. The other images show the corresponding radar scans for the predicted location of the compared methods for the same query. A green frame indicates correct predictions, while a red frame indicates incorrect predictions. The results of the method presented here are labeled "Method".

[0083] Fig. Figure 5 shows a comparison of the location detection at each location of the nuScenes for individual radar scans for (left) the network (ANN) as presented in this revelation, and (right) Autoplace without temporal coding.

[0084] The comparison of our method with the state of the art is shown in Table 1 and Table 2; the qualitative results are in Fig. 4 and Fig.5. We evaluate them against the feature-based LiDAR methods ScanContext (Wang, Han and Wang, Chen and Xie, Lihua, "Intensity scan context: Coding intensity and geometry relations for loop closure detection", ICRA, 2020) and RCS ScanContext (Mulran), a learning-based LiDAR method KPPR, a frequency-based scan radar approach FFT-RadVLAD (Gadd, Matthew and Newman, Paul, "Open-RadVLAD: Fast and Robust Radar Place Recognition", RACONF, 2024), and a cluster-based approach RadVLAD (Open-RadVLAD) with and without RCS values, as well as a learning-based automotive radar solution Autoplace (Auto). We place particular emphasis on Autoplace. The original implementation uses 7 aggregated radar point clouds, and its LSTM-based temporal coding (TE) considers 3 consecutive aggregated scans. In addition to the results in their paper, we show their results for 1 Sweep with and without temporal coding in nuScenes for comparison.The best results for 1 and 7 sweeps are shown in bold and underlined.

[0085] The presented method achieves top results for both datasets, comparable to multi-scan approaches, and uses a more compact scene descriptor. In nuScenes, the small number of points per radar scan makes it difficult for non-learning-based methods to extract useful patterns from the environment. The discrete nature of the projected point clouds also poses a challenge for spinning radar frequency-based approaches such as FFT-RadVLAD. For the Radar4D dataset, our method achieves a similar R@1 with a single scan as Autoplace with TE and 3 scans. The experiment demonstrates that our compact descriptor remains highly informative for position detection. Table 2: Comparison with the state of the art in the 4D radar dataset. R@1 % [%] R@1 / 5 / 10 [%] Output dimming. KPPR 99.4 95.4 / 99.2 / 99.2 256 RadVLAD 99.3 91.9 / 96.3 / 97.3 32768 RadVLAD RCS 99.3 94.6 / 97.4 / 98.2 32768 FFT-RadVLAD 50.4 20.0 / 26.5 / 30.8 32768 FFT-RadVLAD RCS 57.3 21.5 / 28.8 / 34.2 32768 ScanContext 87.7 79.7 / 86.6 / 87.4 1200 ScanContext RCS 94.7 94.0 / 94.7 / 94.7 1200 car 99.9 94.4 / 99.4 / 99.8 9216 Auto+TE* 99.5 97.7 / 99.0 / 99.1 4096 Our 100.0 97.1 / 99.6 / 99.8 256 TE*: Temporal multiscan coding, introduced in Autoplace.

[0086] Ablation studies: The second set of experiments supports our claim that integrating RCS into the network and estimating the importance of each point leads to improved location accuracy. We conduct the experiments using the nuScenes test set. In Table 3, we evaluate how each component contributes to the final result and how it affects runtime performance during inference.

[0087] The most important modules are the scan encoder, which accepts only point coordinates (x, y, z) as inputs, or point coordinates with an additional channel for the RCS (x, y, z, RCS); the proposed RCS network (RCSN); and the point meaning module (PIM). We also experiment with multiplication ⊗ and addition ⊕ of the features in (3). Furthermore, we have tested our training strategy, which uses five positive samples for each query. We also show how scan aggregation affects runtime performance.

[0088] We can observe how much each component contributes to the final result, with the greatest effect being caused by the RCS network module. We also observe that addition is preferred over multiplication by the PIM module, as it increases the useful points for location detection without affecting the rest of the point cloud. Furthermore, the runtime is only minimally affected by the implementation of the RCS network and PIM modules (< 20 ms), but increases significantly with scan aggregation. Adding the RCS as an additional channel leads to a slight performance improvement compared to the RCS network module alone. However, this significantly increases the number of parameters per layer, resulting in a runtime increase of almost 30%.

[0089] In Table 4, we test the influence of the distance parameter δ and the weight of the loss function γ introduced by our point meaning estimation module. Large radii δ lead to false correspondences, while small radii result in no correspondence being found, leading to an equal weighting of all points. This demonstrates how performance can be improved by focusing on the points important for location detection and less on noise points. Additionally, γ varies the influence of the point meaning estimation module and how the multi-target loss from (10) affects the final result. Table 3: Ablation studies of the network modules on nuScenes. Transmitter RCSN PIM 5 Pos. Aggr. R@1 / 5 [%] Runtime [ms] (x, y, z) ⊕ 77.8 / 81.9 269 (x, y, z) 62.7 / 76.8 150 (x, y, z, RCS) 63.2 / 78.0 172 (x, y, z) 70.7 / 80.8 159 (x, y, z) ⊕ 63.1 / 76.5 159 (x, y, z) ⊗ 70.9 / 81.0 168 (x, y, z) ⊕ 73.3 / 82.2 167 (x, y, z) ⊕ 74.3 / 81.6 167 (x, y, z, RCS) ⊕ 74.8 / 82.3 229

[0090] Scan encoder (Encoder) that uses only point coordinate inputs (x, y, z) and with RCS as an additional channel (x, y, z, RCS) our RCS network (RCSN), our point meaning module (PIM), our training strategy with five positive samples per query (5 Pos) and the aggregation of seven scans (Aggr.).

[0091] Conclusion: The presented idea achieves location detection using single scans from standard automotive radar sensors, without relying on additional GNSS or odometers. We have proposed a novel point-based neural network architecture that encodes local and global scene information into a single compressed descriptor. We achieve this by encoding the local information using point convolution and combining it with the scene's RCS data. Furthermore, we integrate an additional module for estimating point importance, which helps the network learn which measurements are useful for location detection. We have implemented and evaluated our approach on various datasets and made comparisons with other existing techniques, all of which support the claims in this report.The experiments indicate that our method achieves high performance in estimating the global location of a vehicle within a map using individual vehicle radar scans, while maintaining a compact scene representation. Table 4: Ablation studies for nuScenes for radius δ and loss weighting factor y from our point evaluation module. δ [m] R@1 / 5 / 10 [%] γ R@1 / 5 / 10 [%] 1.0 71.4 / 81.1 / 83.6 0.00 71.0 / 80.0 / 82.6 2.0 71.5 / 80.2 / 82.2 0.10 71.6 / 80.6 / 83.1 3.0 71.7 / 80.6 / 83.0 0.50 73.3 / 82.2 / 84.6 4.0 70.8 / 80.7 / 82.5 1.00 71.4 / 81.1 / 83.6

[0092] It should be noted that the example implementations mentioned above can be varied by replacing one, several, or all of the described modules with comparable methods, which may also be based on machine learning or an analytical method, such as an expert system, or on statistical analysis, such as a Gaussian kernel model or lookup tables. Likewise, each module can have more or fewer layers than the number specified here. Similarly, the parameter values ​​can be varied to adapt the method to a specific use case.

[0093] Overall, the examples show how radar-based location detection can be performed with a single scan.

Claims

[1] Methods for detecting a location, comprising: • Receiving sample data X from a vehicle radar device, wherein the sample data X comprises 3D coordinates of N radar scan points of a point cloud, • Encoding the received scan data X into D corresponding features F e for N' sample points X' from the point cloud, where N' <N, characterized by • Estimating the meaning of a point F v for each sample point, where the point meaning indicates the degree to which the features at the sample point originate from an existing object at the location and not from interference echoes, reflections or sensor noise of the radar device, • Generating extended features F e+v by combining the features F e and their point meaning F v for each sample point • Compressing the extended features F e+v to generate a D-dimensional descriptor vector f szene , • pairwise comparison of the descriptor vector f szene with stored descriptor vectors, where each stored descriptor vector is linked to location data, • Signaling the location data associated with the stored descriptor vector for which the comparison meets a predefined match criterion. [2] The method of claim 1, further comprising • for each of the N sampling points, a corresponding radar cross-sectional area value, RCS value σ i with i = 1, ...,N from the radar device, • Combining the obtained RCS values ​​σ i to generate a D-dimensional RCS feature vector f rcs , where • generating the descriptor vector f szene the integration of the RCS feature vector f rcs into the descriptor vector f szene includes. [3] Method according to claim 2, wherein the combination of the obtained RCS values ​​σ ito generate the RCS feature vector f rcs includes: • Map each of the obtained RCS values ​​σ i into a D-dimensional RCS feature intermediate vector f σi = MLP r (σ i ) by a multilayer perceptron, MLP, (where r stands for RCS) • Calculating the RCS feature vector f rcs as the average vector of the N intermediate RCS feature vectors f σi . [4] Method according to claim 2 or 3, wherein an artificial neural network, ANN, is operated by a processor circuit of a vehicle and / or a backend server of that vehicle, and • the coding into features F e is performed by an encoder module of the ANN, • the estimation of the point meaning F v is carried out by a scoring module of the ANN, where the scoring module is a multilayer perceptron MLP. vincludes (where v stands for valence) • the combination to form the RCS feature vector f rcs is carried out by an RCS network, RCSN, of the ANN, where the RCSN is the MLP r comprises, where the encoder module and the point evaluation module are combined by a first algebraic operator to generate the extended features F e+v are linked on the basis of an algebraic combination of the corresponding features and the integration of the RCS feature vector f rcs into the descriptor vector f szene is performed by a second algebraic operator. [5] Method according to claim 4, wherein the extended features F e+v into an extended feature vector f VLAD This can be represented by a NetVLAD module that links the first algebraic operator with the second algebraic operator. [6] Method according to claim 5, wherein the second algebraic operator is the descriptor vector fszene generated as ƒscene=ƒVLAD+ƒrcs‖ƒVLAD+ƒrcs‖2. [7] Method according to any one of claims 4 to 6, wherein the encoder module comprises a neural network, NN, implementing 3D point folding with L layers, each layer I intermediate features from its preceding layer I-1 based on a layer-specific point folding kernel g l of spherical shape with radius r l to derive the respective intermediate sample points of this layer I, where for the input layer I = 1 the preceding layer I-1 are the scan data X and for the output layer I = L the intermediate sample points are used as the scan points X' and their intermediate features as the D features. [8] Method according to any of the preceding claims, wherein the conformity criterion includes the conditions that • A distance measure L2 of the compared vectors must yield the smallest value of all compared vectors and / or • The distance measurement must be smaller than a predefined threshold. [9] Method according to any of the preceding claims, wherein at least one of the following properties in a respective neighborhood of the respective sample point is evaluated to estimate the point significance: • a volumetric density of the scan points or sample points, • a count of the scan points or sample points, • a geometric arrangement of the scan points or sample points. [10] Method according to any of the preceding claims, wherein the coordinates of the sampling points correspond to the coordinates of a subset of the sampling points selected by raster scanning. [11] Method according to one of the preceding claims, wherein the sampling data X are taken from only one radar scan for each detection, wherein the sampling data X are obtained from a non-rotating vehicle radar chip of the radar device. [12] Method according to any of the preceding claims, wherein the compression of the extended features produces a descriptor vector of size D of less than 1000 elements, preferably less than 512 elements. [13] Method according to any one of the preceding claims, further comprising: Entering the descriptor vector f szene as a new location in a database of stored descriptor vectors if the match criterion remains violated for all stored descriptor vectors. [14] Processor circuit comprising instructions which, when executed by the processor circuit, cause the processor circuit to perform a method according to any of the preceding claims. [15] Navigation system for a vehicle, wherein the navigation system is adapted to recognize a location according to the method according to any one of claims 1 to 13. [16] Motor vehicle with a radar device for generating sampling data that is correlated with received radar waves and that is connected to a processor circuit according to claim 14.

Citation Information

Patent Citations

  • METHOD FOR DETECTING OBJECTS WITH INITIAL MOTION

    DE102023106974A1

  • Method for processing a point cloud

    DE102023206446A1