Feature extractor, method for using and for training a feature extractor, and driver assistance system
Patent Information
- Application Number
- CN202480088449.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-21
- Filing Date
- 2024-12-17
- Publication Date
- 2026-09-29
AI Technical Summary
然而,由此丢失在谱中包含的其他信息
[0015]然后,在后续进程中,将所述特征向量作为附加信息用于对象探测或路线规划或其他的下游处理过程。
Smart Images

Figure CN122847714A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a feature extractor, a method for using a corresponding feature extractor, and a method for training a corresponding feature extractor. Furthermore, this invention also relates to a driver assistance system for a vehicle equipped with at least one radar sensor, utilizing a feature extractor. Background Technology
[0002] Most radar sensors provide a point cloud as their output signal, consisting of detected reflections. However, internally, the radar sensor receives the spectrum of electromagnetic radar waves and derives the point cloud from these waves. For example, this involves searching for maxima in the spectrum, which represent reflections at locations where objects actually exist with high probability.
[0003] By providing a preprocessed point cloud, the amount of data to be transmitted can be kept small. However, this results in the loss of other information contained in the spectrum.
[0004] DE 10 2016 002 208 A1 describes a method for analyzing and evaluating radar data from at least one radar sensor in a motor vehicle, as well as a motor vehicle.
[0005] DE 10 2019 219 144 A1 describes a method and system for image analysis evaluation of object recognition in a vehicle environment. Summary of the Invention
[0006] Against this backdrop, the present application proposes a feature extractor according to the independent claims, a method for using the corresponding feature extractor, and a method for training the corresponding feature extractor. Advantageous extensions and improvements to the proposed solution are given in the specification and described in the dependent claims.
[0007] Invention Advantages Points in a point cloud provided by a radar sensor are described by vectors consisting of multiple numerical values. These values are, in particular, the coordinates of the point, such as polar coordinates consisting of (two-dimensional or three-dimensional) direction and distance to the point. Additionally, the radial relative velocity between the radar sensor and the object represented by that point can be mapped.
[0008] To date, information about the electromagnetic spectrum region around this point has not been included in the point cloud, or vector.
[0009] In the proposed scheme, this type of information is extracted from the spectrum by the proposed feature extractor, particularly from a spectrally defined local area (Ausschnitt) within the region at that point. Here, the feature extractor does not search for pre-given features or patterns. Instead, the feature extractor learns autonomously and unsupervised: what information it can and cannot obtain from the spectrum of a particular sensor.
[0010] For machine learning, training data is traditionally required. Traditionally, the training data is pre-labeled, that is, attributes are assigned to this training data, so that "artificial intelligence" knows what is contained in the training data.
[0011] However, by pre-defining the learning direction through labeling, content not included as a label is not learned and is therefore ignored. Thus, the success of learning depends primarily on the labeled training data. Furthermore, providing labeled training data is time-consuming and therefore expensive.
[0012] The proposed scheme in this application does not predefine what should be learned when training the feature extractor. In other words, the training objective remains largely open, only predetermined: some new content, not precisely specified beforehand, should be learned based on the provided training data. Here, training is conducted according to the principle that similar content should be categorized similarly, and different content should be categorized differently.
[0013] Training is performed using unlabeled training data. Here, local features are extracted from the spectrum and provided to a feature extractor. The feature extractor extracts specific features, or feature vectors, from these local features. During training, a self-running checking process verifies whether the features were extracted based on the following conditions: identical content should be classified similarly, and different content should be classified differently. If these conditions are not met, the parameters of the feature extractor are adjusted until they are satisfied.
[0014] Then, the trained feature extractor is applied to a local portion of the spectrum from which the point cloud has been derived. Here, a local portion of the spectrum is provided to the feature extractor for each point, or the feature extractor reads such a local portion from the spectrum. The feature extractor then provides a feature vector for each point. This feature vector has a limited length or a limited number of values. With these feature vectors, the file size of the point cloud is increased only slightly. Therefore, the supplemented point cloud can continue to be transmitted via a data line with limited bandwidth or transmission rate.
[0015] Then, in subsequent processes, the feature vector is used as additional information for object detection, route planning, or other downstream processing.
[0016] The proposed solution provides additional information about points in a point cloud at a cost-effective rate. Information not labeled in the labeled training data can be derived from the feature vectors through autonomous unsupervised training.
[0017] According to a first aspect of the invention, a feature extractor is proposed, the feature extractor being configured to read a sub-region of a spectrum detected by the sensor for each point of a point cloud provided by the sensor, extract a feature vector from each sub-region, and merge the feature vector with the corresponding point of the point cloud.
[0018] According to a second aspect of the invention, a method is proposed for using a feature extractor according to the first aspect, wherein a point cloud of a sensor is read in, and for each point of the point cloud, a sub-region of the sensor's base spectrum is read in, wherein a feature vector is extracted from each sub-region and the feature vector is merged with the corresponding point of the point cloud, wherein a point cloud having the feature vector is provided.
[0019] According to a third aspect of the invention, a method is proposed for training a feature extractor according to the first aspect, wherein unlabeled training data consisting of sub-regions of the spectrum of a sensor is read in, a feature vector is extracted from each of the sub-regions, and the extracted feature vectors are evaluated using a self-supervised cost function, wherein the parameters of the feature extractor are optimized using the result of the evaluation.
[0020] According to a fourth aspect of the invention, a driver assistance system for a vehicle equipped with at least one radar sensor is provided, the driver assistance system comprising: a feature extractor according to an embodiment of the first aspect of the invention, the feature extractor being configured to read a sub-region of a spectrum detected by the radar sensor for each point of a point cloud provided by the radar sensor, extract a feature vector from each sub-region, and merge the feature vector with the corresponding point of the point cloud; and a control device configured to control the functions of the vehicle, taking into account the points of the point cloud provided by the feature extractor and merged with the feature vector.
[0021] It can be considered that the conception of embodiments of the present invention is based particularly on the ideas and understandings described below.
[0022] A feature extractor can be a software module implemented on the hardware of a sensor. The feature extractor can be configured, in particular, for pattern recognition. "Which patterns the feature extractor recognizes" can be determined through unsupervised training of the feature extractor.
[0023] The spectrum can contain raw data from the sensor. The sensor can be a radar sensor, such that the raw data is radar sensor raw data. However, the spectrum can also be analyzed in the form of raw data from other sensors or other sensor types. The spectrum can be different for different sensors or sensor types. The feature extractor can be trained specifically for a particular sensor or sensor type. The feature extractor can be trained specifically for a particular spectrum. Thus, the imaging or detection characteristics of the corresponding sensor or sensor type can be considered.
[0024] The structure of the feature vector can be predefined. Therefore, feature extractors trained for different sensors can provide feature vectors that can be used in common.
[0025] The feature extractor can be configured to compress spectral features of a subregion into a feature vector. More information can exist in the spectrum than in the point cloud. When searching for points, the spectral features may not be adequately represented. The spectral features can be analogous to background information, allowing for different interpretations of point vectors that are otherwise identical.
[0026] The feature extractor can be configured to read in point-centered sub-regions. Points from the point cloud can be arranged in the center of each sub-region. These sub-regions can overlap.
[0027] The feature extractor can be configured to read in a sub-region of the same shape for each point. By using sub-regions of the same shape, the feature extractor can analyze and evaluate each sub-region in the same way. This makes the feature vectors comparable.
[0028] Feature extractors can be or include autonomously and unsupervised trained neural networks. Neural networks, in particular, can be convolutional neural networks (CNNs). Neural networks are particularly good at compressing or abstracting information.
[0029] Neural networks can be multi-layered. The last layer can be a summation layer, which is constructed to aggregate all remaining information from the preceding layers into a feature vector. The summation layer can ensure that feature vectors are of the same type.
[0030] Point clouds with feature vectors can be provided via an interface with average bandwidth. For example, a sensor can be connected to a central control device via this interface. The interface can have lower bandwidth compared to the bandwidth that might be required to transmit the entire spectrum.
[0031] The feature extractor can be implemented on the sensor's signal processing hardware. The feature extractor may require limited resources. This eliminates the need for dedicated, accelerated hardware to process the entire spectrum.
[0032] During training, parameters can be changed until the evaluation result is optimal. This can be done iteratively. If the result worsens, the direction of parameter changes can be reversed and / or the step size of the changes reduced.
[0033] The method is preferably implemented by a computer, and can be implemented, for example, in software or hardware or in a hybrid form of software and hardware, such as in a control device.
[0034] Another advantage is a computer program product or computer program having program code that can be stored on a machine-readable carrier or storage medium (such as semiconductor memory, hard disk memory or optical memory) and used to execute, implement and / or manipulate the steps of the method according to one of the above embodiments, especially when the program product or program is implemented on a computer or device.
[0035] It should be noted that some of the possible features and advantages of the present invention are described herein with reference to different embodiments. Those skilled in the art will recognize that features of the control device and methods can be combined, adapted, or interchanged in a suitable manner to obtain other embodiments of the invention. Attached Figure Description
[0036] Embodiments of the present invention are described below with reference to the accompanying drawings, which should not be construed as limiting the invention.
[0037] Figure 1 An illustration of a feature extractor according to one embodiment is shown; and Figure 2 An illustration shows the training of a feature extractor according to one embodiment.
[0038] These figures are illustrative only and are not drawn to scale. The same reference numerals indicate the same or equivalent features. Detailed Implementation
[0039] Figure 1A diagram of a feature extractor 100 according to one embodiment is shown. The feature extractor 100 is configured to extract feature vectors 108 from sub-regions 102 of a spectrum 106 detected by a sensor 104. Here, each sub-region in the sub-region 102 is assigned to a point 110 of a point cloud 112 derived from the spectrum 106. Points 110 are respectively contained within sub-regions 102. The feature vectors 108 are assigned to points 110 of the point cloud 112 and provided for further processing.
[0040] In one embodiment, sensor 104 is a radar sensor and detects electromagnetic spectrum 106. Here, spectrum 106 is a spectrum (range-Doppler spectrum). Coordinates in spectrum 106 correspond, for example, to a defined distance and Doppler (relative) velocity. Spectrum 106 represents, for example, the relative velocity of the detected object relative to sensor 104 by frequency shift. Similarly, phase shift, for example, can represent the distance of the object to sensor 104.
[0041] The feature extractor 100 compares the patterns in the sub-region 102 with the feature patterns learned in an unsupervised manner and maps the identified patterns into the corresponding feature vector 108.
[0042] Figure 2 The illustration shows the training of a feature extractor 100 according to one embodiment. The feature extractor 100 substantially corresponds to... Figure 1 The feature extractor 100 is used in this process. Here, unlabeled sub-regions 102 of the spectrum of a determined sensor or sensor type are provided to the feature extractor 100 as training data 200. The feature extractor 100 extracts feature vectors 108 respectively. The feature vectors are evaluated by a self-supervised cost function 202, wherein the parameters 204 of the feature extractor 100 are continuously changed until the cost of the cost function 202 is minimized.
[0043] Then, as in Figure 1 As shown, a trained feature extractor 100 is used in sensor 104 to assign corresponding feature vectors 108 to points 110 of the point cloud 112 of sensor 104. The points 110 supplemented with feature vectors 108 are then transmitted from the sensor to control device 206, where a target task 208, such as object recognition, is performed. Control device 206 may be located in a vehicle and may be part of a driver assistance system. Accordingly, the target task 208 may ultimately be the control of vehicle functions.
[0044] Here, only a line with average bandwidth is needed for transmission to control device 206, since only the supplemented point cloud 112 is transmitted instead of the entire spectrum.
[0045] In one embodiment, feature extractors 100 trained for different sensors 104 are implemented on different sensors 104. The supplemented point cloud 112 is then transmitted to a common control device 206, and the point cloud is analyzed and evaluated to achieve the target task 208.
[0046] The possible configurations of the invention are summarized again below, or shown in slightly different terms.
[0047] We propose an unsupervised, data-driven method for extracting spectrum-based feature vectors for radar sensors.
[0048] Advanced driver assistance systems (ADAS) and autonomous driving (AD) require an accurate representation of the vehicle's environment. For this purpose, radar sensors are used in addition to camera devices. These sensors provide measurements in the form of point clouds (radar reflections, positions). Each point is characterized by polar coordinates (distance, azimuth) and other properties, referred to below as manually designed features such as signal strength, radar cross section (RCS), elevation angle, and more. Using algorithms for environmental detection (perception), the position, attitude, category, and other properties, if necessary, of relevant objects (e.g., cars, trucks, or pedestrians) are extracted from these point clouds.
[0049] With the introduction of deep learning, traditional perception algorithms are increasingly being replaced by data-driven neural networks, such as object detection networks. While object detection networks are already common in LiDAR sensors, radar object detection networks remain relatively uncommon due to their limited number of points. Other deep learning-based applications, such as semantic segmentation (including classification of drivable areas or visibility checks) and localization, are also becoming increasingly technically feasible. For this purpose, radar measurements in the form of point clouds are often used because point clouds exhibit good abstraction while requiring less data.
[0050] However, in the literature, schemes utilizing information from radar spectra sometimes yield better results. Spectra are a more primitive data format than reflection, and can include additional information discarded during point cloud extraction.
[0051] Directly processing the spectrum is computationally and storage-intensive, and therefore impractical in practice. In particular, the transmission of the spectrum from the sensor to the central control device may require very high data rates.
[0052] Currently, most radar sensors only use manually designed features. Here, we consider a scheme that extracts features from the spectrum.
[0053] So far, supervised learning (with so-called labels) has been used to train on target tasks, such as object detection, where point clouds and local (patches) portions of the spectrum (for each reflection) are used as input. Here, an intermediate representation, i.e., an intermediate representation of the local spectrum, can eventually be extracted and used as a feature vector.
[0054] Manually designed features, such as radar cross section (RCS), have some advantages and disadvantages. These features enable simple utilization and good interpretation. However, the design of other features and their adaptation to new sensors and algorithms require significant manual work. The persuasiveness of features and their resulting application potential are limited because hand-designed features may result in information loss.
[0055] Data-driven and supervised feature extraction from spectral data also has its advantages and disadvantages. These features are highly persuasive and thus lead to improvements on the target task. However, because feature extraction is learned end-to-end in a supervised manner over the target task, labels are required. This introduces significantly more overhead. Furthermore, since these features are learned in conjunction with the target task, they may only be applicable to that specific task and not generally applicable to a wide range of tasks.
[0056] Current schemes for feature extraction are largely based on human decisions, such as feature design or the selection of the target task for a data-driven scheme.
[0057] Unlike manual feature design, the proposed scheme automatically learns to extract features. Furthermore, feature extraction is achieved in a general (without pre-defined target task) and unsupervised (without labels) manner.
[0058] Therefore, as a result, feature extraction from spectral data is learned in an automated and unsupervised manner, enabling the features to be used in different algorithms and target tasks. Through unsupervised learning, the proposed approach allows for the cost-effective and scalable improvement and expansion of future sensors. This is a significant advantage compared to supervised data-driven approaches, which not only require time-consuming and costly labels but also depend on the target task, such as object detection.
[0059] At its core is a module for feature extraction from the spectrum, specifically for each reflection point. Unlike the hand-created features used so far (such as signal strength, variance of distance / velocity estimates, or extended scales of peaks in the spectrum), the described module is a neural network. This neural network is trained in a self-supervised manner and is independent of subsequent target tasks. The network is trained to extract compressed feature vectors of regions of interest from the spectrum for each reflection. Training is performed in a self-supervised manner, for example, using a contrastive loss function and therefore does not require manual labels.
[0060] In the application, the module for feature extraction can then run on individual radar sensors, while downstream algorithms for the target task, such as object detection networks, run on other sensors or on a central control unit, where data from multiple radar sensors are processed as needed. The interface (the data transmitted between the sensors and the central unit) then consists of a point cloud (as before), where each point additionally contains learned (trained) features from the spectrum.
[0061] The proposed solution offers significant advantages. While transmitting point clouds from sensors to a central control device, the interface still requires a low data rate and, additionally, includes spectral information. For the (self-supervised) learning of the extraction network, only the data recorded by the sensors (spectral and reflectance) is necessary. The target task and its labels are not required. Through self-supervised learning, the learned features are independent of their subsequent use, enabling their application in a wide range of applications and algorithms. Since (manual) labels are not required for training, the generation of training data is cost-effective and time-efficient.
[0062] exist Figure 1 A simple example is shown below. Here, a two-dimensional block is extracted from the range-Doppler power spectrum, or kL power spectrum. This operation is performed for each location such that the range-Doppler bin at that location is located in the center of the corresponding block. The block has, for example, a 7×7 bin size. The feature extractor is constructed from a neural network, such as a neural network consisting of multiple 2D convolutional layers and pooling layers. The final layer of the feature extractor is a pooling layer on all remaining grid cells of the local grid, thereby generating a unique feature vector for each location. The feature vector may, for example, contain 32 learned features, each encoded in 16 bits.
[0063] The point cloud with the learned feature vectors can then be further processed by common networks, such as those used for object detection.
[0064] During training, the feature extractor is trained using a self-supervised cost function (Loss) derived from so-called pretext-tasks, such as... Figure 2 The upper part is shown. Here, manual or hand-labeled data is not necessary because the self-supervised cost function does not require these labels. A widely adopted approach for self-supervised learning is to utilize a so-called contrastive loss. This contrastive loss, or other self-supervised cost function, can be used for training. Then, the self-supervised loss is used to adjust all trainable weights of the network using gradient descent (backpropagation).
[0065] During inference time (in the product), a feature extractor is used in the sensor to solve the desired target task, such as object detection. Here, the feature extractor and the target task algorithm can run on different devices, such as... Figure 2 The lower half is shown. For example, a feature extractor can run on a radar sensor, and an object detector can run on a central control device.
[0066] In a system with multiple different radar sensors, a separate feature extractor can be used for each sensor type, and the object detector aggregates point clouds with features learned from all sensors.
[0067] In one embodiment, the method described above is applied in a simplified manner. Here, the region from the spectrum can be a mapping different from a two-dimensional range-Doppler mapping. Instead of a power spectrum, amplitude and phase, or a complex-valued spectrum, can also be used. The region in the spectrum can be different from two dimensions, i.e., it can be one-dimensional (e.g., range only or Doppler only), three-dimensional (e.g., range, Doppler, and azimuth), or multi-dimensional. The region in the spectrum can be different from a rectangular block. The region can use arbitrary spectral units (Zellens) in the spectrum. The feature extractor can be composed of arbitrary layers of a neural network. The network used for the final target task is not limited to a single task, such as object detection, but can also include semantic segmentation, localization, other tasks, or a combination of multiple tasks.
[0068] Furthermore, it can be extended. Multiple networks can be trained for different tasks using the same feature extractor. This feature extractor can then be used for multiple tasks. The tasks of the network following the feature extractor are not limited to perception, but can also include tasks such as planning and functionalization. Additionally, modulation parameters, antenna calibration data, and HF settings (transmit power, HF filters, etc.) can be utilized to account for or compensate for sensor and modulation dependencies.
[0069] Instead of feature extraction for each reflection, a (highly compressed) feature representation can be learned over the entire spectrum. The goal is to enable target tasks to be performed based on the radar spectrum, where feature extraction by the neural network runs on the sensor and the actual task runs on a central control device. Furthermore, the application is not limited to radar data. The method can be applied to any sensor with a point cloud as its output representation, such as lidar, RGB-D cameras, etc. Then, instead of the spectrum, other raw data representations can be chosen. For example, the received time signal of each ray from a time-of-flight lidar sensor could be used.
[0070] exist Figure 1 The diagram illustrates an exemplary feature extraction process on a radar sensor based on a local view of the spectrum for all reflections. Here, a block (sub-region) from the associated spectrum is used for each reflection to extract features.
[0071] exist Figure 2 The diagram illustrates the training process and subsequent use of the trained feature extractor. During training, the extractor is trained using blocks from the spectrum. Here, a self-supervised cost function is utilized to solve a so-called pre-task. The pre-task does not require labels. After training, the feature extractor is used on one or more sensors to extract additional features from the associated spectrum in addition to reflections. This point cloud, or these point clouds, are then used for any target task, such as object detection.
[0072] The proposed architecture can be used in future driver assistance systems and automated driving functions. Other applications include all radar sensors utilizing machine learning, such as fixed radar sensors for traffic monitoring or radar sensors for bicycles. Furthermore, the proposed solution can also be used in the field of robotics (e.g., obstacle recognition for autonomous lawnmowers).
[0073] Finally, it should be noted that terms such as "having" or "comprising" do not exclude other elements or steps, and terms such as "a" do not exclude multiple. Reference numerals in the claims should not be considered limiting.
Claims
1. A feature extractor (100) configured to read a sub-region (102) of a spectrum (106) detected by the sensor (104) for each point (110) of a point cloud (112) provided by the sensor (104), extract a feature vector (108) from each sub-region (102), and merge the feature vector (108) with the corresponding point (110) of the point cloud (112).
2. The feature extractor (100) according to claim 1, wherein the feature extractor is configured to compress the spectral features of the sub-region (102) into the feature vector (108).
3. The feature extractor (100) according to any one of the preceding claims, the feature extractor being configured to read in a sub-region (102) centered on the point (110).
4. The feature extractor (100) according to any one of the preceding claims, the feature extractor being configured to read a sub-region (102) of the same shape for each of the points (110).
5. The feature extractor (100) according to any one of the preceding claims, wherein, The feature extractor (100) is a neural network trained autonomously and unsupervised.
6. The feature extractor (100) according to claim 5, wherein, The neural network is multi-layered, with the last layer being a summarizing layer configured to summarize all remaining information from the preceding layers into the feature vector (108).
7. A method for using a feature extractor (100) according to any one of claims 1 to 6, wherein, The point cloud (112) of the sensor (104) is read in, and for each point (110) of the point cloud (112), a sub-region (102) of the spectrum (106) of the sensor (104) as the basis is read in, wherein a feature vector (108) is extracted from each sub-region (102) and the feature vector is merged with the corresponding point (110) of the point cloud (112), wherein a point cloud (112) with the feature vector (108) is provided.
8. The method according to claim 7, wherein, The point cloud (112) with the feature vector (108) is provided via an interface with average bandwidth.
9. The method according to any one of the preceding claims, wherein, The feature extractor (100) is implemented on the signal processing hardware of the sensor (104).
10. A method for training a feature extractor (100) according to any one of claims 1 to 6, wherein, Unlabeled training data (200) consisting of sub-regions (102) of the spectrum (106) of the sensor (104) is read in, a feature vector (108) is extracted from each sub-region (102), and the extracted feature vectors (108) are evaluated using a self-supervised cost function (202), wherein the parameters (204) of the feature extractor (100) are optimized using the results of the evaluation.
11. The method according to claim 10, wherein, Change the parameter (204) until the evaluation result reaches the optimal level.
12. A driver assistance system for a vehicle equipped with at least one radar sensor, the driver assistance system comprising: The feature extractor (100) according to any one of claims 1 to 6 is configured to read a sub-region (102) of the spectrum (106) detected by the radar sensor for each point (110) of the point cloud (112) provided by the radar sensor, extract a feature vector (108) from each sub-region (102), and merge the feature vector (108) with the corresponding point (110) of the point cloud (112); and A control device (206) configured to control the functions of the vehicle, taking into account the points (110) of the point cloud (112) provided by the feature extractor and merged with the feature vector (108).
Citation Information
Patent Citations
Method for evaluating radar data of at least one radar sensor in a motor vehicle and motor vehicle
DE102016002208A1
Method and system for image evaluation for object recognition in the vicinity of vehicles
DE102019219144A1