Cross-modal four-dimensional radar denoising method and device based on lidar supervision
By constructing a cross-modal four-dimensional radar denoising model based on lidar supervision, and aligning lidar point clouds with four-dimensional radar point clouds to generate supervised point masks, the problem of distinguishing noise points in four-dimensional radar point clouds is solved, thereby improving the perception capability and robustness of autonomous driving systems.
Patent Information
- Application Number
- CN202411840206.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-12-13
AI Technical Summary
Existing technologies struggle to effectively distinguish between noise points and valid points in four-dimensional radar point clouds, leading to a high false detection rate by sensors and impacting the robustness and versatility of autonomous driving systems.
A cross-modal four-dimensional radar denoising method based on lidar supervision is adopted. A denoising model is constructed through a feature encoder, a noise predictor and a matching module. The lidar point cloud is aligned with the four-dimensional radar point cloud to generate a supervised point mask, and the denoising network is trained to distinguish between noise and valid points.
It improves the perception capability of four-dimensional radar in high-noise environments, enhances the robustness and perception accuracy of autonomous driving systems, and reduces the false detection rate.
Smart Images

Figure CN119689430B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the fields of automatic driving technology, environment perception technology and millimeter wave radar denoising technology, and particularly relates to a cross-modal four-dimensional radar denoising network method based on laser radar supervision, a device, an electronic equipment, a computer readable storage medium and a computer program product. BACKGROUND
[0002] The environment perception module is crucial in automatic driving and plays a pivotal role in the subsequent operation of the planning, decision-making and control modules. At present, most automatic driving systems use cameras and laser radars as the main sensors, but these sensors are not ideal in adverse weather conditions. Millimeter wave radar sensors use relatively long wavelengths for detection and have strong anti-interference ability, and they have attracted widespread attention due to their environmental adaptability and good reliability. In the field of automatic driving, the mainstream radar sensors are mainly divided into three-dimensional radars and four-dimensional radars. Compared with three-dimensional radars, four-dimensional radars only exist in millimeter wave radars, and four-dimensional radars can provide spatial three-dimensional information such as distance, azimuth angle, elevation angle and Doppler velocity, so that more spatial information can be used to cope with various driving scenarios.
[0003] Compared with laser radars, the point clouds collected by millimeter wave radars exhibit strong noise, which is particularly prominent in four-dimensional radars. Strong noise significantly increases the false detection rate of the sensor, so whether the effective points and noise points in the four-dimensional radar point cloud can be accurately distinguished is a major challenge in the perception strategy of four-dimensional radars. In the early stage, a large number of three-dimensional radar denoising methods were proposed: some traditional optimization-based methods, such as the moving least squares method and low-rank learning, use geometric projection and self-similarity for denoising, but such means are difficult to handle complex, noisy and highly irregular radar data. It is worth mentioning that neural network-based methods have achieved impressive results under low noise conditions, but they often cannot handle denser, more variable noise levels and larger-scale point clouds. Some workers focus on modality fusion and try to alleviate the limitations of radars by integrating laser radar information, but they have not been explored on four-dimensional radars, where noise and sparsity are more obvious, limiting their robustness and universality in real-world scenarios. SUMMARY
[0004] In view of the deficiencies of the prior art, as shown in the background, the present application proposes a cross-modal four-dimensional radar denoising method based on laser radar supervision, which comprises the following steps: Figure 3
[0005] An initial step, a denoising model composed of a feature encoder, a noise predictor and a matching module is constructed, and four-dimensional radar point clouds with labeled bounding boxes and real classes are obtained as training data;
[0006] The training step extracts aggregate features of the training data by the feature encoder; the noise predictor obtains a predicted point cloud according to the aggregate features, and each position data point in the predicted point cloud has a predicted category; the matching module obtains a predicted point mask after aligning the predicted point cloud with the training data, and inputs the predicted point mask into a classification model to obtain a target region in the training data and a corresponding predicted category of the target region; a boundary loss is constructed according to the target region and the bounding box, and a category loss is constructed according to the predicted category and the real category; and the denoising model is trained according to the boundary loss and the category loss.
[0007] The denoising step inputs the four-dimensional radar point cloud to be denoised into the denoising model after training to obtain a four-dimensional radar denoising result point cloud.
[0008] The laser radar supervised cross-modal four-dimensional radar denoising method, wherein the training data comprises: laser radar point cloud and four-dimensional radar point cloud photographed in the same scene.
[0009] The training step comprises: aligning the laser radar point cloud with the four-dimensional radar point cloud, generating a supervised point mask in the form of a voxel based on the laser radar point cloud, matching the predicted point mask and the supervised point mask to obtain a cross-modal supervision loss, and training the denoising model according to the boundary loss, the category loss and the cross-modal supervision loss.
[0010] The laser radar supervised cross-modal four-dimensional radar denoising method, wherein the feature encoder applies multiple encoding layers to gradually extract the aggregate features from the point cloud.
[0011] The laser radar supervised cross-modal four-dimensional radar denoising method, wherein the denoising step comprises: analyzing the four-dimensional radar denoising result point cloud by the classification model to obtain a classification to which each point in the four-dimensional radar denoising result point cloud belongs, and controlling a vehicle to complete an automatic or assisted driving task, such as assisting in avoiding pedestrians or an automatic emergency braking AEB task, based on the classification to which each point in the four-dimensional radar denoising result point cloud belongs.
[0012] As shown in Figure 4 The present application also provides a laser radar supervised cross-modal four-dimensional radar denoising device, which comprises:
[0013] An initial module, a denoising model composed of a feature encoder, a noise predictor and a matching module, four-dimensional radar point cloud with labeled bounding boxes and real categories as training data;
[0014] The training module extracts aggregated features of the training data; the noise predictor obtains a predicted point cloud according to the aggregated features, each position data point in the predicted point cloud having a predicted category; the matching module obtains a predicted point mask by aligning the predicted point cloud with the training data, and inputs the predicted point mask into a classification model to obtain a target region in the training data and a corresponding predicted category of the target region; a boundary loss is constructed according to the target region and the bounding box, and a category loss is constructed according to the predicted category and the real category; and the denoising model is trained according to the boundary loss and the category loss.
[0015] The denoising module inputs the four-dimensional radar point cloud to be denoised into the denoising model after training to obtain a four-dimensional radar denoising result point cloud.
[0016] The training data includes laser radar point clouds and four-dimensional radar point clouds photographed in the same scene.
[0017] The training module includes aligning the laser radar point clouds and the four-dimensional radar point clouds, generating a supervised point mask in the form of voxels from the laser radar point clouds, matching the predicted point mask and the supervised point mask to obtain a cross-modal supervision loss, and training the denoising model according to the boundary loss, the category loss and the cross-modal supervision loss.
[0018] The denoising module includes analyzing the four-dimensional radar denoising result point cloud by using a classification model to obtain a classification to which each point in the four-dimensional radar denoising result point cloud belongs, and controlling a vehicle to complete an automatic or assisted driving task based on the classification to which each point in the four-dimensional radar denoising result point cloud belongs.
[0019] The present application also provides an electronic device comprising the cross-modal four-dimensional radar denoising device, and the electronic device is connected with an information display device configured to display the four-dimensional radar denoising result point cloud in a display parameter, attribute or through an artificial intelligence model set by a user.
[0020] The present application also provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the cross-modal four-dimensional radar denoising method.
[0021] The present application also provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the steps of the cross-modal four-dimensional radar denoising method.
[0022] As can be seen from the above solutions, the present application has the following advantages:
[0023] The present application finds that the noise is mainly due to the energy leakage around the strong target and the limited resolution of the radar through the analysis of the four-dimensional radar noise. Therefore, to improve the perception ability of the four-dimensional radar, it is necessary to identify the points in the point cloud related to the actual object. The present application is inspired from the latest progress in multi-modal fusion, because the lidar can accurately describe the real position of the object in the physical world, the present application hopes to improve the detection ability of the four-dimensional radar based model by normatively constraining the four-dimensional radar by the lidar. The present application proposes a cross-modal learning strategy, which utilizes the lidar point cloud in the training process, so that the detection model can implicitly identify the noise. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 The millimeter wave radar denoising diagram of the present application;
[0025] Figure 2 The processing device diagram of the present application;
[0026] Figure 3 The method flowchart of the present application;
[0027] Figure 4 The device module diagram of the present application;
[0028] Figure 5 The first electronic equipment structure schematic diagram of the present application;
[0029] Figure 6 The first electronic equipment application environment structure schematic diagram of the present application;
[0030] Figure 7 The second electronic equipment structure schematic diagram of the present application.
[0031] Reference signs:
[0032] A-First electronic equipment;
[0033] B-Cross-modal four-dimensional radar denoising device;
[0034] C-Data acquisition equipment;
[0035] D-Information display equipment;
[0036] 1000-Second electronic equipment;
[0037] I-Computing unit;
[0038] II-ROM;
[0039] III-RAM;
[0040] IV-Bus;
[0041] V-Interface;
[0042] VI-Input unit;
[0043] VII - an output unit;
[0044] VIII - a storage medium;
[0045] IX - a communication unit. DETAILED DESCRIPTION
[0046] It should be noted that the relationship terms such as first and second, and the like, are used only to differentiate one entity or operation from another, and do not necessarily require or imply such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus including a list of elements does not only include those elements, but also includes other elements not explicitly listed, or further includes elements inherent in such process, method, article or apparatus.
[0047] Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus including the element.
[0048] The processor of the present application is the control center of the electronic device, which can be one processor or a collective term of multiple processing elements. For example, it can be one or more central processing units (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present application, such as one or more digital signal processors (DSP), or one or more field programmable gate arrays (FPGA).
[0049] Optionally, the processor can perform various functions of the electronic device by running or executing software programs stored in the memory, and calling data stored in the memory.
[0050] In a particular implementation, as an example, the processors can include one or more CPUs. Each of the processors can be a single-CPU or a multi-CPU. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). The electronic device can include a server, a desktop computer, a notebook computer, a smart phone, a tablet computer, an embedded computer, and the like, where the embedded computer includes a vehicle and a robot, and the like.
[0051] The memory is used to store the software program for implementing the scheme of the present application, and is controlled by the processor to execute. The specific implementation can refer to the method embodiments described above, and will not be repeated here.
[0052] It should be noted that the structure of the electronic device shown in the drawings of the present application does not constitute a limitation thereon, and the actual knowledge structure recognition device can include more or fewer components than shown, or combine certain components, or different component arrangements.
[0053] The above embodiments can be implemented in whole or in part by software, hardware (such as a circuit), firmware, or any combination thereof. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the flow or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid state disk.
[0054] It should also be understood that, in the specification, terms "and / or" merely describes an associated relationship, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, B exists alone, where A and B can be singular or plural. In addition, the character " / " generally represents an "or" relationship between the front and rear associated objects, but it can also represent an "and / or" relationship, which can be understood in the context.
[0055] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be singular or plural.
[0056] It should also be understood that the order of the above processes in various embodiments of the present application does not mean the order of execution, and the execution order of the processes should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0057] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the above-described device embodiments are only schematic, and the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other form.
[0058] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.
[0059] In addition, the functional units in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.
[0060] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0061] The present application finds that the noise is mainly due to the energy leakage around the strong target and the limited resolution of the radar through the analysis of the four-dimensional radar noise. Therefore, to improve the perception ability of the four-dimensional radar, it is necessary to identify the points in the point cloud related to the actual object. The present application is inspired by the latest progress in multi-modal fusion, because the lidar can accurately describe the real position of the object in the physical world, the present application hopes to improve the detection ability of the four-dimensional radar based model by regulating the four-dimensional radar with the lidar. The present application proposes a cross-modal learning strategy, which uses the lidar point cloud in the training process to enable the detection model to implicitly identify the noise. Specifically, the present application includes the following key technical points:
[0062] Technical point 1, the present application proposes a millimeter wave radar denoising network based on hierarchical feature aggregation and prediction matching mechanism, which uses hierarchical feature aggregation to construct a multi-scale feature extraction module of four-dimensional millimeter wave radar, and uses the fitting ability of neural network to effectively distinguish the noise and effective point cloud in the millimeter wave radar point cloud.
[0063] Technical point 2, the core network of the present application includes three parts: feature encoder, noise predictor and matching module. The feature encoder: gradually extracts and aggregates features from the radar point cloud through a hierarchical feature learning architecture; the noise predictor: estimates the noise and effectiveness of each point in the radar point cloud, generates a prediction mask, and the mask score reflects the effectiveness of the point cloud; the matching module: based on the KD-Tree structure, the filtered point cloud and the original point cloud are matched to ensure the coverage range of the effective point cloud.
[0064] Technical point 3, the matched point cloud is processed by the backbone network and the detection head to generate class prediction and position prediction. The binary cross entropy loss function and the mean square error loss function are used to supervise the training of the network, wherein the loss function is used to optimize the predicted bounding box position and class.
[0065] Technical point 4, use laser radar point cloud as a supervision signal in the training process to denoise the four-dimensional millimeter wave radar point cloud, the mechanism includes: the laser radar point cloud is processed by the mask generator to generate a supervised point mask; the mask is matched with the predicted mask of the radar point cloud, the loss function is calculated to guide the training of the denoising network.
[0066] Technical point 5, in the inference stage, because the trained denoising network has effective denoising ability, laser radar data is no longer needed. By eliminating the need for laser radar data in the inference process, the cross-modal processing device maintains the inference efficiency and practicality of perception.
[0067] In order to make the above features and effects of the present application more clear and easy to understand, the following embodiments are given, and the detailed description is as follows with the aid of the accompanying drawings. This specification discloses one or more embodiments containing the features of the present application. The disclosed embodiments are only for illustration. The protection scope of the present application is not limited to the disclosed embodiments, and the present application is defined by the appended claims.
[0068] The object of the present application is to solve the high level noise problem of four-dimensional millimeter wave radar in environmental perception, and a cross-modal four-dimensional radar supervision method and device based on laser radar point cloud is proposed, which effectively distinguishes noise and effective point cloud by aligning laser radar point cloud and four-dimensional radar point cloud. Specifically, it includes three parts: constructing the core module for point cloud denoising; introducing the cross-modal supervision mechanism of laser radar-four-dimensional millimeter wave radar; 3, the full loss function design involving mask calculation. The above steps are described in detail as follows.
[0069] 1, construct a millimeter wave radar denoising network under the mechanism of hierarchical feature aggregation and prediction matching
[0070] The denoising network designed in the present application is composed of three parts: feature encoder noise predictor and matching module The present application adopts a hierarchical feature learning architecture to design the feature encoder. Given a high-noise radar point cloud x as input, the feature encoder applies n encoding layers to gradually extract aggregated features from the point cloud. Each layer represents a transformation of the input, with the goal of discarding noise while retaining useful information, i.e., feature extraction is performed on each point cloud. The difference between "noise points" and "non-noise points" is small, and the feature encoder acts to enlarge the difference between "noise point" features and "non-noise point" features, with the goal of preparing for subsequent noise removal and effective information retention. The connection between layers captures the complex interactions in hidden features h. n The output feature h
[0071]
[0072] The extracted point features h n are passed to the noise predictor The neural network of this module is responsible for estimating the noise and valid points in the radar point cloud.
[0073]
[0074] M radar is the predicted mask, indicating the location of the valid radar point cloud. The data size of the mask is the same size as the point cloud, and the mask score reflects the validity of the point cloud. The higher the score, the higher the confidence of the point cloud, and vice versa. The predicted mask M radar Input matching module The matching module aligns the point cloud filtered by the noise predictor P with the original radar point cloud, ensuring that the receptive field is sufficient to cover most of the valid point cloud. The present invention uses a KD-Tree structure to implement the matching module.
[0075]
[0076] The matched point cloud P match is then processed by the backbone network and detection head to obtain the detection result.c cls and y reg are the class prediction and position prediction respectively. The present invention uses binary cross-entropy loss BCE(·) and mean square error MSE(·) regression to supervise network training,
[0077] L cls = BCE(c cls ,c gt )
[0078] L reg = MSE(y reg -y gt )
[0079] where y gt is the true bounding box, and c gt is the true class of the target region, such as vehicle, pedestrian, lane line, etc.
[0080] The network aims to improve the reliability of radar point cloud in high noise environment, which is crucial for robust perception of autonomous driving systems.
[0081] 2. Cross-modal supervision mechanism using lidar-millimeter wave radar:
[0082] The present application introduces a new cross-modal supervision mechanism, which uses lidar point cloud to supervise the denoising process of four-dimensional millimeter wave radar point cloud during training, while eliminating the dependence on lidar data during inference.
[0083] During the training phase, both lidar and four-dimensional radar point clouds are input data. The lidar point cloud is processed by a mask generator, which aligns the lidar data with the four-dimensional radar data. This module generates a supervised point mask in the form of voxels, which is used to guide the denoising process of the radar point cloud. At the same time, the radar point cloud is processed by the denoising network designed in the previous subsection, which produces a predicted mask that can distinguish between valid point clouds and noise point clouds. Then the predicted point mask and the supervised point mask obtained from the lidar data are matched, and the two masks will be used to calculate the L msk where N is the number of points, and L1 represents the smoothL1 loss.
[0084]
[0085] During the inference phase, since the trained denoising network has effective denoising capability, lidar data is no longer needed for calculation. By eliminating the need for lidar data during inference, this cross-modal framework maintains the efficiency and practicality of radar-based perception, while benefiting from the superior accuracy of lidar supervision during training.
[0086] 3. Full loss function design involving mask calculation:
[0087] For network training of a three-dimensional target detection model, the present application proposes a loss function involving mask calculation:
[0088] L total = alpha * L cls + beta * L reg + gamma * L msk
[0089] L total is composed of three parts: the classification loss of the classifier (L cls ), the regression loss (L reg ), and the valid voxel loss calculated based on the predicted mask and the supervised mask (L msk ). L cls uses focal loss, L reg and L msk use smoothL1 function. The coefficients alpha, beta and gamma are set as hyperparameters, with sizes of 1, 2 and 10 respectively. The present application uses gradient descent algorithm to optimize the weights to minimize the loss.
[0090] Embodiment 1
[0091] Figure 1This invention provides a millimeter-wave radar denoising network design method based on hierarchical feature aggregation and prediction matching mechanism, the steps of which are as follows:
[0092] S11: Receives raw millimeter-wave radar point cloud as an input module;
[0093] S12: Construct a feature extraction module under a hierarchical feature aggregation mechanism to extract point cloud features;
[0094] S13: Predict noisy point clouds using a neural network noise prediction module based on an MLP structure;
[0095] S14: Use the KD-Tree-based point cloud matching module to match valid point clouds;
[0096] S15: Outputs the radar point cloud after noise filtering as an output module.
[0097] Example 2
[0098] This invention provides a processing device for a cross-modal four-dimensional radar supervision method based on lidar point clouds, such as... Figure 2 As shown, the device includes: a lidar point cloud data reading module S21, a cross-modal ground truth point cloud mask generation module S22, a millimeter-wave radar denoising network under a hierarchical feature aggregation and prediction matching mechanism S23, a backbone network and detection head module for processing radar point clouds S24, a cross-modal mask loss calculation module S25, a millimeter-wave radar point cloud data reading module S26, and an overall loss calculation module S27.
[0099] The laser radar point cloud data reading module S21: the module is divided according to the training set and the test set, and the laser radar point cloud data is read; the cross-modal true value point cloud mask generation module S22: the module receives the laser radar point cloud as input, generates a supervised point mask as the supervision information of the radar point cloud; the millimeter wave radar denoising network under the hierarchical feature aggregation and prediction matching mechanism S23: the module receives the original millimeter wave radar point cloud as input, extracts the point cloud features through hierarchical feature aggregation, and predicts the noise point cloud based on the point cloud features, so as to obtain the point cloud filtered out of noise; the backbone network and detection head module for processing radar point cloud S24: the module extracts features and detects targets from the point cloud filtered out of noise, so as to obtain the position information of the target; the cross-modal mask loss calculation module S25: the module calculates the loss in the training stage, and calculates the smoothL1 loss of the predicted point mask and the supervised point mask obtained from the laser radar data; the millimeter wave radar point cloud data reading module S26: the module is divided according to the training set and the test set, and the millimeter wave radar point cloud data is read; the overall loss calculation module S27: the module is used to calculate the classification loss, the regression loss of the classifier and the effective voxel loss calculated based on the prediction mask and the supervision mask, and finally the final loss result is obtained by weighting and summing the loss through the hyperparameter.
[0100] The processing device provided by the embodiment of the present application based on the cross-modal four-dimensional radar supervision method of the laser radar point cloud has the same technical features as the cross-modal four-dimensional radar denoising network method based on the laser radar supervision, and therefore can also realize the functions described above, which will not be described here.
[0101] The following is a system embodiment corresponding to the above-mentioned method embodiment. The technical details mentioned in the above-mentioned embodiment are still valid in this embodiment. In order to reduce repetition, they will not be described here. Correspondingly, the technical details mentioned in this embodiment can also be applied to the above-mentioned embodiment.
[0102] As shown in Figure 4 The present application also provides a cross-modal four-dimensional radar denoising device based on laser radar supervision, which comprises:
[0103] An initial module, a denoising model composed of a feature encoder, a noise predictor and a matching module, a four-dimensional radar point cloud with labeled bounding boxes and real classes as training data;
[0104] The training module extracts aggregated features of the training data; the noise predictor obtains a predicted point cloud according to the aggregated features, each position data point in the predicted point cloud having a predicted category; the matching module obtains a predicted point mask by aligning the predicted point cloud with the training data, removes noise, and inputs the predicted point mask into a classification model to obtain a target region in the training data and a corresponding predicted category; a boundary loss is constructed according to the target region and the bounding box, and a category loss is constructed according to the predicted category and the real category; and the denoising model is trained according to the boundary loss and the category loss.
[0105] The denoising module inputs the four-dimensional radar point cloud to be denoised into the denoising model after training to obtain a four-dimensional radar denoising result point cloud.
[0106] The cross-modal four-dimensional radar denoising device based on laser radar supervision, wherein the training data comprises: laser radar point clouds and four-dimensional radar point clouds photographed in the same scene;
[0107] The training module comprises: aligning the laser radar point cloud with the four-dimensional radar point cloud, generating a supervised point mask in the form of a voxel from the laser radar point cloud, matching the predicted point mask and the supervised point mask, and obtaining a cross-modal supervision loss; and training the denoising model according to the boundary loss, the category loss, and the cross-modal supervision loss.
[0108] The cross-modal four-dimensional radar denoising device based on laser radar supervision, wherein the denoising module comprises: analyzing the four-dimensional radar denoising result point cloud by using a classification model to obtain a classification to which each point in the four-dimensional radar denoising result point cloud belongs, and controlling a vehicle to complete an automatic or assisted driving task based on the classification to which each point in the four-dimensional radar denoising result point cloud belongs.
[0109] As shown in Figure 5 The present application also proposes a first electronic device A comprising the cross-modal four-dimensional radar denoising device.
[0110] As shown in Figure 6 The first electronic device A can be connected to a data acquisition device C and an information display device D through a wired or wireless information transmission scheme, the data acquisition device C being used to acquire four-dimensional radar point clouds and / or laser point clouds, and the information display device D being used to display the four-dimensional radar denoising result point cloud obtained by the present application.
[0111] The information display device D can process the data output by the first electronic device A based on an information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be manually preset, for example, the data output by the first electronic device A is visually displayed, which can be displayed according to the display parameters and / or attributes set by the user, for example, the display data range, the display attributes can be, for example, the display font, the color, whether to scroll and play, etc. The user can present the information specified by the user, and the user can understand the information more timely without accessing the secondary page or scrolling the page, saving the operation of the user. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the user's key attention information according to the user's previous use habits, such as watching time, click times, editing times, etc., and then automatically present the user with rich and necessary key information.
[0112] The application further provides a computer program product, which comprises a computer program, the computer program can be stored on a readable storage medium, and the computer program can execute the laser radar supervision based cross-modal four-dimensional radar denoising method provided by the above-mentioned method when executed by a processor.
[0113] The present application also proposes, in another implementation, a storage medium VIII for storing a computer program for executing the laser-radar supervision based cross-modal four-dimensional radar denoising method. It should be understood that the storage medium in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, and not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct Rambus RAM (DR RAM).
[0114] Figure 7 A schematic block diagram of a second electronic device 1000 that can be used to implement embodiments of the present application is shown. The second electronic device 1000 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The second electronic device 1000 can also represent various forms of mobile devices such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document. The second electronic device 1000 can be the same as or different from the first electronic device A.
[0115] The second electronic device 1000 includes a computing unit I that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory II (ROM) or a computer program loaded into a random access memory (RAM) III from a storage medium VIII. In the RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, the ROM II, and the RAM III are connected to each other through a bus IV. An input / output (I / O) interface V is also connected to the bus IV.
[0116] A plurality of components in the second electronic device 1000 are connected to the I / O interface V, including an input unit VI such as a keyboard, a mouse, and the like, an output unit VII such as various types of displays, a speaker, and the like, a storage medium VIII such as a magnetic disk, an optical disk, and the like, and a communication unit IX such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0117] The computing unit I can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit I performs various methods and processes described above, such as the method steps S1-S3. For example, in some embodiments, the methods can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM II and / or the communication unit IX. When the computer program is loaded into the RAM III and executed by the computing unit I, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit I can be configured to perform the methods by any other appropriate means, such as by means of firmware.
[0118] While the embodiments of the present application have been disclosed as above, they are not limited to only the applications listed in the specification and the embodiments, and can be fully applied to various fields suitable for the present application, and additional modifications can be easily made by those skilled in the art, and thus the present application is not limited to specific details and the figures shown and described herein, without departing from the general concept defined by the claims and the equivalent scope.
Claims
1. A cross-modal four-dimensional radar denoising method based on lidar supervision, characterized in that, The method comprises the following steps: An initial step, a denoising model composed of a feature encoder, a noise predictor and a matching module is constructed, and four-dimensional radar point clouds with labeled bounding boxes and real classes are obtained as training data; A training step, the feature encoder extracts aggregated features of the training data; The noise predictor obtains a predicted point cloud according to the aggregated features, and each position data point in the predicted point cloud has a predicted class; the matching module aligns the predicted point cloud with the training data to obtain a predicted point mask without noise, inputs the predicted point mask into a classification model to obtain a target region in the training data and a corresponding predicted class of the target region; a boundary loss is constructed according to the target region and the bounding box, and a class loss is constructed according to the predicted class and the real class; the denoising model is trained according to the boundary loss and the class loss; A denoising step, the four-dimensional radar point cloud to be denoised is input into the trained denoising model to obtain a four-dimensional radar denoising result point cloud; The training data comprises laser radar point clouds and four-dimensional radar point clouds captured under the same scene. The training step comprises the following steps: aligning the laser radar point clouds with the four-dimensional radar point clouds, generating a supervised point mask in the form of voxels from the laser radar point clouds, matching the predicted point mask with the supervised point mask to obtain a cross-modal supervised loss, and training the denoising model according to the boundary loss, the class loss and the cross-modal supervised loss.
2. The LIDAR supervision based cross-modal four-dimensional radar de-noising method of claim 1, wherein, The feature encoder applies multiple encoding layers to gradually extract the aggregated features from the point cloud.
3. The LIDAR supervision based cross-modal four-dimensional radar de-noising method of claim 1, wherein, The denoising step comprises the following steps: analyzing the four-dimensional radar denoising result point cloud by using the classification model to obtain the classification to which each point in the four-dimensional radar denoising result point cloud belongs, and controlling a vehicle to complete an automatic or assisted driving task based on the classification to which each point in the four-dimensional radar denoising result point cloud belongs.
4. A cross-modal four-dimensional radar denoising device based on lidar supervision, characterized in that, The method comprises the following steps: An initial module, a denoising model composed of a feature encoder, a noise predictor and a matching module is constructed, and four-dimensional radar point clouds with labeled bounding boxes and real classes are obtained as training data; A training module, the feature encoder extracts aggregated features of the training data; The noise predictor obtains a predicted point cloud according to the aggregated features, and each position data point in the predicted point cloud has a predicted class; the matching module aligns the predicted point cloud with the training data to obtain a predicted point mask without noise, inputs the predicted point mask into a classification model to obtain a target region in the training data and a corresponding predicted class of the target region; a boundary loss is constructed according to the target region and the bounding box, and a class loss is constructed according to the predicted class and the real class; the denoising model is trained according to the boundary loss and the class loss; A denoising module, the four-dimensional radar point cloud to be denoised is input into the trained denoising model to obtain a four-dimensional radar denoising result point cloud; The training data comprises laser radar point clouds and four-dimensional radar point clouds captured under the same scene. The training module comprises: aligning the lidar point cloud with the four-dimensional radar point cloud, and generating a supervised point mask in the form of voxels from the lidar point cloud; matching the predicted point mask and the supervised point mask to obtain a cross-modal supervision loss; and training the denoising model according to the boundary loss, the category loss, and the cross-modal supervision loss.
5. The LIDAR supervision based cross-modal four-dimensional radar de-noising apparatus as claimed in claim 4, wherein, The denoising module comprises: analyzing the four-dimensional radar denoising result point cloud by a classification model to obtain the classification to which each point in the four-dimensional radar denoising result point cloud belongs, and controlling the vehicle to complete an automatic or assisted driving task based on the classification to which each point in the four-dimensional radar denoising result point cloud belongs.
6. An electronic device, comprising: The electronic device or the information display device connected thereto is used to display the four-dimensional radar denoising result point cloud in a user-set display parameter, attribute, or through an artificial intelligence model. 7.A computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the steps of the cross-modal four-dimensional radar denoising method of any one of claims 1-3.
8. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the cross-modal four-dimensional radar denoising method of any one of claims 1-3.