A 4D radar multimodal dataset construction and evaluation method and device

By building a multi-sensor synchronous acquisition platform and sensor parameter calibration, the problems of closed 4D radar dataset processing flow and difficult information quality assessment were solved, the construction and evaluation of high-precision multimodal datasets were achieved, and the perception capability of the autonomous driving system in complex environments was improved.

CN120491047BActive Publication Date: 2025-09-26DONGHAI LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510983037.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-26
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

The existing 4D radar dataset processing process is closed, the radar information quality is difficult to evaluate, and there is a lack of controllable collection and evaluation mechanism for parameters such as different radar emission modes and frequency settings. The dataset has limitations in scene diversity, time synchronization mechanism and annotation accuracy, making it difficult to meet the verification requirements of multimodal fusion perception algorithms in complex dynamic environments.

Method used

Build a multi-sensor synchronous acquisition platform that includes different types of sensors, such as 4D radar, visible light camera, infrared camera, lidar and inertial navigation unit, and collect multimodal data through unified time synchronization and spatial layout design; calibrate the parameters and synchronize the time of the sensors, build a multimodal dataset, and evaluate it through point cloud generation testing and target detection calculation average accuracy.

Benefits of technology

It improves the scene diversity and robustness of the dataset, provides high-precision multimodal annotation, supports in-depth research and evaluation of radar perception performance, and ensures the effective application of the dataset in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491047B_ABST
    Figure CN120491047B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for constructing and evaluating a 4D radar multimodal dataset. The method provided in the present application includes: constructing a multi-sensor synchronous acquisition platform; collecting multimodal data in multiple acquisition scenarios based on the multi-sensor synchronous acquisition platform to obtain a multimodal dataset; performing a point cloud generation test on the multimodal dataset, generating a point cloud using the multimodal dataset, and calculating the average accuracy of the point cloud based on the correspondence between the predicted points and the real points in the point cloud; performing target detection on the multimodal dataset, and calculating the average accuracy of the multimodal dataset under multiple recall thresholds; and evaluating the performance of the multimodal dataset based on the average accuracy and average precision of the point cloud. The method and device for constructing and evaluating a 4D radar multimodal dataset provided in the present application are used to solve the problems of closed processing procedures of existing 4D radar datasets and difficulty in evaluating radar information quality, and to construct a 4D radar multimodal dataset with an open radar processing procedure and high-precision multimodal annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of 4D radar dataset construction and evaluation, and in particular to a method and device for constructing and evaluating a 4D radar multimodal dataset. Background Art

[0002] With the rapid development of autonomous driving technology in recent years, multi-sensor fusion perception has become a key path to achieving high-level autonomous driving systems. 4D radar, due to its high resolution, strong robustness, and all-weather operation, offers significant advantages over visible light cameras and LiDAR, particularly in complex environments such as rain, snow, fog, and strong sunlight. It is gradually becoming a core sensor in autonomous driving perception systems. To promote the research and implementation of 4D radar perception capabilities, building multimodal datasets, particularly those based on 4D radar, has become a research hotspot. These datasets can effectively support the modeling, training, and evaluation of key tasks such as object detection, point cloud generation, and multimodal fusion.

[0003] Currently, several public 4D radar datasets are being used for autonomous driving perception research. For example, Astyx is the earliest public 4D radar dataset, but its small size makes it difficult to train deep learning models. While RADial expands the data size, it only provides 2D annotations, limiting its application in 3D perception. VoD and TJ4DRadSet introduce annotations for 3D object detection and tracking tasks, but their radar data sources are still based on the manufacturer's default point cloud output, lacking flexible processing capabilities. The K-Radar dataset further alleviates this problem by providing underlying radar data in tensor format. Dual Radar compares the performance of two different radar devices on the same task, indirectly revealing the profound impact of radar information quality on perception performance.

[0004] However, current datasets still have many shortcomings. First, the radar processing pipelines in most datasets are closed or semi-open, making it difficult for researchers to customize the processing and analysis of raw radar data, limiting in-depth research on radar information quality. Second, there is a lack of controllable collection and evaluation mechanisms for different radar transmission modes, frequency settings, and other parameters, making it difficult to systematically analyze the impact of radar configuration on downstream perception. Third, existing datasets also have certain limitations in terms of scene diversity, time synchronization mechanisms, and annotation accuracy, making it difficult to meet the verification requirements of multimodal fusion perception algorithms in complex dynamic environments.

[0005] Therefore, there is an urgent need for a method to solve the problems of closed processing procedures of existing 4D radar datasets and difficulty in evaluating radar information quality. It is also necessary to construct a 4D radar multimodal dataset with an open radar processing process and high-precision multimodal annotation to support in-depth research and evaluation of radar perception performance. Summary of the Invention

[0006] In view of this, the present application provides a 4D radar multimodal dataset construction and evaluation method and device to solve the problems of closed processing flow of existing 4D radar datasets and difficulty in evaluating radar information quality. It constructs a 4D radar multimodal dataset with an open radar processing flow and high-precision multimodal annotation to support in-depth research and evaluation of radar perception performance.

[0007] Specifically, this application is implemented through the following technical solutions:

[0008] In a first aspect, the present application provides a method for constructing and evaluating a 4D radar multimodal dataset, the method comprising:

[0009] Constructing a multi-sensor synchronous acquisition platform; the multi-sensor synchronous acquisition platform includes at least two sensors of different types;

[0010] Collecting multimodal data in multiple acquisition scenarios based on the multi-sensor synchronous acquisition platform to obtain a multimodal data set;

[0011] Performing a point cloud generation test on the multimodal dataset, generating a point cloud using the multimodal dataset, and calculating an average accuracy of the point cloud based on correspondence between predicted points and real points in the point cloud;

[0012] Performing target detection on the multimodal dataset, and calculating the average precision of the multimodal dataset under multiple recall thresholds;

[0013] The performance of the multimodal dataset is evaluated based on the point cloud average precision and average accuracy.

[0014] A second aspect of the present application provides a 4D radar multimodal dataset construction and evaluation device, the device comprising a construction module, an acquisition module, a calculation module, and an evaluation module;

[0015] Wherein, the building module is used to build a multi-sensor synchronous acquisition platform; the multi-sensor synchronous acquisition platform includes at least two different types of sensors;

[0016] The acquisition module is used to acquire multimodal data in multiple acquisition scenarios based on the multi-sensor synchronous acquisition platform to obtain a multimodal data set;

[0017] The computing module is configured to perform a point cloud generation test on the multimodal dataset, generate a point cloud using the multimodal dataset, and calculate an average accuracy of the point cloud based on correspondence between predicted points and real points in the point cloud;

[0018] The calculation module is further configured to perform target detection on the multimodal dataset and calculate the average precision of the multimodal dataset under multiple recall thresholds;

[0019] The evaluation module is used to evaluate the performance of the multimodal dataset based on the point cloud average accuracy and average precision.

[0020] The 4D radar multimodal dataset construction and evaluation method and device provided in this application, on the one hand, through the design of a multi-sensor synchronous acquisition platform, different types of sensor data are integrated to construct a multimodal 4D radar dataset. First, by using different types of sensors, different information dimensions of the scene can be obtained. For example, radar can provide distance and speed information of objects, while cameras provide high-resolution image information, and lidar is better at providing accurate three-dimensional spatial data. Secondly, the synchronous acquisition of sensors enables multimodal data to have consistent timestamps and spatial alignment, thereby avoiding the time and space offset problems in common multi-sensor fusion, and ensuring high quality and high consistency of data. The construction of this dataset not only enhances the diversity of scenes, but also provides support for various changing conditions such as different environments, lighting, weather, etc., effectively improving the robustness of autonomous driving perception systems in complex environments. Through this 4D radar multimodal dataset, more complex and changeable autonomous driving scenarios can be simulated, providing more comprehensive and realistic data support for the training, evaluation and verification of autonomous driving systems. Secondly, to ensure that the constructed 4D radar multimodal dataset meets the requirements of autonomous driving systems, particularly in terms of object detection and point cloud generation accuracy, two evaluation metrics were designed: the average point cloud precision (APP) for the point cloud generation test and the average object detection precision (APP) for the object detection test. First, the point cloud generation test primarily measures the accuracy of the point cloud data generated by the radar system by calculating the match between predicted points and real points in the point cloud. APP provides a direct reflection of the radar system's ranging accuracy and point cloud quality, directly impacting the autonomous driving system's ability to detect and recognize objects. Second, the average object detection precision (APP) evaluates the dataset's object detection performance at multiple recall thresholds. Specifically, by analyzing the variation in precision at different recall thresholds, the dataset's performance in object detection and tracking tasks is assessed. APP helps understand the dataset's responsiveness to varying accuracy requirements, particularly system performance under varying detection accuracy requirements. These two evaluation metrics not only comprehensively assess the dataset's effectiveness in real-world applications but also help optimize data acquisition and processing strategies, thereby better supporting the performance of autonomous driving systems in various scenarios. These two indicators complement each other. The former focuses on the accuracy of point cloud data, while the latter focuses on target detection capabilities. Combining the two can comprehensively and deeply evaluate the quality of the dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Flowchart of the 4D radar multimodal dataset construction and evaluation method provided in Example 1 of the present application;

[0022] Figure 2 This is a structural diagram of the 4D radar multimodal dataset construction and evaluation device provided in Example 2 of the present application. DETAILED DESCRIPTION

[0023] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this application.

[0024] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a," "the," and "the" used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0025] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0026] Specific embodiments are given below to introduce the technical solutions of the present application in detail.

[0027] Example 1

[0028] Figure 1 This is a flowchart of the 4D radar multimodal dataset construction and evaluation method provided in Example 1 of this application. Figure 1 The method provided in this embodiment may include:

[0029] S101. Build a multi-sensor synchronous acquisition platform.

[0030] Specifically, the multi-sensor synchronous acquisition platform is a system platform that integrates multiple different sensor types (such as 4D radar, lidar, visible light camera, infrared camera, inertial navigation unit, etc.). It features a unified time synchronization mechanism and spatial layout design, and is used to simultaneously collect multimodal perception data in the same spatiotemporal environment. The core purpose of the multi-sensor synchronous acquisition platform is to provide a realistic, accurate, and diverse data foundation for tasks such as autonomous driving and intelligent perception. It is particularly suitable for dataset construction and perception algorithm research.

[0031] It should be noted that the multi-sensor synchronous acquisition platform includes at least two different types of sensors. The multi-sensor synchronous acquisition platform includes a closed 4D radar sensor, an open 4D radar sensor, and multiple auxiliary sensors. The closed 4D radar sensor is a point cloud output radar with a closed internal signal processing process, and is used to provide structured point cloud data. The open 4D radar sensor is used to output ADC raw data, radar heat map, and point cloud data, and has an open internal processing process, and perceives 4D radar through a custom signal processing method. The auxiliary sensors include an RGB camera, an infrared camera, a lidar, an inertial measurement unit, and a GPS module. The auxiliary sensors are used to provide image information, depth information, posture information, and geographic reference information, supporting multimodal data acquisition and alignment.

[0032] Specifically, the closed 4D radar sensor (OCULiiRadar) is a commercial-grade radar with a closed signal processing pipeline. It only outputs data in point cloud format, and information from intermediate processing stages cannot be accessed. The closed 4D radar sensor provides structured point cloud data, suitable for traditional point cloud-based object detection and tracking tasks. Due to its relatively stable point cloud, it can be used for algorithm baseline testing and robustness evaluation. The open 4D radar sensor is a development-level sensor with an open signal processing pipeline. It can output ADC raw data, radar heatmaps, and final point clouds, and supports custom processing paths. The open 4D radar sensor supports perception task research starting from the lowest level of raw data, providing exploration space for research on radar information quality and feature extraction methods. The auxiliary sensor module includes an RGB camera and an infrared camera to provide visual modality data, supplementing information in daylight / low-light conditions or high-contrast scenes, respectively, and supporting the capture of color and thermal features. The auxiliary sensor module also includes a lidar, which provides high-precision spatial depth information, serving as a point cloud reference to assist in radar point cloud verification and alignment. The auxiliary sensor module also includes an inertial measurement unit and a GPS module, which are used to obtain the posture changes and geographic positioning information of the multi-sensor synchronous acquisition platform, providing a reference for motion compensation and spatial synchronization.

[0033] The method provided in this embodiment simultaneously deploys a closed 4D radar sensor, an open 4D radar sensor, and multiple auxiliary sensors in a multi-sensor synchronous acquisition platform. This enables full-link perception data acquisition, from low-level signals to high-level semantics, significantly improving the accuracy, robustness, and diversity of environmental perception. The closed 4D radar sensor, with its stable and structured point cloud output, is suitable for training and evaluating baseline perception models, providing highly reliable input for traditional point cloud processing methods. The open 4D radar sensor breaks through the "black box" limitations and outputs multi-level data, including ADC raw signals, radar heat maps, and point clouds. This allows researchers to independently design radar processing pipelines for problems such as target detection, velocity estimation, and reflection feature extraction, thereby studying the impact of different processing strategies on radar information quality and downstream task performance. In addition, RGB cameras provide clear visible light images, supporting semantic recognition and target classification; infrared cameras enhance imaging capabilities at night or in low-light conditions; lidar provides high-precision depth information, a crucial tool for 3D spatial annotation and structural restoration; and inertial measurement units and GPS modules provide high-frequency attitude estimation and precise geographic location, providing the foundation for multi-sensor data alignment and path playback. The integrated deployment of these three sensor types not only expands information complementarity and redundancy between sensors but also supports the development of multimodal fusion algorithms, improving the system's overall performance in target detection, tracking, and recognition in complex scenarios. It also lays a solid foundation for systematic research on the potential of radar perception and the construction of high-quality, multi-dimensionally annotated datasets.

[0034] During implementation, the sensor type is first selected based on the perception task requirements. Next, mounting brackets are designed and fabricated, and each sensor is installed on a data collection vehicle or platform according to the forward perception field of view, ensuring maximum field of view overlap. Furthermore, each sensor is assigned a unique identifier, equipped with a unified power supply and communication interface, and connected via a time synchronization module that supports PPS (Prolonged Precision Sampling) signal synchronization. Finally, a unified data acquisition system is established to write sensor data in real time to a local storage device, generating data records in a unified format.

[0035] S102: Collect multimodal data in multiple acquisition scenarios based on the multi-sensor synchronous acquisition platform to obtain a multimodal data set.

[0036] Specifically, multimodal data refers to data from different sensor types, typically derived from different sensing methods or detection principles. Based on the description above, different sensor types include 4D radar, visible light cameras, infrared cameras, lidar, inertial measurement units, and GPS modules, each collecting different types of information. For example, 4D radar provides radar point cloud data, RGB cameras provide image data, infrared cameras provide temperature difference data, lidar provides point cloud and depth data, and IMUs and GPS provide position and geographic information.

[0037] It's important to note that multimodality means that this data represents different perception modes and may come from different physical quantities (such as optical signals, thermal signals, depth information, and velocity information). By fusing these different types of data, we can fully reflect the characteristics of the environment. The data from these sensors is complementary, compensating for their respective shortcomings and improving the accuracy and robustness of the perception system. Multiple acquisition scenarios refer to data collection in different environments or scenarios, such as varying weather conditions, lighting conditions, road conditions, or vehicle speeds. Each acquisition scenario may present different perception challenges, resulting in different characteristics of the sensor data. By collecting multimodal data from multiple scenarios, we ensure the diversity of the dataset, provide more abundant training samples, and thus improve the generalization ability of the model trained with the dataset to different real-world application environments.

[0038] In a specific implementation, the multimodal data in multiple acquisition scenarios is collected based on the multi-sensor synchronous acquisition platform to obtain a multimodal data set, including:

[0039] (1) Configure the target sensor's emission mode based on the radar point cloud configuration requirements.

[0040] Specifically, radar point cloud configuration requirements refer to the requirements for the quality, density, distribution range, accuracy, and update frequency of radar-generated point cloud data in different application scenarios. For example, in autonomous driving, in order to accurately identify high-speed moving objects in the distance, high-resolution, high-frame-rate point cloud data may be required; while in close-range obstacle detection, more attention may be paid to the density of points and vertical angle coverage. Target sensors generally refer to radar sensors used to collect point cloud data, such as open 4D radar sensors and closed 4D radar sensors. These sensors can configure their operating parameters at the software and hardware levels according to application needs. The emission mode refers to the way the radar sensor emits electromagnetic waves, including but not limited to transmission power, beam shape, frame period, scanning frequency, modulation method (such as FMCW modulation), antenna layout, etc. These parameters directly affect the resolution, range, and accuracy of the point cloud data generated by the radar.

[0041] In specific implementation, the configuration of the target sensor's emission mode based on the radar point cloud configuration requirements includes: determining the target sensor for the emission mode configuration based on the analysis of the requirements for point cloud resolution, speed detection capability and detection distance; obtaining multiple emission modes supported by the target sensor; determining the emission order of each emission mode based on the performance indicators and priorities of different emission modes; determining emission mode parameters based on the emission mode, and configuring the emission mode parameters to the target sensor in sequence based on the emission order to complete the emission mode setting.

[0042] Specifically, the key performance requirements for the radar point cloud, including range resolution, velocity detection capability, and detection range, are first analyzed based on the mission requirements. Based on these requirements, the target sensor to be configured is selected. Furthermore, the target sensor's supported transmission modes, such as Mode 1 through Mode 4 and MixMode, are queried, each with different performance priorities. Subsequently, a transmission priority is set based on the degree of alignment between each transmission mode's performance metrics (such as bandwidth, detection range, and velocity resolution) and the mission requirements, and the transmission order is determined. Next, key transmission parameters for each transmission mode (such as frequency modulation bandwidth, transmission time, pulse count, and TDMA time slot configuration) are extracted from the existing configuration and configured on the sensor control interface according to the transmission order, completing the entire transmission mode setup.

[0043] For example, in one embodiment, an urban autonomous driving test scenario requires high-precision detection of close-range obstacles while also maintaining the ability to monitor pedestrians and vehicles at moderate speeds. In this case, Mode 2, which supports high range resolution, can be selected as the primary transmission mode, supplemented by Mode 3, which offers superior speed detection capabilities. After analysis, Mode 2 is prioritized over Mode 3, and parameters such as the frequency modulation bandwidth (e.g., Mode 2 uses a high bandwidth of 3 GHz, while Mode 3 uses a medium bandwidth of 1.5 GHz), TDMA encoding, and transmission duration are extracted in that order. These parameters are then distributed to the radar sensor, causing it to operate in Mode 2 first, followed by Mode 3, during acquisition. This results in radar point cloud data with both high range and velocity resolution, meeting the required target detection accuracy and responsiveness.

[0044] The method provided in this embodiment, by setting the target sensor's transmission mode based on radar point cloud configuration requirements, can effectively match radar system performance with specific mission requirements, improving the adaptability and quality of radar data in multiple scenarios. Specifically, this configuration method accurately selects transmission modes and sets the transmission sequence based on a comprehensive analysis of point cloud resolution, velocity detection capability, and detection range, enabling the system to flexibly respond to different perception tasks or environmental changes. For example, high-bandwidth, high-resolution transmission modes are prioritized in scenarios requiring precise obstacle identification, while modes with excellent velocity resolution are enabled for high-speed target monitoring. By pre-setting the transmission priority and parameter configuration sequence, the performance potential of each transmission mode can be maximized within the hardware's constraints, avoiding sensor resource waste or information redundancy. Furthermore, automated parameter distribution based on unified configuration logic significantly improves system operational efficiency and deployment flexibility, providing a foundation for building a scalable and customizable radar acquisition system and supporting the stable output of high-quality data for subsequent tasks such as multimodal fusion, point cloud reconstruction, and target detection.

[0045] (2) Parameter calibration and time synchronization are performed on the sensors in the multi-sensor synchronous acquisition platform.

[0046] It should be noted that since the multi-sensor synchronous acquisition platform contains multiple types of sensors, different types of sensors have natural differences in working mechanism, response range, perception accuracy and installation method. These differences will cause deviations in the collected data in dimensions such as spatial position, scale, and response intensity. For example, due to the different physical perception mechanisms of radar, camera and infrared sensors, the information collected has errors in spatial alignment. Even for sensors of the same type, manufacturing errors, installation angles or different usage environments will cause offsets in internal and external parameters. If parameter calibration is not performed, these errors will directly affect the accuracy of data fusion and reduce the performance of downstream tasks such as point cloud reconstruction and target detection. Therefore, parameter calibration is required for multiple types of sensors.

[0047] In a specific implementation, the parameter calibration of the sensors in the multi-sensor synchronous acquisition platform includes: matching the sensors to determine the sensor pair that needs to be extrinsic calibrated; determining the corresponding calibration target for the sensor pair; collecting synchronous data of the sensor pair in the same scene, extracting the characteristic coordinates of the calibration target from the synchronous data, and constructing a corresponding spatial point set; based on the spatial point set, using a nonlinear least squares optimization algorithm to optimize and calculate the extrinsic parameter matrix, the extrinsic parameter matrix represents the spatial transformation relationship between the sensor pairs; performing spatial transformation on the synchronous data based on the extrinsic parameter matrix, and establishing a coordinate transformation relationship between each sensor based on the extrinsic parameter matrix; based on the coordinate transformation relationship, spatially aligning and fusing data from different sensors.

[0048] Specifically, the sensor pairs requiring extrinsic calibration are first determined based on the position, viewpoint, and data features of each sensor in the multi-sensor synchronous acquisition platform, or by matching overlapping areas of images or point clouds. For example, a visible light camera and lidar are selected as a pair for calibration, as they typically acquire data in the same scene and have overlapping viewpoints. Furthermore, suitable calibration targets are selected for each sensor pair. For extrinsic calibration of cameras and lidar, a black and white checkerboard pattern is used as the camera calibration target, and a corner reflector or calibration plate is used as the radar target. Next, synchronized data acquisition is performed within the same scene. A high-precision synchronization mechanism is implemented to ensure consistent timestamps across all sensors and that data is collected at the same time. After acquiring synchronized data from each sensor, the feature coordinates of the calibration targets are extracted. For cameras, image processing algorithms (such as corner detection) are used to extract the coordinates of the checkerboard corners; for lidar, the corresponding coordinates are obtained by extracting feature points from the point cloud. This step generates two sets of spatial points: one in the camera's image coordinate system and the other in the radar's point cloud coordinate system. After obtaining two sets of feature points, a nonlinear least-squares optimization algorithm (such as the Levenberg-Marquardt algorithm) is used to calculate the extrinsic parameter matrix. The extrinsic parameter matrix includes a rotation matrix and a translation vector, representing the spatial transformation relationship between the two sensor coordinate systems. Finally, the optimized extrinsic parameter matrix is ​​used to spatially transform the data obtained from each sensor. Specifically, the camera image coordinates are converted to the radar coordinate system, or the radar point cloud data is converted to the camera coordinate system. After completing the data transformation using the extrinsic parameter matrix, the system establishes a coordinate transformation relationship between each sensor. Based on this calculated coordinate transformation relationship, the data from different sensors (such as camera, radar, lidar, etc.) is aligned and fused.

[0049] It should be noted that because different types of sensors typically have different sampling frequencies, built-in clock accuracy, and data acquisition mechanisms, these sensors may experience time deviations or delays in actual operation. For example, a camera may capture images at a lower frequency, while a radar may scan at a higher frequency. This results in the images and point cloud data they acquire in the same physical scene not being completely synchronized. Without time synchronization, the fusion of data from these different sensors may result in data misalignment, affecting subsequent analysis and processing results. Therefore, time synchronization is necessary for different types of sensors to resolve the time differences between different sensor data and ensure that data from all sensors can be accurately compared and fused at the same time, thereby improving the overall accuracy and reliability of the multi-sensor system.

[0050] In specific implementation, the time synchronization of the sensors in the multi-sensor synchronous acquisition platform includes: resetting the internal clocks of all sensors based on the GPS clock; the timestamps of the reset sensors are consistent; using the pulse signal per second as a reference to adjust the acquisition time of each sensor; the acquisition time of each sensor is consistent after adjustment; during the data acquisition process, based on the scanning characteristics of each sensor, the inherent time offset between sensors is identified; based on the inherent time offset, the acquisition data frame is matched with the latest data frame of the sensor with reference to the acquisition data frame to construct a synchronization frame; all sensor data in the synchronization frame are filtered; the time offset between all sensors in the synchronization frame does not exceed a preset millisecond.

[0051] Specifically, the GPS receiver's clock signal is used to synchronize the internal clocks of all sensors. Using the standard time derived from the GPS clock as a reference, the sensor clocks can be adjusted through a hardware interface to ensure that all sensor clocks are consistent, avoiding time inconsistencies caused by clock drift. During a time reset, each sensor's internal clock is synchronized with the GPS signal, ensuring that the timestamps of data collected by all sensors are consistent. Each second, a pulse signal arrives, and each sensor receives a synchronization signal. This signal serves as a synchronization trigger. Based on the pulse signal, the acquisition clocks of each sensor are adjusted, synchronizing the sampling periods of each sensor. Based on the pulse signal, each sensor's time is synchronized to the same time base. Furthermore, inherent time offsets are detected and analyzed by comparing the timestamps of sensor data frames. By comparing the timestamps of data collected by multiple sensors, time deviations between sensors are identified and quantified. By obtaining the latest data frame from each sensor, the data collected by all sensors are aligned by timestamp, generating a "synchronization frame" containing the synchronized data from all sensors. Each sensor's data frame is matched based on its acquisition time, ensuring that all data frames correspond to the same instant in time. After the synchronization frame is created, the sensor data is filtered. This filtering process can use a simple low-pass filter or a more complex Kalman filter to smooth the synchronized data and eliminate small errors caused by incomplete time synchronization. During the filtering process, to ensure synchronization accuracy and efficient data fusion, a maximum time offset tolerance is set to limit the time offset between any two sensors within the synchronization frame. If the time offset of a sensor exceeds this value, software synchronization will compensate for it and adjust the data timestamp to ensure precise alignment of the final synchronized data.

[0052] The method provided in this embodiment, through parameter calibration and time synchronization for different sensor types, has a crucial impact on the construction and evaluation of the entire multimodal dataset. First, parameter calibration ensures that the coordinate systems of different sensors are precisely aligned, which is crucial for the construction of multimodal datasets. In multimodal datasets, data from different sensor types often need to be fused to form a dataset containing multiple sensory information. If the spatial and pose differences between these sensors are not calibrated, the fused data will increase errors, affecting the quality of the multimodal dataset. Accurate parameter calibration can eliminate these differences, ensuring that the data collected by each sensor can be fused within a unified spatial reference frame, thereby improving the accuracy and consistency of the dataset. In terms of time synchronization, it ensures the consistency of data collected by all sensors along the time axis. In multimodal datasets, data must be aligned not only spatially but also temporally. If the data from different sensors is time-shifted, especially in dynamic environments, the data collected by the sensors will not reflect the scene at the same moment, which will seriously affect the timeliness and authenticity of the dataset. For example, in autonomous driving applications, on-board cameras, radar, and lidar need to work together to generate a real-time, updated environmental perception model. If the time is not synchronized, the system may incorrectly process changes in the scene, resulting in wrong decisions. Through time synchronization, the output of all sensors is unified to the same time base, ensuring that each frame of data accurately reflects the environmental information at the same time point, which not only improves the real-time performance of the dataset, but also improves the reliability of data evaluation. Overall, the implementation of parameter calibration and time synchronization plays a fundamental role in the construction of multimodal datasets, and enables the constructed multimodal datasets to more accurately reflect real-world scenes, improve the quality of the dataset, and ensure the reliability and effectiveness of the dataset in subsequent applications, such as target detection, behavior recognition, environmental perception and other tasks. Through precise calibration and synchronization, multimodal datasets can more realistically reflect the scenes perceived by multiple sensors, making the performance evaluation results more credible and effective.

[0053] (3) Determine a collection scenario based on environmental conditions, and collect multimodal data using the multi-sensor synchronous collection platform in the collection scenario.

[0054] Specifically, environmental conditions include geographic location, time of day, weather conditions, and lighting conditions. Geographic location refers to selecting different road scenarios, such as urban roads, highways, and rural roads. These scenarios significantly impact sensor performance, especially for visual sensors (such as cameras and lidar). Urban roads may present complex traffic conditions, building obstructions, and significant lighting variations. Rural roads, on the other hand, may prioritize natural environments and minimize traffic interference. Time refers to selecting different time periods for data collection, such as daytime and nighttime. Lighting conditions vary significantly between daytime and nighttime, which can affect visual sensor performance. Especially at night, where light levels are low, infrared or thermal imaging sensors may be required for supplemental use. Weather conditions refer to selecting different weather conditions, such as sunny, cloudy, and rainy. Weather conditions have a significant impact on sensors. For example, lidar performs poorly in rainy or foggy conditions, and visual sensors may also be affected by raindrops and fog. Radar sensors, on the other hand, are more stable in inclement weather. Lighting refers to the fact that varying lighting conditions can affect the performance of visual sensors. Especially during periods like morning and dusk, the angle and intensity of sunlight can lead to significant errors in the image information captured by the sensor.

[0055] In specific implementation, first analyze the environmental conditions and match them to each category. After analyzing the environmental conditions, develop scene selection criteria. Based on the different environmental conditions, determine the type and representativeness of the scenes. For example, select scenes with significant environmental variations to ensure dataset diversity. For example, the complexity of urban roads and the homogeneity of highways may represent different scene characteristics. To cover different lighting conditions, data can be collected at different times of the day, including both daytime and nighttime, ensuring that the dataset includes data from sunny days as well as low-light conditions and even darkness. Based on weather conditions, select scenes under different weather conditions, such as sunny and rainy days, to evaluate their impact on sensor performance. Furthermore, select the appropriate sensor type based on the environmental conditions. For example, in low-light or nighttime scenarios, infrared cameras or lidar may be preferred to complement visual sensors, ensuring that data from different sensors can be effectively collected in various environments. In complex weather conditions (such as rain), radar sensors are more effective than visual sensors, so sensors with strong rain and fog penetration capabilities can be selected for data collection. During data collection, the collection strategy can be dynamically adjusted to suit different environmental conditions. For example, by assessing weather changes in real time, if it rains, the collection sensors can be adjusted to increase reliance on radar and infrared sensors and reduce the use of visual sensors to cope with the impact of environmental changes.

[0056] (4) Labeling and classifying the multimodal data to construct a multimodal dataset.

[0057] In a specific implementation, the multimodal data is labeled and classified to construct a multimodal dataset, including: selecting representative sequence data to be labeled based on the diversity of the multimodal data and the complexity of the scene; determining the sampling frequency of each sequence based on the consistency differences of different scenes, and selecting key frames based on the sampling frequency; determining the labeling range of each key frame based on the vehicle's driving environment and the distribution of target objects; and labeling each category of objects with bounding boxes within the labeling range of the key frame. After the labeling is completed, professional labelers conduct multiple rounds of verification to construct a multimodal dataset based on the labeled information.

[0058] Specifically, multimodal data is first analyzed for factors such as location, time, and weather to select data covering a variety of environmental conditions. By comparing the characteristics of sensor data from different environments, representative sequences that cover most driving environments and exhibit scene diversity are selected. Furthermore, the sampling frequency of each sequence is dynamically adjusted based on the characteristics (density) of each scene. For urban roads, where objects are densely packed around the vehicle, a sampling frequency of 3Hz or 5Hz may be appropriate, while for rural roads, a lower sampling frequency is used. Keyframes are periodically selected from each sequence based on the set frequency. Motion analysis and distribution statistics of target objects in the scene, combined with sensor performance, determine an appropriate annotation area. To ensure completeness and coverage, the annotation range should cover objects ahead, while also taking into account the vehicle's width to ensure effective annotation within 100 meters to the left and right. Within the annotation range of each keyframe, objects are annotated and corresponding bounding boxes are generated. Manual annotators extract the position and dimensions of each object from the scene using the lidar point cloud and camera images. Object bounding boxes are annotated using a 3D coordinate system, and the object's motion properties are annotated based on the relative position of the vehicle and the object. For each object, annotators determine its 3D bounding box based on the overlapping areas of sensor data. All annotations are performed within the lidar point cloud data space to ensure consistency. Finally, based on the visual information and semantic understanding of the sensor data, annotators classify the objects and assign corresponding attributes. After the initial annotation, algorithms automatically detect possible errors or deviations. After this initial inspection, other annotators review the annotation results and make corrections.

[0059] S103 , performing a point cloud generation test on the multimodal dataset, generating a point cloud using the multimodal dataset, and calculating an average accuracy of the point cloud based on correspondence between predicted points and real points in the point cloud.

[0060] Specifically, the purpose of the point cloud generation test is to evaluate the accuracy of the multimodal dataset when generating point clouds. Multimodal datasets combine different types of sensor data, such as radar, camera, and lidar. The information provided by these sensors has different characteristics and complementary advantages and disadvantages in the perception scene. Since each sensor has different characteristics, they perceive the environment in different ways. Therefore, this information needs to be fused into accurate point cloud data. The point cloud generation test verifies the effect of the multimodal dataset in generating point clouds by comparing the difference between the generated point cloud and the real point cloud and calculating the average accuracy of the point cloud. Through the point cloud generation test, it is possible to determine how the multimodal data in the multimodal dataset affects the accuracy of the point cloud, providing a basis for subsequent tasks such as 3D target detection and environmental perception, and ensuring that the multi-sensor fusion method can effectively improve perception performance.

[0061] In a specific implementation, the point cloud generation test is performed on the multimodal dataset, a point cloud is generated using the multimodal dataset, and the average accuracy of the point cloud is calculated based on the correspondence between the predicted points and the real points in the point cloud, including: generating a point cloud for the multimodal dataset, obtaining a predicted point set and a real point set from the generated point cloud; for a first point in the predicted point set, based on the distance between the first point and the second point in the real point set, determining a first target point to be included in the predicted point set; for a second point in the real point set, based on the distance between the second point and the first point in the predicted point set, determining a second target point to be included in the real point set; respectively calculating a first number of the first target points and a second number of the second target points; calculating a first ratio of the first number to the number of elements in the predicted point set, calculating a second ratio of the second number to the number of elements in the real point set, and performing an integration operation on the first ratio with respect to the second ratio over a first interval to obtain the average accuracy of the point cloud.

[0062] Specifically, Point Cloud Average Precision (PAP) is a metric that measures the degree of match between the generated point cloud and the ground-truth point cloud and is commonly used to evaluate point cloud generation quality. It quantifies the accuracy of the generated point cloud by comparing the correspondence between the predicted point set (i.e., the point cloud generated from a multimodal dataset) and the ground-truth point set (i.e., the actual measured or annotated point cloud data). PAP reflects the consistency between the generated point cloud and the ground-truth point cloud. A higher PAP value indicates a closer match between the generated point cloud and the ground-truth point cloud, and thus better generation quality.

[0063] In specific implementation, the best matching point is first determined by calculating the distance between each first point in the predicted point set and each second point in the real point set. For each predicted point, the real point with the smallest distance is selected as the matching point, and vice versa. Furthermore, by calculating the number of predicted points and real points that meet the matching conditions, the number of matching points is determined, and the number of target points in the predicted point set and the number of target points in the real point set are respectively obtained. The ratio of the number of target points in the predicted point set to the total number of predicted points (the first ratio) and the ratio of the number of target points in the real point set to the total number of real points (the second ratio) are calculated. Finally, the first ratio and the second ratio are integrated within a certain interval to obtain the average accuracy of the point cloud.

[0064] The calculation process of the average accuracy of the point cloud can be expressed as the following formula:

[0065] ;

[0066] ;

[0067] ;

[0068] Among them, the is the first target point; is the first point in the prediction point set; is the prediction point set; is the second point in the real point set; is a real point set; is the distance between the first point and the second point; is the neighborhood radius; is the second target point; is the average accuracy of the point cloud; is the first number of the first target point; is the number of elements in the prediction point set; is the second number of the second target point; is the number of elements in the real point set.

[0069] The method provided in this embodiment accurately quantifies the quality of a multimodal dataset during the point cloud generation process by calculating the average point cloud accuracy. First, by comparing the predicted point set with the actual point set, the generated point cloud data can intuitively reveal the spatial consistency and matching degree of the point cloud data generated by different sensors. Distance-based calculations during the matching process ensure that each point is correctly classified as a target point in the predicted point set or a target point in the actual point set. This process effectively reflects the relative accuracy of multi-sensor data fusion. Furthermore, by calculating the ratio of the target point to the number of elements in each dataset, the matching degree of each data source can be accurately quantified, thereby preventing the overall point cloud accuracy from being affected by the poor quality of a single sensor data. Finally, the integral operation over the first interval combines the differences in these ratios to provide a more stable and comprehensive point cloud average accuracy. This accuracy assessment method provides a detailed quality analysis of multimodal datasets, not only helping to identify which sensors generate data with higher accuracy, but also optimizing the weight of each sensor during the data fusion process, further improving the overall performance of the dataset. Through this process, more reliable data support can be provided for high-precision perception tasks in fields such as autonomous driving, ensuring the accuracy and consistency of multimodal sensor data in complex environments.

[0070] S104: Perform target detection on the multimodal dataset, and calculate the average precision of the multimodal dataset under multiple recall thresholds.

[0071] Specifically, in applications such as autonomous driving, object detection is a key perception task. Its core function is to accurately identify and locate objects of interest from complex sensor data. The purpose of performing object detection on multimodal datasets is to evaluate their applicability and effectiveness in real-world perception tasks. Object detection typically combines multimodal information, such as images and point cloud data, using algorithms to identify objects such as vehicles, pedestrians, and obstacles, and output their locations and categories.

[0072] Furthermore, the recall threshold is the criterion used to measure whether a model has identified all true objects in a detection task, essentially balancing missed detections with false detections. Different recall thresholds correspond to different tolerance levels, and the model's performance at multiple recall thresholds reflects its stability under varying detection sensitivities. By evaluating object detection performance at multiple recall thresholds, the Average Precision (AP) can be calculated—the average precision at different recall levels—as a comprehensive indicator of detection performance. A higher AP indicates a more accurate model across various object scales, environments, and recall conditions.

[0073] In a specific implementation, target detection is performed on the multimodal dataset, and the average precision of the multimodal dataset under multiple recall thresholds is calculated, including: determining a recall threshold set; the recall threshold set includes multiple recall thresholds; target detection is performed on the multimodal dataset, and the precision corresponding to each recall threshold is calculated according to the detection result; for each recall threshold in the recall threshold set, a target recall threshold whose recall threshold is not less than a preset recall threshold is screened out, the precision corresponding to the target recall threshold is determined, and the maximum precision among the precisions is taken as the target precision; the target precision corresponding to each recall threshold in the recall threshold set is accumulated, and the accumulated value is divided by the number of elements in the recall threshold set to obtain the average precision.

[0074] Specifically, when determining the recall threshold set, in one possible implementation, the recall threshold value range is first set, typically [0, 1]. This range is then evenly divided into several discrete recall thresholds based on the desired precision granularity. For example, if the step size is 0.1, the resulting recall threshold set is {0.0, 0.1, 0.2, …, 1.0}. The recall threshold set can be adjusted based on the precision requirements of the task; the smaller the step size, the more detailed the assessment. In another possible implementation, the recall threshold set is dynamically generated based on the key points of the actual recall rate distribution in the detection results. Furthermore, target detection is performed on the multimodal dataset, and the detection results are obtained and compared with the ground truth annotations. The model's precision at each recall threshold is calculated based on that recall threshold. All detection boxes are traversed, sorted by detection confidence, and matched with the ground truth objects in sequence to obtain the corresponding precision at different recall rates. Then, for each recall threshold in the recall threshold set, among all points where the recall threshold is greater than or equal to the preset recall threshold, the accuracy value with the highest accuracy within the interval is selected as the target accuracy of the recall threshold. Finally, the target accuracy values ​​corresponding to all recall thresholds are accumulated and divided by the number of recall thresholds to obtain the final average accuracy, which is used as an evaluation indicator for the overall performance of target detection. For example, in one embodiment, the recall threshold set can be expressed as: ,The recall threshold set has a step size of 0.025 and contains a total of 40 recall thresholds.

[0075] The calculation process of average precision can be expressed as the following formula:

[0076] ;

[0077] ;

[0078] Among them, the is the target accuracy corresponding to each recall threshold; is the recall threshold; is the preset recall threshold; is the accuracy corresponding to all recall thresholds that are not less than the preset recall threshold; is the average precision; is the number of elements in the recall threshold set.

[0079] The method provided in this embodiment can evaluate the comprehensive performance of target detection models supported by multimodal datasets from a more fine-grained and comprehensive perspective. Specifically, in a multimodal environment, different types of sensors provide a variety of heterogeneous information such as structure, texture, depth, and temperature. Although the fusion of this information enhances the accuracy and robustness of perception, it also increases the complexity and uncertainty of data processing. Therefore, a single evaluation metric is no longer able to accurately measure the actual performance of the detection model. To address this problem, this embodiment first sets a set of recall thresholds. This set is usually divided according to predefined intervals to ensure that all detection sensitivity levels from low recall to high recall are covered. At each recall threshold, the corresponding accuracy is calculated to reveal the model's recognition accuracy for the target at that recall level. Subsequently, by screening all target recall thresholds that are not lower than a preset recall threshold and extracting the maximum accuracy as the representative performance of that threshold, the optimal detection effect at each detection sensitivity is refined to avoid the negative impact of low-quality detection results on the overall evaluation. Finally, the average precision is calculated by averaging the maximum precision corresponding to all target recall thresholds. This metric not only balances detection completeness (via recall) and accuracy (via precision) but also enhances the model's performance stability across different task sensitivities. For multimodal datasets, this evaluation method can not only reflect the actual contribution of different modal fusion strategies to object detection, identifying which modal combinations are most advantageous; it can also expose potential issues caused by multimodal information redundancy or conflict, providing guidance for improving data preprocessing, feature fusion, model design, and other aspects. Furthermore, this method can provide a quantitative basis for model selection in deployment scenarios. For example, in scenarios with high safety requirements, detection accuracy with high recall may be more important, while in resource-constrained scenarios, detection strategies with high precision and low false positive rates may be preferred. Therefore, by analyzing average precision based on multiple recall thresholds, we can more scientifically and systematically explore the true value and potential of multimodal datasets in object detection.

[0080] S105 : Evaluate the performance of the multimodal dataset based on the point cloud average precision and the average accuracy.

[0081] In specific implementations, the point cloud average precision and average accuracy are first normalized using linear normalization or Z-score standardization. Since the numerical ranges and calculation methods of point cloud average precision and average accuracy may differ, to ensure their comparability within a unified evaluation system, both metrics must be normalized to ensure that their contributions to the final evaluation results are dimensionally consistent. Furthermore, appropriate weights are assigned to point cloud average precision and average accuracy based on the specific evaluation objectives or use cases. For example, if the multimodal dataset is primarily used for geometric reconstruction (such as 3D map generation), a higher weight can be assigned to point cloud average precision. If the application focuses more on semantic recognition (such as object detection), a higher weight should be assigned to average accuracy. Once the weights are set, they can be adjusted based on data-driven methods (e.g., tuning using validation set performance). A comprehensive scoring metric is then calculated using a weighted summation based on the set weights. Finally, this comprehensive scoring metric is used to quantitatively evaluate the performance of the multimodal dataset and to compare different multimodal datasets and evaluate performance changes before and after optimization.

[0082] The method provided in this embodiment combines the two metrics, point cloud average precision and mean average accuracy, to evaluate the performance of a multimodal dataset. This method comprehensively reflects the perception capabilities of a multimodal dataset from the two dimensions of geometric structure restoration and semantic recognition. Point cloud average precision reflects the accuracy of spatial structure restoration during point cloud generation based on multimodal information. It assesses the degree of match between the predicted point cloud and the real point cloud, reflecting the dataset's ability to support high-quality point cloud construction. Mean average accuracy is a core metric in object detection, used to measure the accuracy and stability of object detection results at different recall rates, reflecting the dataset's support for multimodal object perception tasks. Combining these two metrics avoids the biased conclusions drawn from a single metric. It can identify whether a dataset possesses detail preservation and high resolution in spatial information representation, and assess its semantic recognition and discrimination capabilities in complex perception tasks. This provides a more scientific and comprehensive basis for comprehensive dataset quality assessment, data augmentation direction selection, perception system design, and model training strategy optimization. This joint evaluation mechanism is particularly important for measuring the gain effect after data fusion, understanding the complementarity between different modalities, and facilitating practical application value in the development of datasets involving new sensors such as 4D radar.

[0083] The method provided in this embodiment, on the one hand, improves data diversity and reliability through multi-sensor collaborative acquisition. Specifically, this embodiment builds a multi-sensor synchronous acquisition platform that integrates a closed 4D radar sensor, an open 4D radar sensor and multiple auxiliary sensors. The closed 4D radar sensor outputs stable structured point cloud data, which is suitable for algorithm baseline testing and robustness evaluation of traditional point cloud processing tasks; the open 4D radar sensor outputs ADC raw data, radar heat map and point cloud data, providing data support for the research of underlying perception tasks. Auxiliary sensors supplement information from different dimensions. Through multi-sensor collaborative acquisition, the dimension of data and information complementarity are greatly expanded, and the accuracy, robustness and diversity of environmental perception are improved. The collected data can cover the full-link perception information from underlying signals to high-level semantics, provide rich samples for the development of multimodal fusion algorithms, make the constructed data set closer to complex real-life scenarios, and improve the generalization ability of the model in different environments.

[0084] Secondly, by optimizing the transmission mode configuration, the requirements of multiple scenarios are met. Specifically, this embodiment configures the target sensor's transmission mode based on the radar point cloud configuration requirements. First, the requirements for point cloud resolution, speed detection capability, and detection range are analyzed to determine the target sensor. For example, in an urban autonomous driving test scenario, an open 4D radar sensor is selected as the target sensor to meet the requirements for high-precision detection of close-range obstacles and monitoring of pedestrians and vehicles at medium speeds. The multiple transmission modes supported by the target sensor are obtained, and the transmission order is determined based on the performance indicators of each transmission mode and the degree of match between the mission requirements. The transmission mode parameters are then determined and configured to the sensor. This method allows the radar operating mode to be flexibly adjusted according to different scenarios, effectively matching radar system performance with specific mission requirements. This avoids sensor resource waste or information redundancy, improves the adaptability and quality of radar data in multiple scenarios, and provides stable and high-quality data support for subsequent tasks such as multimodal fusion, point cloud reconstruction, and target detection, thereby enhancing system operational efficiency and deployment flexibility.

[0085] Third, data quality is ensured through precise calibration and synchronization. Specifically, sensor parameters are calibrated and time synchronized. For parameter calibration, the system identifies the sensor pairs to be calibrated, selects appropriate calibration targets, collects synchronized data, extracts the characteristic coordinates of the calibration targets, and constructs a spatial point set. A nonlinear least-squares optimization algorithm is used to calculate the extrinsic parameter matrix, enabling inter-sensor coordinate transformation and spatial data alignment and fusion. For time synchronization, the internal clocks of all sensors are reset based on the GPS clock. The acquisition time is adjusted using a pulse-per-second signal. Inherent inter-sensor time offsets are identified and compensated for. Synchronized frames are constructed and data is filtered to ensure that sensor time offsets within synchronized frames do not exceed a preset millisecond. Precise calibration and synchronization ensure spatial and temporal consistency of the data, avoiding spatial misalignment and temporal asynchrony during multi-sensor data fusion. This improves the accuracy and timeliness of the dataset, enabling the constructed multimodal dataset to more accurately reflect real-world scenarios and enhances the reliability of data evaluation, providing a reliable data foundation for subsequent dataset-based applications (such as object detection, behavior recognition, and environmental perception).

[0086] Fourthly, innovative evaluation metrics are used to comprehensively assess dataset performance. Specifically, dataset performance is evaluated by calculating the average precision of the point cloud through point cloud generation testing and the average precision under multiple recall thresholds in the object detection task. The average precision of the point cloud is calculated by comparing the predicted point set of the generated point cloud with the true point set, calculating the number of matching points based on distance, and integrating the corresponding ratio. This reflects the degree of match between the generated point cloud and the true point cloud, measuring the quality of the multimodal dataset during the point cloud generation process and revealing the spatial consistency and fusion accuracy of point cloud data from different sensors. In the object detection task, a set of recall thresholds is set, and the precision under each recall threshold is calculated. The maximum precision is selected and taken as the target precision, and the average precision is accumulated to obtain the average precision. This metric evaluates model performance at different detection sensitivity levels, balancing detection completeness and accuracy, and reflects the effectiveness of the multimodal dataset in supporting the object detection task. These two metrics comprehensively reflect the dataset's perception capabilities from the two dimensions of geometric structure restoration and semantic recognition, avoiding the one-sidedness of a single metric. It provides a scientific basis for dataset quality assessment, optimization direction selection (such as determining data enhancement strategies) and model training strategy formulation (such as adjusting model parameters and selecting appropriate modal combinations). Especially in the development of 4D radar multimodal datasets, it helps to deeply understand the data fusion effect and the practical application value of different modalities.

[0087] Example 2

[0088] Corresponding to the aforementioned embodiment of a 4D radar multimodal dataset construction and evaluation method, the present application also provides an embodiment of a 4D radar multimodal dataset construction and evaluation device.

[0089] Figure 2This is a schematic diagram of the structure of the 4D radar multimodal dataset construction and evaluation device provided in Example 2 of this application. Figure 2 , the device provided in this embodiment includes a construction module 210, a collection module 220, a calculation module 230 and an evaluation module 240;

[0090] The construction module 210 is used to construct a multi-sensor synchronous acquisition platform; the multi-sensor synchronous acquisition platform includes at least two different types of sensors;

[0091] The acquisition module 220 is configured to acquire multimodal data in multiple acquisition scenarios based on the multi-sensor synchronous acquisition platform to obtain a multimodal data set;

[0092] The calculation module 230 is configured to perform a point cloud generation test on the multimodal dataset, generate a point cloud using the multimodal dataset, and calculate an average accuracy of the point cloud based on the correspondence between predicted points and real points in the point cloud;

[0093] The calculation module 230 is further configured to perform target detection on the multimodal dataset and calculate the average precision of the multimodal dataset under multiple recall thresholds;

[0094] The evaluation module 240 is configured to evaluate the performance of the multimodal dataset based on the point cloud average precision and the average accuracy.

[0095] The device of this embodiment can be used to perform Figure 1 The steps, specific implementation principles and implementation processes of the method embodiment shown are similar and will not be repeated here.

[0096] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0097] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0098] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A 4D radar multimodal dataset construction and evaluation method, characterized by: The method comprises: Constructing a multi-sensor synchronous acquisition platform; the multi-sensor synchronous acquisition platform includes at least two sensors of different types; Collecting multimodal data in multiple acquisition scenarios based on the multi-sensor synchronous acquisition platform to obtain a multimodal data set; The method comprises the following steps: configuring the emission mode of the target sensor based on the radar point cloud configuration requirements; performing parameter calibration and time synchronization on the sensors in the multi-sensor synchronous acquisition platform; determining an acquisition scenario based on environmental conditions, and using the multi-sensor synchronous acquisition platform to collect multimodal data in the acquisition scenario; annotating and classifying the multimodal data to construct a multimodal dataset; Performing a point cloud generation test on the multimodal dataset, generating a point cloud using the multimodal dataset, and calculating an average accuracy of the point cloud based on correspondence between predicted points and real points in the point cloud; Performing target detection on the multimodal dataset, and calculating the average precision of the multimodal dataset under multiple recall thresholds; wherein, a recall threshold set is determined; the recall threshold set includes multiple recall thresholds; target detection is performed on the multimodal dataset, and the precision corresponding to each recall threshold is calculated based on the detection result; for each recall threshold in the recall threshold set, a target recall threshold whose recall threshold is not less than a preset recall threshold is screened out, the precision corresponding to the target recall threshold is determined, and the maximum precision among the precisions is taken as the target precision; the target precision corresponding to each recall threshold in the recall threshold set is accumulated, and the accumulated value is divided by the number of elements in the recall threshold set to obtain an average precision; The performance of the multimodal dataset is evaluated based on the point cloud average precision and average accuracy.

2. The method according to claim 1, characterized in that The point cloud generation test is performed on the multimodal dataset, a point cloud is generated using the multimodal dataset, and an average accuracy of the point cloud is calculated based on the correspondence between the predicted points and the real points in the point cloud, including: Generating a point cloud for the multimodal data set, and obtaining a predicted point set and a true point set from the generated point cloud; For a first point in the predicted point set, determining a first target point to be included in the predicted point set based on a distance between the first point and a second point in the real point set; For a second point in the real point set, determining a second target point to be included in the real point set based on a distance between the second point and the first point in the predicted point set; respectively calculating a first number of the first target points and a second number of the second target points; A first ratio of the first quantity to the number of elements in the predicted point set is calculated, a second ratio of the second quantity to the number of elements in the actual point set is calculated, and an integration operation is performed on the first ratio with respect to the second ratio over a first interval to obtain an average accuracy of the point cloud.

3. The method according to claim 1, characterized in that Configuring the emission mode of the target sensor based on the radar point cloud configuration requirements includes: Determine the target sensor for the emission mode configuration based on the analysis of point cloud resolution, speed detection capability, and detection distance requirements; Acquire multiple emission modes supported by the target sensor; Determine the launch order of each launch mode based on the performance indicators and priorities of different launch modes; Emission mode parameters are determined based on the emission mode, and the emission mode parameters are sequentially configured to the target sensor based on the emission sequence to complete the emission mode setting.

4. The method according to claim 1, wherein The multi-sensor synchronous acquisition platform includes a closed 4D radar sensor, an open 4D radar sensor and multiple auxiliary sensors; The closed 4D radar sensor is a point cloud output radar with a closed internal signal processing process. The closed 4D radar sensor is used to provide structured point cloud data. The open 4D radar sensor is used to output ADC raw data, radar heat map and point cloud data. The internal processing flow is open and 4D radar perception is achieved through a custom signal processing method. The auxiliary sensors include an RGB camera, an infrared camera, a lidar, an inertial measurement unit and a GPS module. The auxiliary sensors are used to provide image information, depth information, posture information and geographic reference information, and support multimodal data acquisition and alignment.

5. The method according to claim 1, wherein The parameter calibration of the sensors in the multi-sensor synchronous acquisition platform includes: Matching the sensors to determine a sensor pair that requires external parameter calibration; For the sensor pair, determining a corresponding calibration target; Collecting synchronous data of the sensor pair in the same scene, extracting characteristic coordinates of the calibration target from the synchronous data, and constructing a corresponding spatial point set; Based on the spatial point set, a nonlinear least squares optimization algorithm is used to optimize and calculate an extrinsic parameter matrix, where the extrinsic parameter matrix represents a spatial transformation relationship between the sensor pairs; Performing spatial transformation on the synchronization data based on the extrinsic parameter matrix, and establishing a coordinate transformation relationship between the sensors based on the extrinsic parameter matrix; Based on the coordinate transformation relationship, the data from different sensors are spatially aligned and fused.

6. The method according to claim 1, wherein The time synchronization of the sensors in the multi-sensor synchronous acquisition platform includes: Reset the internal clocks of all sensors based on the GPS clock; the timestamps of all sensors after reset are consistent; Using the pulse signal per second as a reference, the acquisition time of each sensor is adjusted; after adjustment, the acquisition time of each sensor is consistent; During data acquisition, the inherent time offset between sensors is identified based on the scanning characteristics of each sensor; In view of the inherent time offset, the collected data frame is used as a reference, and the collected data frame is matched with the latest data frame of the sensor to construct a synchronization frame; All sensor data within the synchronization frame are filtered; and the time offset between all sensors within the synchronization frame does not exceed a preset millisecond.

7. The method according to claim 1, characterized in that The labeling and classification of the multimodal data to construct a multimodal dataset includes: Select representative sequence data to be annotated based on the diversity of the multimodal data and the complexity of the scenarios; Determine the sampling frequency of each sequence based on the consistency differences of different scenes, and select key frames based on the sampling frequency; Determine the annotation range of each key frame based on the vehicle's driving environment and the distribution of target objects; Bounding box annotation is performed on objects of each category within the annotation range of the key frame. After the annotation is completed, multiple rounds of verification are performed by professional annotators to construct a multimodal dataset based on the annotation information.

8. A 4D radar multimodal data set construction and evaluation device, characterized in that: The device is applied to the method according to any one of claims 1 to 7, and the device includes a construction module, a collection module, a calculation module and an evaluation module; Wherein, the building module is used to build a multi-sensor synchronous acquisition platform; the multi-sensor synchronous acquisition platform includes at least two different types of sensors; The acquisition module is used to acquire multimodal data in multiple acquisition scenarios based on the multi-sensor synchronous acquisition platform to obtain a multimodal data set; The method comprises the following steps: configuring the emission mode of the target sensor based on the radar point cloud configuration requirements; performing parameter calibration and time synchronization on the sensors in the multi-sensor synchronous acquisition platform; determining an acquisition scenario based on environmental conditions, and using the multi-sensor synchronous acquisition platform to collect multimodal data in the acquisition scenario; annotating and classifying the multimodal data to construct a multimodal dataset; The computing module is configured to perform a point cloud generation test on the multimodal dataset, generate a point cloud using the multimodal dataset, and calculate an average accuracy of the point cloud based on correspondence between predicted points and real points in the point cloud; The calculation module is further configured to perform target detection on the multimodal dataset and calculate the average precision of the multimodal dataset under multiple recall thresholds; wherein, a recall threshold set is determined; the recall threshold set includes multiple recall thresholds; target detection is performed on the multimodal dataset, and the precision corresponding to each recall threshold is calculated based on the detection result; for each recall threshold in the recall threshold set, a target recall threshold whose recall threshold is not less than a preset recall threshold is screened out, the precision corresponding to the target recall threshold is determined, and the maximum precision among the precisions is taken as the target precision; the target precision corresponding to each recall threshold in the recall threshold set is accumulated, and the accumulated value is divided by the number of elements in the recall threshold set to obtain an average precision; The evaluation module is used to evaluate the performance of the multimodal dataset based on the point cloud average accuracy and average precision.

Citation Information

Patent Citations

  • Millimeter wave radar and vision fused three-dimensional target detection method based on attention mechanism

    CN114708585A

  • Laser radar point cloud ground detection method and device

    CN116413740A