Multi-modal data space alignment method and roadside sensing system
By projecting and adjusting the 3D projection frame data of multimodal data under time synchronization conditions, the problem of spatial alignment of multimodal data is solved, and the synchronous display of multimodal data in the same physical scene is realized, thereby improving the accuracy of environmental perception.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies cannot achieve spatial synchronization of multimodal data, resulting in visual data and radar data being misaligned at the same physical moment, affecting the accuracy of environmental perception.
By acquiring time-synchronized image data and initial point cloud data, a 3D data frame is projected onto the camera imaging plane using a pre-calibrated intrinsic and extrinsic parameter matrix. The projection frame data is then adjusted based on the matching results to achieve spatial alignment between the image data and the point cloud data.
It achieves dual alignment of multimodal data in time and space, ensuring that visual data and radar data are displayed synchronously in the same physical scene, thus improving the accuracy of environmental perception.
Smart Images

Figure CN121639752A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a multimodal data spatial alignment method and a roadside perception system. Background Technology
[0002] With the development of vehicle-road cooperative driving, autonomous driving, and advanced driver assistance systems, multi-sensor fusion technology has become a core means of environmental perception. Among them, the fusion result of data collected by cameras and LiDAR is widely used in target detection and tracking due to its advantages of combining rich texture information and accurate 3D geometric data. However, high-quality multimodal data fusion requires not only temporal synchronization but also spatial synchronization to ensure that the multimodal data pertains to the same scene at the same physical moment.
[0003] Existing technologies offer hardware-based and software-based time synchronization methods; however, these methods can only achieve relatively accurate time synchronization, but cannot achieve effective spatial synchronization. Therefore, how to achieve spatial synchronization of multimodal data is a problem that needs to be solved. Summary of the Invention
[0004] The purpose of this application is to provide a multimodal data spatial alignment method and a roadside perception system, which can achieve dual alignment of multimodal data in time and space, so as to realize the effect of synchronous display of multimodal data in the same physical scene.
[0005] The embodiments of this application are implemented as follows: A first aspect of this application provides a multimodal data spatial alignment method, the method comprising: Acquire time-synchronized image data and initial point cloud data, wherein the initial point cloud data is pre-annotated with a 3D data frame of at least one target object; Acquire at least one candidate point cloud data frame that is temporally adjacent to the initial point cloud data; Based on the pre-calibrated intrinsic and extrinsic parameter matrices, the initial point cloud data and the three-dimensional data frames corresponding to each target object in each candidate point cloud data are projected onto the imaging plane of the camera to obtain the first three-dimensional projection frame data corresponding to each target object in the initial point cloud data and the second three-dimensional projection frame data corresponding to each target object in each candidate point cloud data. Based on the matching results of the first 3D projection frame data and image data corresponding to the target object, and the matching results of the second 3D projection frame data and image data corresponding to the target object, the first 3D projection frame data corresponding to the target object is adjusted to obtain the target 3D projection frame data corresponding to the target object. The initial point cloud data is then adjusted based on the target 3D projection frame data to make the initial point cloud data and image data spatially aligned.
[0006] As one possible implementation, acquiring at least one frame of candidate point cloud data that is temporally adjacent to the initial point cloud data includes: Get the current timestamp of the initial point cloud data; Determine the set of timestamps with the current timestamp as the midpoint, and the number of timestamps in the timestamp set is a preset number; Obtain at least one frame of candidate point cloud data with timestamps from the timestamp set.
[0007] As one possible implementation, based on the matching results of the first 3D projection frame data corresponding to the target object and the image data, and the matching results of the second 3D projection frame data corresponding to the target object and the image data, the first 3D projection frame data corresponding to the target object is adjusted to obtain the target 3D projection frame data corresponding to the target object, including: Object recognition is performed on image data to obtain at least one target object in the image data and the corresponding recognition box data for each target object; Based on the matching results between the first 3D projection frame data and the recognition frame data, determine whether the first 3D projection frame data is accurate; If not, the first three-dimensional projection frame data is adjusted based on the matching results of each second three-dimensional projection frame data and the recognition frame data to obtain the target three-dimensional projection frame data.
[0008] As one possible implementation, the accuracy of the first 3D projection frame data is determined based on the matching result between the first 3D projection frame data and the recognition frame data, including: Calculate the first area intersection-union ratio between the first 3D projection frame data and the recognition frame data; If the first area intersection-union ratio is greater than or equal to a preset threshold, then the data of the first three-dimensional projection frame is determined to be accurate.
[0009] As one possible implementation, the first 3D projection frame data is adjusted based on the matching results between the second 3D projection frame data and the recognition frame data, including: Calculate the second area intersection-union ratio between each second 3D projection frame data and the recognition frame data; The three-dimensional projection frame data to be used is determined based on the second area intersection-union ratio between each second three-dimensional projection frame data and the recognition frame data. The first 3D projection frame data is adjusted based on the 3D projection frame data to be used.
[0010] As one possible implementation, the 3D projection frame data to be used is determined based on the second area intersection-union ratio between each second 3D projection frame data and the recognition frame data, including: The second area intersection-union ratio between each second 3D projection frame data and the recognition frame data is traversed to determine the maximum area intersection-union ratio from multiple second area intersection-union ratios, and the second 3D projection frame data corresponding to the maximum area intersection-union ratio is used as the 3D projection frame data to be used.
[0011] As one possible implementation, the first 3D projection frame data is adjusted based on the 3D projection frame data to be used, including: Replace the first 3D projection frame data with the 3D projection frame data to be used.
[0012] As one possible implementation, object recognition is performed on image data to obtain at least one target object in the image data and the corresponding bounding box data for each target object, including: Based on a preset visual recognition algorithm, object recognition is performed on image data to determine at least one target object in the image data and the location data of each target object. Based on the location data of each target object, determine the bounding box data corresponding to at least one target object in the image data.
[0013] As one possible implementation, the initial point cloud data is adjusted based on the target 3D projection bounding box data to spatially align the initial point cloud data with the image data, including: Based on the 3D data frame of at least one target object in the initial point cloud data, determine the set of point clouds to be corrected corresponding to each target object; Based on the target 3D projection frame data corresponding to each target object, determine the target geometric parameters corresponding to each target object. The target geometric parameters include: center coordinates, size, orientation angle, and motion speed. Based on at least one candidate point cloud data frame that is temporally adjacent to the initial point cloud data, determine the motion result of each target object; Based on the motion results of each target object, the point cloud set to be corrected for each target object is subjected to distortion correction processing to obtain the adjusted point cloud set corresponding to each target object; The adjusted point cloud sets corresponding to each target object are aligned according to the target geometric parameters of each target object so that the initial point cloud data and image data are spatially aligned.
[0014] A second aspect of this application provides a roadside perception system, which includes a controller, a camera, and a lidar device. The camera and lidar device are communicatively connected to the controller, which is used to execute the steps of the multimodal data spatial alignment method described in the first aspect.
[0015] A third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the multimodal data spatial alignment method described in the first aspect above.
[0016] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal data space alignment method described in the first aspect.
[0017] The beneficial effects of the embodiments of this application include: This application provides a multimodal data spatial alignment method. It acquires image data and initial point cloud data under time synchronization conditions, and obtains multiple candidate point cloud data frames temporally adjacent to the initial point cloud data. At least one 3D data frame corresponding to a target object is pre-annotated in the initial point cloud data and the multiple candidate point cloud data frames. Based on the camera's calibration intrinsic and extrinsic parameter matrices, the 3D data frames corresponding to each target object in the initial point cloud data and the 3D data frames corresponding to each target object in each candidate point cloud data are projected onto the camera's 2D imaging plane to obtain first 3D projection frame data corresponding to each target object in the initial point cloud data and second 3D projection frame data corresponding to each target object in each candidate point cloud data. The first 3D projection frame data corresponding to each target object in the initial point cloud data is then aligned with... The bounding box data of each target object in the image data are paired one by one. Then, the second 3D projection box data corresponding to each target object in each candidate point cloud data is paired one by one with the bounding box data of each target object in the image data. Based on the matching results, the 3D projection box data that best matches the bounding box data of each target object in the image data is selected from the first 3D projection box data and multiple second 3D projection box data. The selected 3D projection box data replaces the first 3D projection box data corresponding to each target object to obtain the target 3D projection box data corresponding to each target object. The initial point cloud set corresponding to each target object in the initial point cloud data is adjusted according to the target 3D projection box data corresponding to each target object to obtain the initial point cloud data that is spatially aligned with the image data. The alignment of image data with initial point cloud data is performed under timestamp synchronization. This involves converting 3D point cloud data into 3D projection bounding box data and matching it with 2D image data. Based on the matching results, candidate point cloud data is used to compensate for the initial point cloud data, resulting in accurate 3D projection bounding box data spatially aligned with the image data. Then, the accurate 3D projection bounding box data is used to correct the initial point cloud data, ensuring that each target object in the adjusted initial point cloud data is spatially aligned with each target object in the image data. This achieves dual alignment of multimodal data in both time and space, enabling the synchronous display of multimodal data within the same physical scene. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the structure of a roadside sensing system provided in an embodiment of this application; Figure 2 A flowchart illustrating the first multimodal data spatial alignment method provided in this application embodiment; Figure 3 A flowchart illustrating the second multimodal data spatial alignment method provided in this application embodiment; Figure 4 A flowchart illustrating the third multimodal data spatial alignment method provided in this application embodiment; Figure 5 A flowchart illustrating the fourth multimodal data spatial alignment method provided in this application embodiment; Figure 6 A flowchart illustrating the fifth multimodal data spatial alignment method provided in this application embodiment; Figure 7 A flowchart illustrating the sixth multimodal data spatial alignment method provided in this application embodiment; Figure 8 A flowchart illustrating the seventh multimodal data spatial alignment method provided in this application embodiment; Figure 9 This is a schematic diagram of an existing method for fusion of radar point clouds and target objects based on timestamp synchronization. Figure 10 This is a schematic diagram of an existing camera-target object fusion method based on timestamp synchronization. Figure 11 A schematic diagram illustrating the spatiotemporal fusion of radar point clouds and target objects, provided as an embodiment of this application; Figure 12 A schematic diagram illustrating the spatiotemporal fusion of a camera and a target object, provided as an embodiment of this application; Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0020] Reference numerals: 10: Roadside sensing system; 101: Controller; 102: Camera; 103: LiDAR device; 1301: Memory; 1302: Processor. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0022] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0023] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0024] Currently, both hardware and software time synchronization methods can achieve relatively precise timestamp synchronization between visual and radar data. However, this timestamp-synchronized data fusion scheme cannot guarantee the alignment of visual and radar data corresponding to dynamic targets in the same space. This results in visual and radar data at the same physical moment not being relevant to the same physical scene.
[0025] Furthermore, the time window for a camera to acquire a frame of image data is relatively short, typically around 2ms, while the time window for a LiDAR device to acquire a frame of point cloud data is longer, typically 50-100ms. This means that each frame of image data acquired by the camera is only used to represent the scene within a short time segment, while each frame of point cloud data acquired by the LiDAR device is used to represent the dynamic scene scanned over a longer period. Moreover, the timestamp of each frame of point cloud data acquired by the LiDAR device may be the start time, end time, or midpoint of the frame. This makes it difficult to guarantee that the image data and point cloud data of different dynamic targets in the same space represent the same physical scene, relying solely on the synchronization of timestamps between image data and point cloud data.
[0026] To address this, this application provides a multimodal data spatial alignment method. This method acquires image data and initial point cloud data under time synchronization conditions, and obtains multiple candidate point cloud data that are temporally adjacent to the initial point cloud data. Both the initial and candidate point cloud data are pre-annotated with at least one 3D data frame of a target object. Based on a pre-calibrated intrinsic and extrinsic parameter matrix, the 3D data frames corresponding to each target object in the initial point cloud data and each target object in the candidate point cloud data are projected onto the camera's imaging plane to obtain first 3D projection frame data corresponding to each target object in the initial point cloud data and second 3D projection frame data corresponding to each target object in the candidate point cloud data. Based on the matching results of the first 3D projection frame data and the image data, and the matching results of the second 3D projection frame data and the image data, the target 3D projection frame data corresponding to each target object is determined. The initial point cloud data is then adjusted based on the target 3D projection frame data of each target object to achieve spatial alignment between the initial point cloud data and the image data. This achieves dual alignment of multimodal data in both time and space, enabling the synchronous display of multimodal data within the same physical scene.
[0027] The multimodal data spatial alignment method provided in the embodiments of this application will be explained in detail below with reference to the accompanying drawings.
[0028] Figure 1 See the schematic diagram of a roadside sensing system provided in this application. Figure 1 The roadside perception system 10 provided in this application embodiment includes: a controller 101, at least one camera 102 and at least one lidar device 103. Each camera 102 and each lidar device 103 is communicatively connected to the controller 101. The controller 101 is used to acquire image data uploaded by each camera 102 and point cloud data uploaded by each lidar device 103, and to spatially align the image data and point cloud data under the premise of time synchronization of image data and point cloud data.
[0029] Specifically, the controller 101 receives image data uploaded by each camera 102 and point cloud data uploaded by each LiDAR device 103. Based on the timestamps of the image data and point cloud data, it determines one frame of image data and one frame of point cloud data with the same timestamp. Simultaneously, it acquires multiple candidate point cloud data frames temporally adjacent to the point cloud data corresponding to the current timestamp. Both the point cloud data and each candidate point cloud data frame pre-annotated with 3D data frames of dynamic targets. Further, the 3D data frames are projected onto the camera's imaging plane according to the camera's calibration intrinsic parameter matrix to obtain 3D projection frame data corresponding to each dynamic target in the point cloud data and 3D projection frame data corresponding to each dynamic target in each candidate point cloud. Simultaneously, visual recognition is performed on the image data to obtain recognition frame data for each dynamic target in the image data. Based on the matching results of the 3D projection frame data and recognition frame data of each dynamic target, the accurate 3D projection frame data corresponding to each dynamic target is determined. The current frame point cloud data corresponding to the current timestamp is adjusted based on the accurate 3D projection frame data to obtain point cloud data spatiotemporally aligned with the image data.
[0030] It should be noted that, in order to improve the road environment perception capability of the roadside perception system 10 and thus provide strong data support for intelligent traffic, at least one camera 102 and at least one lidar device 103 are often deployed at each intersection of roads in the urban traffic network to collect scene images of each intersection in real time, and improve the accuracy of prediction of complex dynamic environments based on the spatiotemporal fusion results of the image data of the camera 102 and the point cloud data of the lidar device 103.
[0031] Figure 2 A flowchart illustrating a multimodal data spatial alignment method provided in this application is shown. This method can be applied to the controller 101 in the roadside sensing system 10 described above. See also... Figure 2 This application provides a multimodal data spatial alignment method, including: S201. Acquire time-synchronized image data and initial point cloud data, wherein the initial point cloud data is pre-annotated with a 3D data frame of at least one target object.
[0032] Optionally, time synchronization means that the timestamp of the image data uploaded by the camera is the same as the timestamp of the initial point cloud data uploaded by the LiDAR device. The image data refers to a series of images captured by the camera deployed on the roadside equipment through full pixel exposure. The timestamp of the image data is based on the time when the camera captures the image data. The initial point cloud data refers to a frame of point cloud data with the same timestamp as the image data that is scanned by the LiDAR device deployed on the roadside equipment in a certain time window.
[0033] Optionally, the initial point cloud data collected by the LiDAR device can be annotated with a 3D data frame in advance using a 3D data frame annotation tool. The 3D data frame annotation tool can be any one of SUSTechPOINTS, LabelCloud, Apollo, Custom Tools, etc., and this application does not make any specific limitation on it.
[0034] Specifically, the annotation tool loads the initial point cloud data collected by the LiDAR device, identifies the target objects that need to be annotated in the initial point cloud data, places 3D data frames according to the identified target objects, obtains the 3D data frames corresponding to each target object contained in the initial point cloud data, and adds labels or attribute information to each 3D data frame to distinguish each target object.
[0035] Optionally, a 3D data frame is a cubic bounding box used to label the position of a target object in 3D space. The target object can be a variety of dynamic targets such as pedestrians and vehicles in the physical scene corresponding to the current timestamp. This application does not make any specific limitations on this.
[0036] In the initial point cloud data, each 3D data frame corresponds to a target object, and each 3D data frame contains multiple basic parameters such as center point coordinates, length, width, height, and heading angle. This application does not impose specific limitations on these parameters.
[0037] S202. Obtain at least one frame of candidate point cloud data that is temporally adjacent to the initial point cloud data.
[0038] Optionally, temporal adjacency means that the timestamps of the candidate point cloud data and the timestamps of the initial point cloud data are temporally adjacent. For example, if the timestamp corresponding to the initial point cloud data is 13:00:00, the candidate point cloud data may be multiple frames of point cloud data before and after the timestamp 13:00:00.
[0039] In the multimodal data spatial alignment method, the initial point cloud data is used as the reference point cloud data corresponding to the image data at the current timestamp. The point cloud data of multiple frames adjacent to the initial point cloud data in the time dimension are used as the compensation point cloud data of the initial point cloud data. Based on the alignment result of the initial point cloud data and the image data, the point cloud data of each target object in the initial point cloud data is corrected so that the scene of each target object is consistent in the time dimension and spatial dimension under the multimodal sensor.
[0040] S203. Based on the pre-calibrated intrinsic and extrinsic parameter matrices, project the initial point cloud data and the three-dimensional data frames corresponding to each target object in each candidate point cloud data onto the imaging plane of the camera to obtain the first three-dimensional projection frame data corresponding to each target object in the initial point cloud data and the second three-dimensional projection frame data corresponding to each target object in each candidate point cloud data.
[0041] Optionally, the pre-calibrated intrinsic and extrinsic parameter matrices refer to the camera's calibration intrinsic and extrinsic parameter matrices, which are aligned with the point cloud data acquired by the LiDAR device. These matrices include the camera's intrinsic parameter matrix and extrinsic parameter matrix. The camera's intrinsic parameter matrix describes the camera's internal geometric and optical characteristics, determining how to project a three-dimensional point onto a specific pixel on the camera's two-dimensional image plane based on the camera coordinate system. The camera's extrinsic parameter matrix describes the camera's position and orientation in the world coordinate system, determining how to transform a point in the world coordinate system to the camera coordinate system.
[0042] Optionally, the imaging plane refers to the two-dimensional imaging plane inside the camera, where light is focused after passing through the camera lens to form an image. By projecting the 3D data frames corresponding to each target object in the initial point cloud data onto the camera's imaging plane based on the camera's intrinsic and extrinsic parameter matrices, the first 3D projection frame data corresponding to each target object in the initial point cloud data can be obtained. Simultaneously, by projecting the 3D data frames corresponding to each target object in each candidate point cloud data onto the camera's imaging plane based on the camera's intrinsic and extrinsic parameter matrices, the second 3D projection frame data corresponding to each target object in each candidate point cloud data can be obtained.
[0043] It should be noted that both the first and second 3D projection frame data are 2D bounding box data. Both are formed by projecting the 3D bounding box of the target object in 3D space onto the 2D imaging plane of the camera. Both the first and second 3D projection frame data serve as a bridge connecting 3D point cloud data and 2D image data.
[0044] Specifically, the coordinates of the eight corner points of the 3D data frame corresponding to each target object in the initial point cloud data are obtained, and the initial point cloud data is transformed from the world coordinate system to the camera coordinate system based on the coordinates of the eight corner points of the 3D data frame. Based on the intrinsic and extrinsic parameter matrix of the camera, the 3D point cloud data is projected onto the 2D image plane of the camera to obtain the 2D projection points corresponding to the 3D point cloud data. Based on the minimum bounding rectangle of each 2D projection point set on the 2D image plane and the camera coordinates after the 3D data frame is transformed, the first 3D projection frame data corresponding to each target object in the initial point cloud data is determined. The second 3D projection frame data corresponding to each target object in the candidate point cloud data is consistent with the above projection process, and will not be described again in this application.
[0045] It should be noted that the first and second 3D projection frame data are not used to characterize the importance of the projection frame data, but only to distinguish the 3D projection frame data of each target object in the initial point cloud data and the 3D projection frame data of each target object in the multi-frame candidate point cloud data. This application does not make any specific limitations on this. It should also be noted that the 3D projection frame data is a comprehensive data structure, which includes: the core 3D geometric parameters of the target object, 2D projection information, attribute information, sensor context semantics, physical attributes, quality indicators, and scene context semantics. Among them, the core 3D geometric parameters include: the position, size, and orientation of the target object; the 2D projection information includes: the bounding box and corner points on the image plane; the attribute information includes: the category, confidence level, and tracking information of the target object; the sensor context semantics includes: camera parameters, timestamps, and camera coordinate system; the physical attributes include: the motion state and material properties of the target object; the quality indicators include: quality annotations and consistency verification results; and the scene context semantics includes: environmental information and multi-view association results. This application does not make specific limitations on these aspects.
[0046] S204. Based on the matching results of the first three-dimensional projection frame data and image data corresponding to the target object and the matching results of the second three-dimensional projection frame data and image data corresponding to the target object, adjust the first three-dimensional projection frame data corresponding to the target object to obtain the target three-dimensional projection frame data corresponding to the target object, and adjust the initial point cloud data according to the target three-dimensional projection frame data so that the initial point cloud data and image data are spatially aligned.
[0047] Optionally, the image data can be pre-processed with visual recognition to obtain the target objects contained in the image data and the corresponding recognition box data of each target object. The visual recognition processing can be implemented by the YOLO recognition algorithm, and this application does not make any specific limitations on it.
[0048] The matching result refers to the pairing of the 3D projection bounding box data of each target object in the point cloud data with the recognition bounding box data of each target object in the image data. In other words, there is a one-to-one correspondence between the 3D data frame data and the 2D recognition bounding box data of each target object. It is worth noting that the matching result between the first 3D projection bounding box data and multiple second 3D projection bounding box data of each target object and the recognition bounding box data of each target object in the image data is determined based on the consistency verification results.
[0049] Specifically, the first 3D projection bounding box data corresponding to each target object in the initial point cloud data is paired one-to-one with the recognition bounding box data corresponding to each target object in the image data. At the same time, the second 3D projection bounding box data corresponding to each target object in each candidate point cloud data is paired one-to-one with the recognition bounding box data corresponding to each target object in the image data to obtain the most accurate 3D projection bounding box data corresponding to each target object in the image data. The first 3D projection bounding box data corresponding to each target object in the initial point cloud data is replaced with the most accurate 3D projection bounding box data corresponding to each target object. The point cloud set corresponding to each target object in the initial point cloud data is adjusted according to the replaced first 3D projection bounding box data to obtain point cloud data that is spatially aligned with the image data.
[0050] It is worth noting that the target 3D projection frame data is the most accurate 3D projection frame data corresponding to the target object. The most accurate 3D projection frame data corresponding to the target object refers to the 3D projection frame data with the highest matching degree with the recognition frame data corresponding to the target object in the image data. The most accurate 3D projection frame data corresponding to each target object comes from the first 3D projection frame data and multiple second 3D projection frame data corresponding to each target object. This application does not make specific limitations on this.
[0051] It should also be noted that spatiotemporal alignment refers to the complete overlap of dynamic targets in two-dimensional image data and three-dimensional point cloud data under time synchronization conditions. The fusion result of visual data acquired by the camera and radar point cloud data acquired by the lidar device at the same physical moment represents the same physical scene.
[0052] In this embodiment, image data and initial point cloud data under time synchronization conditions are acquired, and multiple candidate point cloud data frames temporally adjacent to the initial point cloud data are acquired. At least one 3D data frame corresponding to a target object is pre-annotated in the initial point cloud data and the multiple candidate point cloud data frames. Based on the camera's calibration intrinsic and extrinsic parameter matrices, the 3D data frames corresponding to each target object in the initial point cloud data and the 3D data frames corresponding to each target object in each candidate point cloud data are projected onto the camera's 2D imaging plane to obtain first 3D projection frame data corresponding to each target object in the initial point cloud data and second 3D projection frame data corresponding to each target object in each candidate point cloud data. The first 3D projection frame data corresponding to each target object in the initial point cloud data and the second 3D projection frame data corresponding to each target object in the candidate point cloud data are then compared with the data in the image data. The bounding box data of the target objects are paired one by one. The second 3D projection box data corresponding to each target object in each candidate point cloud data is paired one by one with the bounding box data of each target object in the image data. Based on the matching results, the 3D projection box data that best matches the bounding box data corresponding to each target object in the image data is selected from the first 3D projection box data and multiple second 3D projection box data corresponding to each target object. The selected 3D projection box data replaces the first 3D projection box data corresponding to each target object to obtain the target 3D projection box data corresponding to each target object. The initial point cloud set corresponding to each target object in the initial point cloud data is adjusted according to the target 3D projection box data corresponding to each target object to obtain the initial point cloud data that is spatially aligned with the image data. The alignment of image data with initial point cloud data is performed under timestamp synchronization. This involves converting 3D point cloud data into 3D projection bounding box data and matching it with 2D image data. Based on the matching results, candidate point cloud data is used to compensate for the initial point cloud data, resulting in accurate 3D projection bounding box data spatially aligned with the image data. Then, the accurate 3D projection bounding box data is used to correct the initial point cloud data, ensuring that each target object in the adjusted initial point cloud data is spatially aligned with each target object in the image data. This achieves dual alignment of multimodal data in both time and space, enabling the synchronous display of multimodal data within the same physical scene.
[0053] In one alternative implementation, see [link to implementation details]. Figure 3 The specific operation of step S202 above can be as follows: S301. Obtain the current timestamp of the initial point cloud data.
[0054] Optionally, the current timestamp refers to the time identifier of the initial point cloud data annotation of the current frame, and the current timestamp is the same as the timestamp of the image data annotation.
[0055] S302. Determine the set of timestamps with the current timestamp as the midpoint, and the number of timestamps in the timestamp set is a preset number.
[0056] Optionally, the timestamp set refers to a set of multiple time-adjacent timestamps obtained by traversing forward and backward with the current timestamp as the midpoint. For example, if the current timestamp of the initial point cloud data is 13:00:00, the timestamps of the five frames of point cloud data traversed forward along the time axis are 12:59:59, 12:59:58, 12:59:57, 12:59:56, and 12:59:55, and the timestamps of the five frames of point cloud data traversed backward along the time axis are 13:00:01, 13:00:00, ...9 If 02, 13:00:03, 13:00; 04, 13:00:05, then the timestamp set would be {12:59:59, 12:59:58, 12:59:57, 12:59:56, 12:59:55, 13:00:00, 13:00:01, 13:00:02, 13:00:03, 13:00; 04, 13:00:05}, and the timestamp set would contain 11 timestamps. This application does not make a specific limitation on this.
[0057] It should be noted that this application takes the time interval of the timestamps labeled in each frame of point cloud data collected by the lidar device as 1 second as an example, which does not mean that the time interval between the timestamps in each frame of point cloud data collected by the lidar device can only be 1 second. This application does not make any specific limitation in this regard.
[0058] Optionally, the number of timestamps in the timestamp set is a preset number, which is a number of timestamps preset by the user. The preset number can be 5, 10, 15, 20, etc., and this application does not make a specific limitation on it.
[0059] S303. Obtain at least one frame of candidate point cloud data with timestamps from the timestamp set.
[0060] Optionally, the multi-frame point cloud data uploaded by the lidar device can be traversed, and point cloud data with timestamps from the timestamp set can be selected as candidate point cloud data.
[0061] In one alternative implementation, see [link to implementation details]. Figure 4 The operation in step S204 above, "adjusting the first three-dimensional projection frame data corresponding to the target object based on the matching results of the first three-dimensional projection frame data and image data corresponding to the target object, and the matching results of the second three-dimensional projection frame data and image data corresponding to the target object, to obtain the target three-dimensional projection frame data corresponding to the target object," can specifically be as follows: S401. Perform object recognition on the image data to obtain at least one target object in the image data and the recognition box data corresponding to each target object.
[0062] Optionally, the image data is input into a visual recognition tool to perform object recognition on the image data, thereby determining the target object contained in the image data and the corresponding bounding box data of the target object. The bounding box data is the bounding box recognition result of the target object output by the visual recognition tool, and includes information such as the target object's location, category, and execution method; this application does not specifically limit this information.
[0063] S402. Based on the matching result between the first three-dimensional projection frame data and the recognition frame data, determine whether the first three-dimensional projection frame data is accurate.
[0064] Optionally, the accuracy of the first three-dimensional projection frame data is determined based on the matching results between the first three-dimensional projection frame data corresponding to each target object in the initial point cloud data and the recognition frame data corresponding to each target object in the image data.
[0065] Specifically, if the first 3D projection box data corresponding to a target object in the initial point cloud data has a high degree of matching with the recognition box data corresponding to the target object in the image data, then the first 3D projection box data corresponding to the target object is accurate, and there is no need to correct the point cloud set of the target object.
[0066] S403. If not, adjust the first three-dimensional projection frame data according to the matching results of each second three-dimensional projection frame data and the recognition frame data to obtain the target three-dimensional projection frame data.
[0067] Optionally, if the matching degree between the first 3D projection box data corresponding to a target object in the initial point cloud data and the recognition box data corresponding to the target object in the image data is low, then the first 3D projection box data corresponding to the target object is inaccurate. In this case, it is necessary to further determine the matching result between the second 3D projection box data corresponding to the target object in the multi-frame candidate point cloud data and the recognition box data corresponding to the target object in the image data. Based on the matching result, the second 3D projection box data with the highest matching degree to the recognition box data corresponding to the target object in the image data is found, and the first 3D projection box data corresponding to the target object is replaced with the second 3D projection box data to obtain the target 3D projection box data of the target object.
[0068] In one alternative implementation, see [link to implementation details]. Figure 5 The specific operation of step S402 above can be as follows: S501. Calculate the first area intersection-union ratio between the first three-dimensional projection frame data and the recognition frame data.
[0069] Optionally, the first area intersection-union ratio is used to evaluate the matching degree between the first three-dimensional projection frame data corresponding to each target object and the recognition frame data corresponding to each target object. The first area intersection-union ratio refers to the ratio of the intersection area of the first three-dimensional projection frame data corresponding to each target object and the recognition frame data corresponding to each target object to the union area of the first three-dimensional projection frame data corresponding to each target object and the recognition frame data corresponding to each target object.
[0070] S502. If the first area intersection-union ratio is greater than or equal to a preset threshold, then the data of the first three-dimensional projection frame is determined to be accurate.
[0071] Optionally, the preset threshold is a matching degree evaluation threshold set by the user in advance. The preset threshold can be 0.7, 0.8, 0.9, etc., and this application does not make specific limitations on it.
[0072] Optionally, when the first area intersection-union ratio between the first three-dimensional projection box data corresponding to a target object in the initial point cloud data and the recognition box data corresponding to the target object in the image data is greater than a preset threshold, it is determined that the first three-dimensional projection box data corresponding to the target object is accurate, and the three-dimensional point cloud data corresponding to the first three-dimensional projection box data does not need to be adjusted, so that the spatiotemporal alignment of the two-dimensional image and the three-dimensional point cloud of the target object can be achieved.
[0073] In one alternative implementation, see [link to implementation details]. Figure 6 The operation of "adjusting the first three-dimensional projection frame data according to the matching results of each second three-dimensional projection frame data and the recognition frame data" in step S403 above can be specifically as follows: S601. Calculate the second area intersection-union ratio between each second three-dimensional projection frame data and the recognition frame data.
[0074] Optionally, the second area intersection-union ratio is used to evaluate the matching degree between the second three-dimensional projection frame data corresponding to each target object and the recognition frame data corresponding to each target object. The second area intersection-union ratio refers to the ratio of the intersection area of the second three-dimensional projection frame data corresponding to each target object and the recognition frame data corresponding to each target object to the union area of the second three-dimensional projection frame data corresponding to each target object and the recognition frame data corresponding to each target object.
[0075] It is worth noting that the second area intersection-union ratio (IUU) between each second 3D projection box data corresponding to each target object and the recognition box data corresponding to each target object in the image data needs to be calculated separately. Based on the second IUU, the matching degree between each second 3D projection box data corresponding to each target object and the recognition box data corresponding to each target object can be evaluated separately.
[0076] S602. Determine the three-dimensional projection frame data to be used based on the second area intersection-union ratio between each second three-dimensional projection frame data and the recognition frame data.
[0077] Optionally, based on the second area intersection-union ratio between each second three-dimensional projection frame data corresponding to the target object and the recognition frame data corresponding to the target object, the second three-dimensional projection frame data with the highest matching degree to the recognition frame data corresponding to the target object is found, and the second three-dimensional projection frame data is used as the three-dimensional projection frame data to be used corresponding to the target object.
[0078] S603. Adjust the first three-dimensional projection frame data according to the three-dimensional projection frame data to be used.
[0079] Optionally, the first three-dimensional projection frame data corresponding to the target object is replaced with the three-dimensional projection frame data to be used corresponding to the target object to obtain the target three-dimensional projection frame data corresponding to the target object.
[0080] In one optional implementation, step S602 can specifically be performed as follows: The second area intersection-union ratio between each second 3D projection frame data and the recognition frame data is traversed to determine the maximum area intersection-union ratio from multiple second area intersection-union ratios, and the second 3D projection frame data corresponding to the maximum area intersection-union ratio is used as the 3D projection frame data to be used.
[0081] Optionally, the maximum area intersection-union ratio refers to the second area intersection-union ratio with the largest value among the second area intersection-union ratios of multiple second three-dimensional projection box data corresponding to the target object and the recognition box data corresponding to the target object.
[0082] In one optional implementation, step S603 can specifically be performed as follows: Replace the first 3D projection frame data with the 3D projection frame data to be used.
[0083] In one alternative implementation, see [link to implementation details]. Figure 7 The specific operation of step S401 above can be as follows: S701. Based on a preset visual recognition algorithm, perform object recognition on the image data to determine at least one target object in the image data and the position data of each target object.
[0084] Optionally, the preset visual recognition algorithm can be any one of the YOLO visual recognition algorithm, SSD visual recognition algorithm, Faster R-CNN visual recognition algorithm, and Cascade R-CNN visual recognition algorithm, and this application does not make any specific limitation on it.
[0085] Optionally, object recognition is performed on the image data based on a preset visual recognition algorithm to determine the target objects contained in the image data and the position data of each target object. Here, position data refers to the camera coordinate information of the target object in the image data.
[0086] S702. Based on the location data of each target object, determine the recognition box data corresponding to at least one target object in the image data.
[0087] Optionally, based on the location data of each target object in the image data, the corresponding recognition box data of each target object in the image data is determined.
[0088] In one alternative implementation, see [link to implementation details]. Figure 8 The operation of "adjusting the initial point cloud data according to the target 3D projection frame data so that the initial point cloud data is spatially aligned with the image data" in step S204 above can be specifically described as follows: S801. Based on the 3D data frame of at least one target object in the initial point cloud data, determine the set of point clouds to be corrected corresponding to each target object.
[0089] Optionally, the point cloud set to be calibrated refers to the point cloud set corresponding to each target object in the initial point cloud data. The point cloud set to be calibrated corresponding to each target object can be a point cloud set that needs to be adjusted or a correct point cloud set that does not need to be adjusted. This application does not make specific limitations on this.
[0090] It should be noted that the set of point clouds to be calibrated for each target object contains multiple initial point cloud data for each target object.
[0091] Optionally, based on the 3D data frames of each target object marked in the initial point cloud data, the set of point clouds to be corrected corresponding to each target object in the initial point cloud data is determined.
[0092] S802. Based on the target 3D projection frame data corresponding to each target object, determine the target geometric parameters corresponding to each target object. The target geometric parameters include: center coordinates, size, orientation angle, and movement speed.
[0093] Optionally, the target geometric parameters corresponding to each target object refer to the correct geometric parameters corresponding to each target object. The target geometric parameters include: center coordinates, dimensions, orientation angle, and motion speed. Among them, the center coordinates refer to the three-dimensional center coordinates of the target object, the dimensions refer to the accurate length, width, and height of the target object, the orientation angle refers to the actual orientation angle of the target object in three-dimensional space, and the motion speed refers to the motion speed of the target object in three-dimensional space.
[0094] S803. Determine the motion result of each target object based on at least one candidate point cloud data that is temporally adjacent to the initial point cloud data.
[0095] Optionally, by inputting multiple candidate point cloud data frames that are temporally adjacent to the initial point cloud data into a preset motion model, the motion trend of each target object in the time domain corresponding to the multiple candidate point cloud data frames and the initial point cloud data can be determined.
[0096] The motion result refers to the trend of motion change of the target object in the time domain corresponding to multiple frames of candidate point cloud data and the initial point cloud data.
[0097] S804. Based on the motion results of each target object, perform distortion correction processing on the point cloud set to be corrected for each target object to obtain the adjusted point cloud set corresponding to each target object.
[0098] Optionally, distortion correction processing can be performed on the point cloud set to be corrected corresponding to each target object based on the motion trend of each target object to obtain the adjusted point cloud set corresponding to each target object. Here, the adjusted point cloud set refers to the set of distortion-corrected point cloud data corresponding to each target object.
[0099] It is worth noting that by performing distortion correction processing on the radar point cloud data corresponding to each target object based on the motion change trend of each target object, the geometric distortion caused by the asynchronous acquisition time of point cloud data and the relative motion of objects in the physical scene can be corrected, thereby restoring the true three-dimensional geometric shape of the target object at a certain physical moment and providing the target object with accurate instantaneous position and instantaneous attitude.
[0100] S805. Align the adjusted point cloud set corresponding to each target object according to the target geometric parameters corresponding to each target object, so that the initial point cloud data and image data are spatially aligned.
[0101] Optionally, based on the accurate target geometric parameters corresponding to each target object, the distortion-free point cloud data corresponding to each target object is rotated, translated, or otherwise processed to align the 3D point cloud data and image data of each target object in the same space, and the radar point cloud data and image data are used to represent the same physical scene.
[0102] Figure 9 and Figure 10 This is a schematic diagram of the fusion of existing timestamp-synchronized multimodal data. (See attached image) Figure 9 and Figure 10 When the image data acquired by the camera and the point cloud data acquired by the lidar device are only synchronized with the timestamp, the dynamic target displayed by the 3D data frame based on the point cloud data annotation and the image data system may have problems such as recognition error and inaccurate 3D data frame annotation.
[0103] Figure 11 and Figure 12 A schematic diagram of spatiotemporally synchronized multimodal data fusion provided in this application is shown below. Figure 11 and Figure 12 Based on the multimodal data spatial alignment method provided in this application, the three-dimensional data frames corresponding to each target object in the corrected three-dimensional point cloud data are aligned with the recognition box data corresponding to each target object in the image data in the same physical scene, thereby improving the data fusion effect of visual data and radar data.
[0104] See Figure 9 and Figure 11 It is known that synchronization based solely on timestamps cannot achieve alignment between radar point cloud data, 3D projection frames, and dynamic targets in the same scene. However, the 3D point cloud data and 3D projection frames corrected by the multimodal data spatial alignment method provided in this application can be synchronized and aligned with the dynamic targets.
[0105] See Figure 10 and Figure 12 It is known that synchronization based solely on timestamps cannot achieve dynamic following between the 3D data frame and the dynamic target object, resulting in spatial misalignment between the 3D data frame and the dynamic target object. However, the multimodal spatial alignment method provided in this application can not only achieve alignment between radar point cloud data and dynamic targets, but also synchronously achieve alignment between the 3D data frame and the dynamic target object.
[0106] Therefore, the multimodal data spatial alignment method provided in this application can achieve better intelligent cooperative driving. This application can break through the constraints of existing timestamp synchronization and further realize the spatiotemporal synchronization of target objects.
[0107] The following describes the electronic device and computer-readable storage medium used to implement the multimodal data spatial alignment method provided in this application. The specific implementation process and technical effects are described above and will not be repeated below.
[0108] Figure 13 This is a schematic diagram of the structure of an electronic device provided in this application. See also: Figure 13 The electronic device includes a memory 1301 and a processor 1302. The memory 1301 stores a computer program that can run on the processor 1302. When the processor 1302 executes the computer program, it implements the steps in any of the above method embodiments.
[0109] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the various method embodiments described above.
[0110] Optionally, this application also provides a program product, such as a computer-readable storage medium, including a program that, when executed by a processor, performs any of the above-described embodiments of the multimodal data space alignment method.
[0111] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) or processor to execute partial steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0113] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A multi-modal data space alignment method, characterized by, The method comprises: acquiring time-synchronized image data and initial point cloud data, wherein the initial point cloud data is pre-labeled with a three-dimensional data frame of at least one target object; acquiring at least one frame of candidate point cloud data adjacent in time to the initial point cloud data; projecting the three-dimensional data frame of each target object in the initial point cloud data and each candidate point cloud data onto the imaging plane of a camera based on the pre-labeled internal and external parameter matrix to obtain first three-dimensional projection frame data corresponding to each target object in the initial point cloud data and second three-dimensional projection frame data corresponding to each target object in each candidate point cloud data; adjusting the first three-dimensional projection frame data corresponding to the target object according to the matching result of the first three-dimensional projection frame data and the image data and the matching result of each second three-dimensional projection frame data and the image data to obtain target three-dimensional projection frame data corresponding to the target object, and adjusting the initial point cloud data according to the target three-dimensional projection frame data to align the initial point cloud data with the image data in space.
2. The multi-modal data space alignment method of claim 1, wherein, The acquiring of the at least one frame of candidate point cloud data adjacent in time to the initial point cloud data comprises: acquiring the current timestamp of the initial point cloud data; determining a timestamp set with the current timestamp as the midpoint, wherein the number of timestamps in the timestamp set is a preset number; acquiring at least one frame of candidate point cloud data with timestamps in the timestamp set.
3. The multi-modal data space alignment method of claim 1, wherein, The adjusting of the first three-dimensional projection frame data corresponding to the target object according to the matching result of the first three-dimensional projection frame data and the image data and the matching result of each second three-dimensional projection frame data and the image data to obtain target three-dimensional projection frame data corresponding to the target object comprises: performing object recognition on the image data to obtain at least one target object in the image data and recognition frame data corresponding to each target object; determining whether the first three-dimensional projection frame data is accurate according to the matching result of the first three-dimensional projection frame data and the recognition frame data; if not, adjusting the first three-dimensional projection frame data according to the matching result of each second three-dimensional projection frame data and the recognition frame data to obtain target three-dimensional projection frame data.
4. The multi-modal data space alignment method of claim 3, wherein, The determining of whether the first three-dimensional projection frame data is accurate according to the matching result of the first three-dimensional projection frame data and the recognition frame data comprises: calculating a first area intersection ratio between the first three-dimensional projection frame data and the recognition frame data; if the first area intersection ratio is greater than or equal to a preset threshold, it is determined that the first three-dimensional projection frame data is accurate.
5. The multi-modal data space alignment method of claim 3, wherein, The adjusting of the first three-dimensional projection frame data according to the matching result of each second three-dimensional projection frame data and the recognition frame data comprises: calculating a second area intersection ratio between each second three-dimensional projection frame data and the recognition frame data, respectively; determine the three-dimensional projection frame data to be used according to the second area intersection union ratios between each of the second three-dimensional projection frame data and the recognition frame data; adjust the first three-dimensional projection frame data according to the three-dimensional projection frame data to be used.
6. The multi-modal data space alignment method of claim 5, wherein, The method further includes: traversing the second area intersection union ratios between each of the second three-dimensional projection frame data and the recognition frame data to determine a maximum area intersection union ratio from the plurality of second area intersection union ratios, and determining the second three-dimensional projection frame data corresponding to the maximum area intersection union ratio as the three-dimensional projection frame data to be used.
7. The multi-modal data space alignment method of claim 5, wherein, The method further includes: replacing the first three-dimensional projection frame data with the three-dimensional projection frame data to be used.
8. The multi-modal data space alignment method of claim 3, wherein, The method further includes: performing object recognition on the image data based on a preset visual recognition algorithm to determine at least one target object in the image data and position data of each target object; determining the recognition frame data corresponding to each target object in the image data according to the position data of each target object.
9. The multi-modal data space alignment method of claim 1, wherein, The method further includes: determining a to-be-corrected point cloud set corresponding to each target object according to the three-dimensional data frame of at least one target object in the initial point cloud data; determining a target geometric parameter corresponding to each target object according to the target three-dimensional projection frame data of each target object, the target geometric parameter including a center coordinate, a size, an orientation angle, and a motion speed; determining a motion result of each target object according to at least one frame of candidate point cloud data adjacent in time to the initial point cloud data; performing de-distortion processing on the to-be-corrected point cloud set of each target object according to the motion result of each target object to obtain an adjusted point cloud set corresponding to each target object; aligning the adjusted point cloud set corresponding to each target object according to the target geometric parameter corresponding to each target object to align the initial point cloud data and the image data in space.
10. A roadside perception system, comprising: The roadside perception system includes a controller, a camera, and a laser radar device, the camera and the laser radar device are both in communication connection with the controller, and the controller is configured to perform the steps of the multi-modal data spatial alignment method according to any one of claims 1-9.