Multi-mode fusion data labeling method based on pose correction
Through the multi-mode fusion data annotation method based on posture correction, the problem of time misalignment of multi-mode fusion perceived annotation data in the autonomous driving system is solved, and the precise annotation of camera and radar data is realized, which reduces the workload and improves the labeling accuracy.
Patent Information
- Application Number
- CN202510343047.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
In autonomous driving and robotic systems, multi-mode fusion perception requires a large amount of labeled data, but due to data time misalignment, the 3D frame of radar data is difficult to accurately project on camera data, resulting in large quantities of labeling and high difficulty.
The multi-mode fusion data labeling method based on posture correction is adopted. By loading camera data, radar data and path information, the radar data are corrected, the position is reconstructed, the data labeling is performed, and the corrected radar labeling results are restored.
Multi-mode fusion labeling of camera data and radar data is realized, which reduces the workload and solves the problems of misalignment of data corresponding targets and difficulty in labeling.
Smart Images

Figure CN120220151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and more specifically, to a multi-modal fusion data annotation method based on pose correction. Background Art
[0002] In autonomous driving and robotic systems, in order to achieve a high perception and recognition ability, cameras and radars are usually used for multi-modal fusion perception. However, to achieve this perception purpose, a large amount of labeled data is usually required for algorithm training. However, since there is usually a certain time misalignment between the camera and radar data received during the movement of the vehicle or robot, therefore, when actually performing multi-modal fusion annotation, it is difficult to accurately project the 3D box of the radar data onto the camera data. Therefore, relatively complex correction of the camera data annotation results is usually required. However, it is very difficult to label the 3D box itself on the 2D camera image data. Therefore, both the workload and the work difficulty are very large.
[0003] Currently, by manually correcting the misaligned multi-modal annotation data, the following problems exist in manual correction: on the one hand, it is difficult to accurately label the 3D target box on the 2D image; on the other hand, the annotation workload is large.
[0004] Therefore, there is an urgent need for a multi-modal fusion data annotation method based on pose correction. Summary of the Invention
[0005] The object of the present invention is to provide a multi-modal fusion data annotation method based on pose correction to solve the above problems in the prior art, and to be able to more accurately achieve the multi-modal fusion annotation of camera data and radar data.
[0006] The present invention provides a multi-modal fusion data annotation method based on pose correction, which includes:
[0007] Loading camera data, radar data, and path information;
[0008] Correcting the radar data based on the camera data and the path information;
[0009] Performing pose reconstruction based on the camera data, the corrected radar data, and the path information;
[0010] Performing data annotation based on the relative pose of the camera and the radar after pose reconstruction;
[0011] Restoring the radar data according to the data annotation result and the radar data correction relationship.
[0012] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, the camera data, radar data, and path information include consecutive multiple frames of data, and the path information includes at least one of IMU information, RTK information, and GPS information.
[0013] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, correcting the radar data based on the camera data and the path information specifically includes:
[0014] Taking the timestamp of the camera as a reference, correcting the time difference between radar data points according to the camera data, radar data, and path information, so that the states of all radar points are corrected to the same state as the timestamp of the camera.
[0015] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, taking the timestamp of the camera as a reference, correcting the time difference between radar data points according to the camera data, radar data, and path information, so that the states of all radar points are corrected to the same state as the timestamp of the camera, specifically includes:
[0016] Performing high-precision SLAM mapping according to the camera data, radar data, and path information, and correcting the radar data in real time to correct the drifted radar data frame to the same state as the corresponding camera data timestamp.
[0017] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, performing pose reconstruction based on the camera data, the corrected radar data, and the path information specifically includes:
[0018] Based on the camera data, the corrected radar data, and the path information, reconstructing the pose between the camera and the radar so that the relative pose between the camera and the radar remains stationary.
[0019] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, performing pose reconstruction based on the camera data, the corrected radar data, and the path information so that the relative pose between the camera and the radar remains stationary, specifically includes:
[0020] Based on the constructed SLAM mapping relationship, reconstructing the pose relationship between the camera data and the corresponding radar frame data.
[0021] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, performing data annotation based on the relative pose between the camera and the radar after pose reconstruction, specifically includes:
[0022] Perform multi-modal fusion annotation based on the pose relationship between the reconstructed camera data and radar data and the corrected radar data.
[0023] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, the multi-modal fusion annotation based on the pose relationship between the reconstructed camera data and radar data and the corrected radar data specifically includes:
[0024] Utilize the exclusive pose relationship between each frame of radar data after pose reconstruction and the corresponding frame of camera data, project the radar data onto the camera data, and establish the corresponding relationship between the radar data and the camera data.
[0025] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, the restoration of the radar data according to the data annotation result and the radar data correction relationship specifically includes:
[0026] Restore the corrected radar annotation result to the corresponding annotation result before correction.
[0027] The multi-modal fusion data annotation method based on pose correction as described above, wherein, preferably, the restoration of the corrected radar annotation result to the corresponding annotation result before correction specifically includes:
[0028] Map the corrected radar data annotation result into the radar data before correction.
[0029] The present invention provides a multi-modal fusion data annotation method based on pose correction, which performs pose reconstruction based on the corrected radar data and performs data annotation based on the relative pose between the camera and radar after pose reconstruction, can more accurately achieve multi-modal fusion annotation of camera data and radar data; can reduce the workload of camera-radar multi-modal fusion by more than 60%; can solve the problems of misalignment of corresponding targets in camera-radar multi-modal fusion data and the difficulty of camera-radar multi-modal fusion data annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described below in conjunction with the drawings, wherein:
[0031] Figure 1 Is a flowchart of an embodiment of the multi-modal fusion data annotation method based on pose correction provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. The description of the exemplary embodiments is merely illustrative and in no way limits the present disclosure or its application or use. The present disclosure may be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to make the present disclosure thorough and complete and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, the components of materials, numerical expressions and values set forth in these embodiments should be construed as merely exemplary and not as limitations.
[0033] The terms "first", "second", and the like used in the present disclosure do not denote any order, quantity, or importance, but are merely used to distinguish different parts. Words such as "including" or "comprising" mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. Terms such as "upper" and "lower" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0034] In the present disclosure, when it is described that a specific component is located between a first component and a second component, there may or may not be an intermediate component between the specific component and the first component or the second component. When it is described that a specific component is connected to other components, the specific component may be directly connected to the other components without an intermediate component, or may not be directly connected to the other components but have an intermediate component.
[0035] All terms used in the present disclosure (including technical terms or scientific terms) have the same meaning as understood by those of ordinary skill in the art to which the present disclosure pertains, unless otherwise specifically defined. It should also be understood that terms defined in a general dictionary, such as those, should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense, unless specifically defined as such here.
[0036] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.
[0037] As Figure 1 shown, in the actual execution process of the multi-modal fusion data annotation method based on pose correction provided in this embodiment, it specifically includes the following steps:
[0038] Step S1: Load camera data, radar data, and path information.
[0039] Among them, the camera data, radar data, and path information include multiple consecutive frames of data. In an embodiment of the present invention, the radar is a lidar. The path information includes at least one of IMU information, RTK information, and GPS information. It should be noted that the present invention does not specifically limit the type of path information, the number of cameras, and the number of radars.
[0040] Step S2: Correct the radar data based on the camera data and the path information.
[0041] Specifically, based on the time stamp of the camera, correct the time difference between radar data points according to the camera data, radar data, and path information, so that the states of all radar points are corrected to the same state as the time stamp of the camera. Since there is usually a certain time difference between the first radar point and the last radar point during the acquisition of radar data, the radar data itself is misaligned relative to the camera data. Therefore, in step S2, the path information and camera data are combined to correct the radar data points, so that the states of all radar points are corrected to the same state as the camera time stamp.
[0042] In an embodiment of the present invention, high-precision SLAM mapping is performed according to the camera data, radar data, and path information, and the radar data is corrected in real time to correct the drifting radar data frame to the same state as the corresponding camera data time stamp.
[0043] Step S3: Perform pose reconstruction based on the camera data, the corrected radar data, and the path information.
[0044] Specifically, based on the camera data, the corrected radar data, and the path information, reconstruct the pose between the camera and the radar so that the relative pose between the camera and the radar remains stationary. Due to reasons such as jitter and time misalignment caused during actual movement, there is usually a certain difference between the relative pose between the camera and the radar and the pose in the stationary state. Therefore, it is necessary to reconstruct the pose between the camera and the radar to ensure sufficient accuracy.
[0045] In an embodiment of the present invention, based on the constructed SLAM mapping relationship, reconstruct the pose relationship between the camera data and the corresponding radar frame data.
[0046] Step S4: Perform data annotation based on the relative pose of the camera and the radar after pose reconstruction.
[0047] Specifically, multi-modal fusion annotation is performed based on the pose relationship between the reconstructed camera data and radar data and the corrected radar data. In an embodiment of the present invention, using the exclusive pose relationship between each frame of radar data after pose reconstruction and the corresponding each frame of camera data, the radar data is projected onto the camera data, and the corresponding relationship between the radar data and the camera data is established. In a specific implementation, when the data volume is sufficient, inverse mapping annotation from the camera to the radar can be achieved. There is an exclusive pose relationship between each frame of corrected radar data and the corresponding each frame of camera data, and this pose relationship can accurately project the radar data onto the camera data and make them correspond one by one. Therefore, the annotation accuracy and quality are very high.
[0048] Step S5, restore the radar data according to the data annotation result and the radar data correction relationship.
[0049] Specifically, restore the corrected radar annotation result to the corresponding annotation result before correction. In an embodiment of the present invention, map the corrected radar data annotation result to the radar data before correction. Since during the annotation process, the radar data used is the relatively ideal data after correction, rather than the actual real data, in order to achieve a high model training effect, it is necessary to restore the corrected radar annotation result to the corresponding annotation result before correction, so as to realize model training using real data. In the specific implementation process, both the corrected radar data and the radar data before correction can be added to the dataset, which can expand the training data and enhance the algorithm performance.
[0050] The multi-modal fusion data annotation method based on pose correction provided by the embodiments of the present invention performs pose reconstruction based on the corrected radar data, and performs data annotation based on the relative pose between the camera and the radar after pose reconstruction, which can more accurately achieve multi-modal fusion annotation of camera data and radar data; can reduce the workload of camera-radar multi-modal fusion by more than 60%; can solve the problems of misalignment of corresponding targets in camera-radar multi-modal fusion data and the difficulty of camera-radar multi-modal fusion data annotation.
[0051] So far, the embodiments of the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details well known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed here based on the above description.
[0052] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified or some technical features can be equivalently replaced without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A multi-modal fusion data labeling method based on posture correction, characterized in that: include: Load camera data, radar data and path information; modifying the radar data based on the camera data and the path information; Performing posture reconstruction based on the camera data, the corrected radar data and the path information; Data annotation is performed based on the relative pose of the camera and radar after pose reconstruction; The radar data is restored based on the data annotation results and the radar data correction relationship.
2. The multi-modal fusion data labeling method based on posture correction according to claim 1 is characterized in that: The camera data, radar data and path information include continuous multiple frames of data, and the path information includes at least one of IMU information, RTK information and GPS information.
3. The multi-modal fusion data labeling method based on posture correction according to claim 1 is characterized in that: The correcting the radar data based on the camera data and the path information specifically includes: Based on the camera's timestamp, the time difference between radar data points is corrected according to the camera data, radar data, and path information, so that the status of all radar points is corrected to the same status as the camera's timestamp.
4. The multi-modal fusion data labeling method based on posture correction according to claim 3 is characterized in that: The method uses the camera's timestamp as a reference, and corrects the time difference between radar data points according to the camera data, radar data, and path information, so that the states of all radar points are corrected to the same state as the camera's timestamp. Specifically, it includes: High-precision SLAM mapping is performed based on camera data, radar data and path information, and radar data is corrected in real time to correct the drifting radar data frame to the same state as the corresponding camera data timestamp.
5. The multi-modal fusion data labeling method based on posture correction according to claim 1 is characterized in that: The pose reconstruction based on the camera data, the corrected radar data and the path information specifically includes: Based on the camera data, the corrected radar data and the path information, the posture between the camera and the radar is reconstructed so that the relative posture between the camera and the radar remains stationary.
6. The multi-modal fusion data labeling method based on posture correction according to claim 5 is characterized in that: The reconstructing the position and posture between the camera and the radar based on the camera data, the corrected radar data and the path information so that the relative position and posture between the camera and the radar remain stationary specifically includes: Based on the constructed SLAM mapping relationship, the pose relationship between the camera data and the corresponding radar frame data is reconstructed.
7. The multi-modal fusion data labeling method based on posture correction according to claim 1 is characterized in that: The data annotation based on the relative posture of the camera and the radar after posture reconstruction specifically includes: Multi-mode fusion annotation is performed based on the pose relationship between the reconstructed camera data and radar data and the corrected radar data.
8. The multi-modal fusion data labeling method based on posture correction according to claim 7 is characterized in that: The multi-mode fusion annotation based on the reconstructed position and posture relationship of the camera data and the radar data and the corrected radar data specifically includes: By utilizing the exclusive pose relationship between each frame of radar data after pose reconstruction and each corresponding frame of camera data, the radar data is projected onto the camera data, and a corresponding relationship between the radar data and the camera data is established.
9. The multi-modal fusion data labeling method based on posture correction according to claim 1 is characterized in that: The restoring of the radar data according to the data annotation result and the radar data correction relationship specifically includes: The corrected radar annotation results are restored to the corresponding annotation results before correction.
10. The multi-modal fusion data labeling method based on posture correction according to claim 9 is characterized in that: The step of restoring the corrected radar annotation result to the corresponding annotation result before the correction specifically includes: Map the corrected radar data annotation results to the radar data before correction.