Multi-source data high-precision positioning method and system based on multi-modal data
By acquiring ultra-high precision point cloud maps and combining them with INS information and vehicle-mounted camera image features, high-precision positioning of multimodal data was achieved, solving the problems of insufficient positioning accuracy and low efficiency in traditional methods, and providing pixel-level precise positioning and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JISHU TECHNOLOGY (WUHAN) CO LTD
- Filing Date
- 2025-12-24
- Publication Date
- 2026-05-01
AI Technical Summary
Existing high-precision positioning methods based on multimodal data from multiple sources suffer from poor positioning accuracy. In particular, when sensor calibration accuracy and mapping capabilities are limited, image data annotation accuracy is limited, and traditional manual annotation is inefficient.
By acquiring and annotating ultra-high precision point cloud maps, coarse localization is performed by combining initial INS information. The semantic features and image features of images captured by vehicle-mounted cameras are matched to achieve semi-automatic annotation and multi-view image localization. A multi-modal data collaborative localization method is adopted, which includes a combination of a semi-automatic data annotation module, a lidar point cloud localization module, and a multi-view image localization module.
It achieves pixel-level high-precision positioning, improves the robustness of positioning in different scenarios with different sensor configurations and perspectives, simplifies the process and eliminates the need for manual interaction, supports low device computing power requirements, and is suitable for all-in-one machines or cloud deployment.
Smart Images

Figure CN121963200A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of high-precision positioning, and in particular to a multi-source high-precision positioning method and system based on multimodal data. Background Technology
[0002] Autonomous driving technology has made rapid progress thanks to the development of deep learning; however, training deep learning models requires massive amounts of data. Manual annotation methods are time-consuming and labor-intensive. To improve annotation efficiency, many companies have begun researching 4D annotation technology. This involves first mapping the time-series vehicle-mounted LiDAR data, then annotating the entire dataset from a 3D perspective, and finally transferring the annotation to a specific moment in time. However, these methods typically only handle small-scale, continuous time-series data. New regions and time periods require re-mapping and re-annotation, resulting in limited efficiency improvements.
[0003] In addition, image data is usually only used as an auxiliary means of point cloud data annotation. The global point cloud is annotated in 3D view, and then 2D annotation is formed by calibrating and projecting it onto the image. The annotation accuracy is seriously affected by the mapping capability and the accuracy of sensor calibration. Summary of the Invention
[0004] The purpose of this invention is to address the problem of poor positioning accuracy in existing high-precision positioning methods based on multimodal data and multi-source data, and to provide a high-precision positioning method and system based on multimodal data and multi-source data.
[0005] The above-mentioned objective of this application is achieved through the following technical solution: S1: Obtain and annotate an ultra-high precision point cloud map; S2: Based on the initial INS information, retrieve the point cloud data of the ultra-high precision point cloud map and perform coarse positioning to obtain the image 3D positioning information; S3: Based on the images captured by the vehicle-mounted camera, the semantic features of the ultra-high precision point cloud map are retrieved according to the calibration and 3D positioning information of the image; the semantic features and the extracted image features are solved to complete the 3D positioning of the image.
[0006] Optionally, step S1 includes: Annotators perform semi-automatic annotation on ultra-high precision point cloud maps using point and line features rendered by a virtual camera. This includes: The annotator selects the region of interest and calculates the normal to the plane points. The virtual camera is rendered at a preset height along the plane point normal, and the extracted 2D image feature points and lines are obtained. The extracted 2D image feature points and lines are back-projected into 3D space. Clicking on the area near the 2D image feature points completes the semi-automatic annotation of the feature points and lines.
[0007] Optionally, step S1 may also include: the road structure elements to be labeled include: traffic lights, road signs and road markings.
[0008] Optionally, step S2 includes: For the point cloud of consumer-grade automotive LiDAR, the point cloud of the ultra-high precision point cloud map is retrieved based on the initial INS value and registered to complete 3D positioning, including: The initial INS information is filtered to obtain smooth trajectory information; The processed trajectory information is used to perform point cloud registration and frame stacking of key frames, and non-key frames are also mutually registered. Use trajectory information to retrieve ultra-high precision point cloud maps; Using trajectory information as initial values, the keyframes after frame stacking are registered with the ultra-high precision point cloud map to obtain the registered pose and whether the current scene is a failure scene. After all keyframes have been registered, the pose of each frame is optimized using the inter-frame registration information and the absolute pose of the non-failure scene as constraints. Finally, the optimized pose is output, which is the 3D positioning information of the image.
[0009] Optionally, step S3 includes: Using camera calibration and 3D image positioning information as the initial pose values of images captured by the vehicle-mounted camera, the semantic features of the ultra-high precision point cloud map of 3D elements in the scene are retrieved. Extract 2D semantic features from images captured by the vehicle-mounted camera; The semantic features of 2D elements and the semantic features of 3D elements are matched according to a preset threshold to obtain 2D-3D features; By using 2D-3D features, the pose optimization problem is iteratively solved to obtain the final high-precision pose, thus completing the 3D localization of the image.
[0010] A high-precision positioning system based on multi-modal data from multiple sources, the system comprising: a semi-automatic data annotation module, a lidar point cloud positioning module, and a multi-view image positioning module; The semi-automatic data annotation module, the lidar point cloud positioning module, and the multi-view image positioning module are connected in sequence. The semi-automatic data annotation module is used to acquire and annotate ultra-high precision point cloud maps. The lidar point cloud positioning module is used to recall point cloud data of the ultra-high precision point cloud map and perform coarse positioning based on the initial INS information to obtain image 3D positioning information. The multi-view image positioning module is used to retrieve the semantic features of the ultra-high precision point cloud map based on the acquired images captured by the vehicle-mounted camera, according to the calibration and image 3D positioning information; and to complete the 3D positioning of the image by solving the semantic features and the extracted image features.
[0011] An electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to enable the electronic device to perform a multi-source data high-precision positioning method based on multimodal data.
[0012] A computer-readable storage medium storing instructions that, when executed, perform a multi-source data high-precision positioning method based on multimodal data.
[0013] The beneficial effects of the technical solution provided in this application are: Based on a customized ultra-high precision point cloud map, for typical vehicle-mounted LiDAR multi-view image time-series data, the system first retrieves the base map point cloud data of relevant areas based on the coarse positioning data provided by the INS system. Then, it performs coarse registration between the vehicle-mounted LiDAR point cloud and the base map point cloud to obtain the vehicle's position information in the base map coordinate system. Utilizing the precise semantic features annotated on the base map, the image is localized to obtain pixel-level precise pose of the multi-view image in the base map coordinate system. Accuracy: Pixel-level annotation accuracy is achieved through collaborative localization at different granularities based on a customized ultra-high precision point cloud map; Generalization: The hierarchical localization strategy effectively improves the robustness of data localization under different sensor configurations and sensor perspectives in different acquisition scenarios; Efficiency: The localization process is streamlined and fully automated, requiring no manual interaction or prior human input, with low requirements for device computing power, and supports deployment on all-in-one machines or cloud devices. Attached Figure Description
[0014] The present application will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a step diagram of an embodiment of this application; Figure 2 This is a system block diagram in an embodiment of this application; Figure 3 This is a schematic diagram of the electronic device structure in the embodiments of this application. Detailed Implementation
[0015] To provide a clearer understanding of the technical features, objectives, and effects of this application, the specific embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0016] The embodiments of this application provide a high-precision positioning method based on multi-modal data from multiple sources.
[0017] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the steps of a multi-source data high-precision positioning method based on multimodal data in an embodiment of this application, including: S1: Obtain and annotate an ultra-high precision point cloud map; S2: Based on the initial INS information, retrieve the point cloud data of the ultra-high precision point cloud map and perform coarse positioning to obtain the image 3D positioning information; S3: Based on the images captured by the vehicle-mounted camera, the semantic features of the ultra-high precision point cloud map are retrieved according to the calibration and 3D positioning information of the image; the semantic features and the extracted image features are solved to complete the 3D positioning of the image.
[0018] Step S1 includes: As one example, traditional annotation work determines the feature points of the annotated elements by visual inspection. In order to ensure the accuracy of the annotation, the annotation area needs to be enlarged several times, and then the annotator clicks to determine the specific 3D point coordinates, which is inefficient.
[0019] Annotators perform semi-automatic annotation on ultra-high precision point cloud maps using point and line features rendered by a virtual camera. This includes: The annotator selects the region of interest and calculates the normal to the plane points. The virtual camera is rendered at a preset height along the plane point normal, and the extracted 2D image feature points and lines are obtained. The extracted 2D image feature points and lines are back-projected into 3D space. Clicking on the area near the 2D image feature points completes the semi-automatic annotation of the feature points and lines.
[0020] As one implementation, the annotator selects the region of interest, calculates the normal to the planar points, and then renders a virtual camera along the normal at a certain height. Pixel color reflects the intensity of the corresponding point; pixels without point projection are rendered as black. The feature point lines of the rendered image represent the actual geometric intensity edges. The extracted 2D image feature point lines are back-projected into 3D space. The annotator does not need to zoom in on the annotation area to precisely click on 3D points; instead, they click on the area near the feature points to complete the selection of feature points. The entire feature annotation is completed by clicking in a certain sequence. The scene elements obtained from this operation serve as the semantic feature basis for subsequent coarse and fine localization.
[0021] Step S1 also includes: the road structure elements to be labeled include: traffic lights, road signs and road markings.
[0022] Step S2 includes: For the point cloud of consumer-grade automotive LiDAR, the point cloud of the ultra-high precision point cloud map is retrieved based on the initial INS value and registered to complete 3D positioning, including: The initial INS information is filtered to obtain smooth trajectory information; The processed trajectory information is used to perform point cloud registration and frame stacking of key frames, and non-key frames are also mutually registered. Use trajectory information to retrieve ultra-high precision point cloud maps; Using trajectory information as initial values, the keyframes after frame stacking are registered with the ultra-high precision point cloud map to obtain the registered pose and whether the current scene is a failure scene. After all keyframes have been registered, the pose of each frame is optimized using the inter-frame registration information and the absolute pose of the non-failure scene as constraints. Finally, the optimized pose is output, which is the 3D positioning information of the image.
[0023] As one example, due to issues such as point cloud map quality, single-frame laser point cloud quality, inertial navigation cumulative error, jumps, and GPS lock loss, simply fusing these positioning information can lead to vehicle positioning deviations and positioning jumps, failing to provide accurate initial values for subsequent processes.
[0024] Step S3 includes: Using camera calibration and 3D image positioning information as the initial pose values of images captured by the vehicle-mounted camera, the semantic features of the ultra-high precision point cloud map of 3D elements in the scene are retrieved. Extract 2D semantic features from images captured by the vehicle-mounted camera; The semantic features of 2D elements and the semantic features of 3D elements are matched according to a preset threshold to obtain 2D-3D features; By using 2D-3D features, the pose optimization problem is iteratively solved to obtain the final high-precision pose, thus completing the 3D localization of the image.
[0025] As one example, due to the influence of time synchronization, point cloud positioning errors, and camera jitter, simply projecting 3D annotations onto the camera image plane using camera calibration and vehicle positioning information to complete 2D annotation is often inaccurate, requiring precise image positioning. Using camera calibration and vehicle positioning information as initial image pose values, precise base map features composed of high-precision 3D semantics are retrieved; 2D semantic features are extracted from the image, and the 2D and 3D features are projected and matched according to a threshold; the matched 2D-3D feature pairs are used to construct a pose optimization problem iteratively solved to obtain the final high-precision pose.
[0026] A high-precision positioning system based on multi-modal data from multiple sources, the system comprising: a semi-automatic data annotation module, a lidar point cloud positioning module, and a multi-view image positioning module; The semi-automatic data annotation module, the lidar point cloud positioning module, and the multi-view image positioning module are connected in sequence. The semi-automatic data annotation module is used to acquire and annotate ultra-high precision point cloud maps. The lidar point cloud positioning module is used to recall point cloud data of the ultra-high precision point cloud map and perform coarse positioning based on the initial INS information to obtain image 3D positioning information. The multi-view image positioning module is used to retrieve the semantic features of the ultra-high precision point cloud map based on the acquired images captured by the vehicle-mounted camera, according to the calibration and image 3D positioning information; and to complete the 3D positioning of the image by solving the semantic features and the extracted image features.
[0027] This application also discloses an electronic device. (See reference...) Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.
[0028] The communication bus 502 is used to enable communication between these components.
[0029] The user interface 503 may include a display screen, and optionally, the user interface 503 may also include a standard wired interface or a wireless interface.
[0030] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0031] This application also discloses a computer-readable storage medium storing multiple instructions adapted for loading by a processor to execute the above-described high-precision positioning method based on multi-modal data from multiple sources.
[0032] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure.
[0033] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A high-precision positioning method based on multi-modal data from multiple sources, characterized in that, The method includes the following steps: S1: Obtain and annotate an ultra-high precision point cloud map; S2: Based on the initial INS information, retrieve the point cloud data of the ultra-high precision point cloud map and perform coarse positioning to obtain the image 3D positioning information; S3: Based on the images captured by the vehicle-mounted camera, the semantic features of the ultra-high precision point cloud map are retrieved according to the calibration and 3D positioning information of the image; the semantic features and the extracted image features are solved to complete the 3D positioning of the image.
2. The high-precision positioning method based on multi-modal data and multi-source data as described in claim 1, characterized in that, Step S1 includes: Annotators perform semi-automatic annotation on ultra-high precision point cloud maps using point and line features rendered by a virtual camera. This includes: The annotator selects the region of interest and calculates the normal to the plane points. The virtual camera is rendered at a preset height along the plane point normal, and the extracted 2D image feature points and lines are obtained. The extracted 2D image feature points and lines are back-projected into 3D space. Clicking on the area near the 2D image feature points completes the semi-automatic annotation of the feature points and lines.
3. The high-precision positioning method based on multi-modal data and multi-source data as described in claim 2, characterized in that, Step S1 also includes: the road structure elements to be labeled include: traffic lights, road signs and road markings.
4. The high-precision positioning method based on multi-modal data and multi-source data as described in claim 1, characterized in that, Step S2 includes: For the point cloud of consumer-grade automotive LiDAR, the point cloud of the ultra-high precision point cloud map is retrieved based on the initial INS value and registered to complete 3D positioning, including: The initial INS information is filtered to obtain smooth trajectory information; The processed trajectory information is used to perform point cloud registration and frame stacking of key frames, and non-key frames are also mutually registered. Use trajectory information to retrieve ultra-high precision point cloud maps; Using trajectory information as initial values, the keyframes after frame stacking are registered with the ultra-high precision point cloud map to obtain the registered pose and whether the current scene is a failure scene. After all keyframes have been registered, the pose of each frame is optimized using the inter-frame registration information and the absolute pose of the non-failure scene as constraints. Finally, the optimized pose is output, which is the 3D positioning information of the image.
5. The high-precision positioning method based on multi-modal data and multi-source data as described in claim 1, characterized in that, Step S3 includes: Using camera calibration and 3D image positioning information as the initial pose values of images captured by the vehicle-mounted camera, the semantic features of the ultra-high precision point cloud map of 3D elements in the scene are retrieved. Extract 2D semantic features from images captured by the vehicle-mounted camera; The semantic features of 2D elements and the semantic features of 3D elements are matched according to a preset threshold to obtain 2D-3D features; By using 2D-3D features, the pose optimization problem is iteratively solved to obtain the final high-precision pose, thus completing the 3D localization of the image.
6. A high-precision positioning system based on multi-modal data and multi-source data, used to implement the high-precision positioning method based on multi-modal data and multi-source data as described in any one of claims 1-5, characterized in that, The system includes: a semi-automatic data annotation module, a lidar point cloud positioning module, and a multi-view image positioning module; The semi-automatic data annotation module, the lidar point cloud positioning module, and the multi-view image positioning module are connected in sequence. The semi-automatic data annotation module is used to acquire and annotate ultra-high precision point cloud maps. The lidar point cloud positioning module is used to recall point cloud data of the ultra-high precision point cloud map and perform coarse positioning based on the initial INS information to obtain image 3D positioning information. The multi-view image positioning module is used to retrieve the semantic features of the ultra-high precision point cloud map based on the acquired images captured by the vehicle-mounted camera, according to the calibration and image 3D positioning information; and to complete the 3D positioning of the image by solving the semantic features and the extracted image features.
7. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a computer, perform the method as described in any one of claims 1-5.