Positioning method, electronic device, movable platform, storage medium, and program product
By combining multi-frame observation data and GNSS data for semantic map 3D reconstruction and association, the problems of large storage space and large computational load of point cloud maps are solved, and lightweight, high-precision and robust positioning services are achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SZ ZHUOYU TECH CO LTD
- Filing Date
- 2025-10-09
- Publication Date
- 2026-04-30
AI Technical Summary
Point cloud maps consume a lot of storage space and require a large amount of computation in autonomous driving, and their positioning accuracy and robustness are insufficient in complex environments.
By acquiring the first offline 3D semantic map and multi-frame observation data, combined with GNSS data, 3D semantic map reconstruction and association are performed. Nonlinear optimization is then performed using multi-source data to generate lightweight, high-precision, and robust positioning results.
It reduces storage space and computational load, improves positioning accuracy and robustness, adapts to positioning needs under different environments and conditions, and has flexibility and adaptability.
Smart Images

Figure CN2025126603_30042026_PF_FP_ABST
Abstract
Description
Positioning methods, electronic devices, mobile platforms, storage media, and application products
[0001] This application claims priority to Chinese Patent Application No. 202411481026.2, filed on October 22, 2024, entitled “Positioning Method, Electronic Device, Mobile Platform, Storage Medium and Program Product”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of positioning technology, and in particular to a positioning method, electronic device, mobile platform, storage medium, and program product. Background Technology
[0003] Currently, autonomous driving technology is being used more and more widely. Positioning plays a crucial role in autonomous driving because it provides accurate location information to downstream planning modules, enabling these modules to plan accordingly and thus achieve autonomous driving.
[0004] In related technologies, point cloud maps are pre-constructed and stored using LiDAR. During autonomous driving, the results of online LiDAR scanning are matched with the stored point cloud maps to obtain the current location information. However, point cloud maps require a large amount of storage space and computation. Summary of the Invention
[0005] This application provides a positioning method, electronic device, mobile platform, storage medium, and program product to solve the problems of large storage space and high computational load of point cloud maps, and to provide lightweight, high-precision, and robust positioning services in complex environments.
[0006] Firstly, this application provides a positioning method, including:
[0007] Acquire a first offline 3D semantic map and multi-frame observation data; the observation data includes road observation data, self-motion estimation data and Global Navigation Satellite System (GNSS) data corresponding to the same time.
[0008] Obtain a second offline 3D semantic map within the target area from the first offline 3D semantic map; the target area is centered on the location indicated by the GNSS data;
[0009] Semantic map 3D reconstruction is performed based on road observation data and self-motion estimation data from multiple frames of observation data to generate an online 3D semantic map;
[0010] The online elements in the online 3D semantic map are associated with the offline elements in the second offline 3D semantic map to obtain a set of associated data pairs;
[0011] Based on the associated data, nonlinear optimization is performed on the self-motion estimation data in the set and multiple frames of the observation data to determine the positioning result.
[0012] Based on the above technical content, by combining road observation data and self-motion estimation data from multiple frames of observation data, the three-dimensional reconstruction of semantic maps can be accurately performed. By using GNSS data to determine the target range, a second offline three-dimensional semantic map is obtained within the target range. The online three-dimensional semantic map is then associated with the second offline three-dimensional semantic map to provide a basis for nonlinear optimization solutions. This method requires relatively little storage space and computational load, and through multi-source data fusion and nonlinear optimization solutions, it can provide lightweight, high-precision, and robust positioning services in complex environments.
[0013] In some embodiments, acquiring multi-frame observation data includes:
[0014] Acquire road observation data, self-motion estimation data, and GNSS data corresponding to each time point within the sliding window;
[0015] The road observation data, self-motion estimation data, and GNSS data corresponding to the same moment are determined as one frame of observation data to obtain multiple frames of observation data.
[0016] Furthermore, data synchronization was achieved by treating observation data at the same moment as a single frame of observation data.
[0017] In some embodiments, the step of performing nonlinear optimization on the self-motion estimation data in the set and multiple frames of the observation data based on the associated data to determine the positioning result includes:
[0018] Using the positioning results at each time point within the sliding window as the state variables to be estimated, an error function is constructed based on the state variables to be estimated, the associated data set, and the self-motion estimation data in the observation data of multiple frames.
[0019] The error function is solved by nonlinear optimization to obtain the positioning result at the current time.
[0020] Furthermore, by using nonlinear optimization to solve the problem, we can make full use of the self-motion estimation data in the associated data set and multi-frame observation data, construct an error function and optimize it, which can significantly improve the accuracy of the positioning results.
[0021] In some embodiments, the set of associated data pairs includes multiple associated data pairs, and the associated data pairs include location information corresponding to the associated online and offline elements, respectively;
[0022] The step of using the positioning results at each time point within the sliding window as the state variable to be estimated, and constructing an error function based on the state variable to be estimated, the associated data set, and the self-motion estimation data in multiple frames of the observation data, includes:
[0023] For each moment within the sliding window, the difference between the self-motion estimation data of the moment and the previous moment is used as the self-motion estimation constraint, and the difference between the state variable to be estimated at the moment and the previous moment and the self-motion estimation constraint is determined as the motion error term at the moment.
[0024] Using the location information corresponding to the online and offline elements already associated in the associated data pair as observation constraints, for each time point within the sliding window, the positioning result at that time point is used as the state variable to be estimated. The difference between the predicted position of the target element based on the state variable to be estimated and the location information is determined as the observation error term at that time point; the target element is the same element referred to by the online element and the offline element.
[0025] The sum of the motion error term and the observation error term at each moment within the sliding window is determined as the error function.
[0026] Furthermore, by using the difference between the self-motion estimation data at each time step and the previous time step as the self-motion estimation constraint, and determining the difference between the state variable to be estimated at each time step and the self-motion estimation constraint as the motion error term, the self-motion estimation data can be effectively used to constrain the positioning results and reduce accumulated errors. By using the location information corresponding to the online and offline elements already associated in the associated data pair as the observation constraint, and determining the difference between the predicted location of the target element based on the state variable to be estimated and the location information as the observation error term, the location information of the associated data pair can be fully utilized. By comprehensively considering the motion error term and the observation error term, the error function can more comprehensively reflect the actual situation, reduce the impact of single data source error on the positioning results, and thus obtain more accurate positioning results.
[0027] In some embodiments, the step of performing semantic map 3D reconstruction based on road observation data and self-motion estimation data from multiple frames of observation data to generate an online 3D semantic map includes:
[0028] The road observation data and self-motion estimation data from multiple frames of the observation data are input into the motion recovery structure algorithm, and the motion recovery structure algorithm is used to perform semantic map 3D reconstruction to generate an online 3D semantic map.
[0029] Furthermore, multi-frame observation data provides rich perspective and detail information. Combined with self-motion estimation data, it can significantly improve the accuracy and completeness of 3D reconstruction. The motion recovery structure algorithm can effectively recover the 3D structure of the scene and generate a high-precision 3D model through multi-view geometric analysis. Through semantic map reconstruction, not only can the geometric information of the scene be obtained, but also semantic tags (such as roads, buildings, pedestrians, etc.) can be combined to provide a richer environmental understanding.
[0030] In some embodiments, associating online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs includes:
[0031] A graph matching algorithm is used to associate online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs.
[0032] Furthermore, graph matching algorithms can utilize graph information for feature matching, which can more accurately identify and associate corresponding features in online and offline maps compared to using only geometric or attribute information.
[0033] In a second aspect, this application provides an electronic device, including: a processor and a memory communicatively connected to the processor;
[0034] The memory stores computer-executed instructions;
[0035] The processor executes computer execution instructions stored in the memory to implement the positioning method as described in any of the first aspects.
[0036] Thirdly, this application provides a mobile platform, including the electronic device described in the second aspect.
[0037] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the positioning method described in any of the first aspects.
[0038] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the positioning method described in any of the first aspects.
[0039] The positioning method, electronic device, mobile platform, storage medium, and program product provided in this application can accurately reconstruct a three-dimensional semantic map by combining road observation data and self-motion estimation data from multiple frames of observation data. It determines the target range using Global Navigation Satellite System (GNSS) data and acquires a second offline three-dimensional semantic map within that range. The online three-dimensional semantic map is then correlated with the second offline three-dimensional semantic map, providing a basis for nonlinear optimization solutions. This method fully utilizes different types of information (road observation data, self-motion estimation data, and GNSS data) from multiple frames of observation data, achieving effective fusion of multi-source information. This effectively reduces errors that may arise from a single data source, improving the reliability and accuracy of positioning. Because this method relies not only on GNSS data but also incorporates other observation data, it can still achieve effective positioning even when GNSS signals are unstable or unavailable. Through nonlinear optimization solutions, it can be dynamically adjusted and optimized based on actual observation data, adapting to positioning needs under different environments and conditions, exhibiting high flexibility and adaptability. Compared to related technologies, this application requires relatively less storage space and computation, and through multi-source data fusion and nonlinear optimization, it can provide lightweight, high-precision, and robust positioning services in complex environments. Attached Figure Description
[0040] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0041] Figure 1 is a schematic diagram illustrating an application scenario according to an exemplary embodiment;
[0042] Figure 2 is a flowchart illustrating a positioning method according to an exemplary embodiment;
[0043] Figure 3 is a flowchart illustrating a positioning method according to another exemplary embodiment;
[0044] Figure 4 is a schematic diagram illustrating a factor graph structure according to an exemplary embodiment;
[0045] Figure 5 is a schematic diagram illustrating a positioning process according to an exemplary embodiment;
[0046] Figure 6 is a schematic diagram of a positioning device according to an exemplary embodiment;
[0047] Figure 7 is a schematic diagram of the structure of an electronic device according to an exemplary embodiment.
[0048] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0049] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0050] The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. In the following descriptions of embodiments, "a plurality of" means two or more, unless otherwise explicitly defined.
[0051] Currently, autonomous driving technology is being applied more and more widely. Localization plays a crucial role in autonomous driving because it is a higher-level pre-task operation used to provide accurate location information to downstream planning modules, enabling these modules to plan based on this location information and thus achieve autonomous driving. The following section uses a vehicle autonomous driving scenario as an example to introduce localization schemes in related technologies.
[0052] One positioning scheme involves pre-constructing and storing a point cloud map using LiDAR. During autonomous driving, the current location is obtained by matching the results of online LiDAR scanning with the stored point cloud map. However, point cloud maps require significant storage space and computational resources. Another positioning scheme combines a high-precision map with real-time kinematic (RTK) carrier phase differential technology. This scheme first determines the vehicle's lane from the high-precision map based on the precise RTK positioning information, then matches the lane lines detected online by the onboard camera with the lane lines on the high-precision map to obtain the vehicle's positioning information. This scheme relies on satellite signals, resulting in large positioning errors in occluded scenarios and inaccurate positioning. Furthermore, it suffers from large matching errors in scenarios with blurred lane lines, also hindering accurate positioning.
[0053] To address the aforementioned technical issues, this application provides a positioning method, electronic device, mobile platform, storage medium, and program product, aiming to provide better positioning results for downstream planning modules.
[0054] Figure 1 is a schematic diagram illustrating an application scenario according to an exemplary embodiment. As shown in Figure 1, the application scenario includes an electronic device 1. The electronic device 1 is used to control the movement of a mobile platform. Exemplarily, the electronic device 1 is disposed on the mobile platform, or the electronic device 1 is disposed independently, and the electronic device 1 is capable of communicating with the mobile platform. Exemplarily, the electronic device 1 is a control terminal for the mobile platform or other device capable of controlling the movement of the mobile platform, such as a domain controller.
[0055] For example, a mobile platform can be a vehicle, a robotic platform (such as a service robot, an exploration robot, a scientific research robot, etc.), a drone, or other equipment, but is not limited to these.
[0056] In one scenario, during the movement of a mobile platform, electronic device 1 determines the positioning result of the mobile platform by executing the positioning method provided in this application, and controls the movement of the mobile platform based on the positioning result.
[0057] For example, if the mobile platform is a vehicle, this application can be applied to navigation assistance scenarios to help realize intelligent driving functions such as highway navigation and urban navigation.
[0058] The positioning method provided in this application is implemented by a positioning device, which can be integrated into an electronic device 1.
[0059] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0060] Figure 2 is a flowchart illustrating a positioning method according to an exemplary embodiment. As shown in Figure 2, the positioning method provided in this embodiment includes the following steps:
[0061] Step S101: Obtain the first offline 3D semantic map and multi-frame observation data; the observation data includes road observation data, self-motion estimation data and GNSS data corresponding to the same time.
[0062] The first offline 3D semantic map is an offline 3D semantic map corresponding to the target environment, which is the current environment of the mobile platform, such as a city or a street. The first offline 3D semantic map includes feature identifiers and location information for offline elements. Offline elements include at least one of traffic elements and road elements. Feature identifiers are used to uniquely identify offline elements to distinguish different elements; for example, feature identifiers may be names or numbers. Traffic elements include at least one of traffic signs, traffic lights, and streetlights. Road elements include at least one of lane lines, road edges, turning signs (such as arrows or signs on the road surface), stop lines, and pedestrian crossings. Location information is used to represent the 3D position of the corresponding element in the target environment. For example, the 3D position is represented by coordinates in a map coordinate system, which can be a global coordinate system or a relative coordinate system. Optionally, the first offline 3D semantic map may also include other information, such as the category of offline elements and road topology information; this application does not limit this.
[0063] Optionally, in this embodiment, a first offline three-dimensional semantic map is pre-constructed and stored, or a high-precision map is downloaded as the first offline three-dimensional semantic map and stored, so that the first offline three-dimensional semantic map can be obtained during positioning.
[0064] For example, constructing the first offline 3D semantic map is specifically implemented as follows: 3D data of the target environment is collected using sensor devices; data from different sensor devices is integrated using multi-sensor fusion technology to obtain more comprehensive and accurate 3D data; the 3D data is preprocessed; features are extracted from the preprocessed 3D data, and semantic segmentation is performed on the 3D data using deep learning algorithms (such as convolutional neural networks), classifying and labeling different categories of objects (such as traffic elements, road elements, etc.); based on the 3D data and semantic segmentation results, a 3D model of the target environment is constructed using a 3D reconstruction algorithm; the semantic segmentation results are combined with the 3D model to generate a 3D semantic map with semantic annotations. The sensor devices include at least one of LiDAR, camera, and Inertial Measurement Unit (IMU), but are not limited to these. Preprocessing can include operations such as denoising, correction, enhancement, and filtering to improve data quality. The 3D reconstruction algorithm can be a Simultaneous Localization and Mapping (SLAM) algorithm, but is not limited to these.
[0065] The multi-frame observation data includes observation data corresponding to each time point within the sliding window, with one time point corresponding to one frame of observation data. The sliding window is a window of fixed length that can slide in one direction; that is, the maximum number of frames of observation data that the sliding window can acquire is fixed, for example, a maximum of 5, 8, or 10 frames. It should be noted that during the sliding process, each time a new frame of observation data is acquired, the oldest frame is removed to ensure that the number of frames in the sliding window remains constant.
[0066] The road observation data includes at least one of traffic element observation data and road element observation data.
[0067] Traffic element observation data includes the element identifier, category, and location information of traffic elements. The element identifier is used to uniquely identify traffic elements to distinguish them from different elements. Categories can be traffic signs, traffic lights, or lampposts, etc. Location information represents the two-dimensional location of the traffic element. For example, if the traffic element observation data is identified from an image, the location information is the two-dimensional pixel coordinates of the traffic element in the image. For example, an image acquisition device (such as a camera or webcam) is installed on a mobile platform. While the mobile platform is in motion, electronic equipment controls the image acquisition device to periodically acquire images and performs image recognition and feature extraction on these images to obtain the traffic element observation data.
[0068] Road element observation data includes element identifiers, categories, attribute information, and location information for road elements. Element identifiers are used to uniquely identify traffic elements to distinguish them from different elements. Categories can be lane lines, road edges, turning signs (such as arrows or signs on the road surface), stop lines, or pedestrian crossings, etc. Attribute information describes the attributes of road elements. For example, if the category of a road element is a lane line, the attribute information is solid line or dashed line; another example is a turning sign, where the attribute information is U-turn, left turn, right turn, or straight ahead. Location information represents the two-dimensional position of the road element. For example, if the road element observation data is periodically acquired using Bird's Eye View (BEV) detection technology, then the location information is the two-dimensional coordinates of the road element in the BEV coordinate system.
[0069] The ego-motion estimation data refers to motion data acquired periodically using an ego-motion estimation algorithm. For example, the motion data includes position, attitude, velocity, and acceleration. Position refers to relative position (the change in position of the movable platform relative to its starting point); attitude includes attitude angles and / or attitude matrices; velocity includes linear velocity and / or angular velocity; and acceleration includes linear acceleration and / or angular acceleration. Optionally, ego-motion estimation can be implemented using data from various sensors, including visual sensors (such as cameras), IMUs, and wheel speed sensors.
[0070] GNSS data includes the location information of a mobile platform in a global coordinate system, which includes longitude, latitude, and altitude coordinates. For example, a GNSS receiving device is installed on the mobile device to receive satellite signals. The electronic device periodically acquires the satellite signals received by the GNSS receiving device and processes the satellite signals to obtain GNSS data.
[0071] Step S102: Obtain a second offline three-dimensional semantic map within the target area from the first offline three-dimensional semantic map; the target area is centered on the location indicated by the GNSS data.
[0072] The first offline 3D semantic map covers a relatively large environmental range, but for positioning, the semantic map within the current location of the mobile platform is sufficient as a reference. Therefore, a portion of the semantic map can be extracted from the first offline 3D semantic map to reduce the computational load of subsequent operations. Furthermore, the location indicated by GNSS data is relatively accurate and can represent the current location of the mobile platform. Therefore, the target range can be defined as the area centered on the location indicated by the GNSS data and with a preset distance as its radius. A second offline 3D semantic map within the target range is then obtained from the first offline 3D semantic map. The target range can be considered as the current location of the mobile platform. The preset distance can be set according to actual needs; this application does not limit it, for example, a preset distance of 50 meters or 100 meters.
[0073] It should be noted that the location information in GNSS data is in the global coordinate system, while the location information of features in the first offline 3D semantic map is in the map coordinate system. Therefore, before determining the target range, it is necessary to convert the GNSS data to GNSS data in the map coordinate system in order to unify the coordinate systems.
[0074] Step S103: Based on the road observation data and self-motion estimation data in the multi-frame observation data, perform semantic map 3D reconstruction to generate an online 3D semantic map.
[0075] The online 3D semantic map includes feature identifiers and location information for online elements. Online elements include at least one of traffic elements and road elements. Feature identifiers are used to uniquely identify online elements to distinguish them from different elements. Location information indicates the 3D position of the online elements in the map coordinate system.
[0076] Optionally, a preset 3D reconstruction algorithm is used to perform semantic map 3D reconstruction on road observation data and self-motion estimation data from multiple frames of observation data to obtain an online 3D semantic map. The preset 3D reconstruction algorithm can be set according to actual needs, and this application does not limit it; examples include Structure from Motion (SfM) algorithm, SLAM algorithm, and deep learning-based 3D reconstruction algorithms.
[0077] Step S104: Associate the online features in the online 3D semantic map with the offline features in the second offline 3D semantic map to obtain a set of associated data pairs.
[0078] The associated data pair set includes multiple associated data pairs. Each associated data pair includes the feature identifier, category, and location information corresponding to the associated online and offline features, respectively.
[0079] The online 3D semantic map is obtained by reconstructing the semantic map from road observation data and self-motion estimation data in multi-frame observation data. The online 3D semantic map can be regarded as the semantic map within the current range of the mobile platform, while the second offline 3D semantic map is the offline 3D semantic map within the target range. The target range can be regarded as the current range of the mobile platform. Therefore, it is very likely that there are elements in these two semantic maps that correspond to the same thing in the real scene.
[0080] Optionally, if the distance between an online element and an offline element is less than a preset distance threshold and they belong to the same category, they are very likely to correspond to the same thing in a real scene. The online element and the offline element can be considered a pair of elements with a relationship, and the element identifier, category, and location information corresponding to the online element and the offline element can be used to form a related data pair. The distance between the two elements can be determined based on the element's location information. For example, taking any offline element in the second offline 3D semantic map as an example, the distance between the two elements is determined based on the location information of the offline element and the location information of online elements in the online 3D semantic map that belong to the same category as the offline element. The preset distance threshold can be set according to actual needs, and this application does not limit it; for example, the preset distance threshold can be 0.2 meters, 0.3 meters, etc.
[0081] Step S105: Based on the associated data, perform nonlinear optimization on the self-motion estimation data in the set and multi-frame observation data to determine the positioning result.
[0082] The localization result includes the current pose of the mobile platform, which includes both position and orientation. Optionally, this localization result is a global localization result.
[0083] Alternatively, the nonlinear optimization solution can be obtained using Gauss-Newton's method, gradient descent, or other nonlinear optimization methods.
[0084] This embodiment, by combining road observation data and self-motion estimation data from multiple frames of observation data, can accurately reconstruct a 3D semantic map. It uses GNSS data to determine the target range and acquires a second offline 3D semantic map within that range. The online 3D semantic map is then correlated with the second offline 3D semantic map, providing a basis for nonlinear optimization solutions. This method fully utilizes different types of information (road observation data, self-motion estimation data, and GNSS data) from multiple frames of observation data, achieving effective fusion of multi-source information. This effectively reduces errors that may arise from a single data source, improving the reliability and accuracy of positioning. Because this method relies not only on GNSS data but also incorporates other observation data, it can still achieve effective positioning even when GNSS signals are unstable or unavailable. Through nonlinear optimization solutions, it can be dynamically adjusted and optimized based on actual observation data, adapting to positioning needs under different environments and conditions, exhibiting high flexibility and adaptability. Compared to related technologies, this application requires relatively less storage space and computational load, and through multi-source data fusion and nonlinear optimization solutions, it can provide lightweight, high-precision, and robust positioning services in complex environments.
[0085] Figure 3 is a flowchart illustrating a positioning method according to another exemplary embodiment. As shown in Figure 3, the positioning method provided in this embodiment is a further refinement based on the positioning method provided in the previous embodiment of this application. The positioning method provided in this embodiment includes the following steps:
[0086] Step S201: Obtain the first offline 3D semantic map.
[0087] Step S202: Obtain the road observation data, self-motion estimation data, and GNSS data corresponding to each time point within the sliding window.
[0088] In this embodiment, the implementation of steps S201-S202 is the same as step S101 in the previous embodiment, and will not be repeated here.
[0089] Step S203: Determine the road observation data, self-motion estimation data and GNSS data corresponding to the same moment as a frame of observation data to obtain multiple frames of observation data.
[0090] Optionally, the road observation data, motion estimation data, and GNSS data acquired in step S202 each carry their own timestamps, thereby allowing road observation data, motion estimation data, and GNSS data corresponding to the same timestamp to be identified as a single frame of observation data. In some scenarios, due to transmission delays or processing delays, the timestamps of road observation data, motion estimation data, and GNSS data collected at the same time may be inconsistent. In such cases, time synchronization processing is required first to align the timestamps before identifying the road observation data, motion estimation data, and GNSS data corresponding to the same timestamp as a single frame of observation data. For example, the time synchronization processing is specifically implemented by using interpolation or filtering methods to process data points with incompletely consistent timestamps.
[0091] In this embodiment, steps S202-S203 are one way to acquire multiple frames of observation data. By treating the observation data at the same moment as a single frame of observation data, data synchronization is achieved.
[0092] This application does not limit the execution order of steps S201 and S202-S203. In some embodiments, step S201 may be executed first, followed by steps S202-S203; steps S202-S203 may be executed first, followed by step S201; or steps S201 and S202-S203 may be executed simultaneously.
[0093] Step S204: Obtain the second offline 3D semantic map within the target area from the first offline 3D semantic map.
[0094] In this embodiment, the implementation of step S204 is the same as that of step S102 in the previous embodiment, and will not be described again here.
[0095] In this embodiment, after step S204, step S103 is executed, which is to perform semantic map 3D reconstruction based on road observation data and self-motion estimation data in multi-frame observation data to generate an online 3D semantic map. In some embodiments, step S103 is specifically implemented as the following step S205.
[0096] Step S205: Input the road observation data and self-motion estimation data from the multi-frame observation data into the motion reconstruction structure algorithm, and use the motion reconstruction structure algorithm to perform semantic map 3D reconstruction to generate an online 3D semantic map.
[0097] Among them, the Structure from Motion (SfM) algorithm is a three-dimensional (3D) reconstruction algorithm. Optionally, step S205 is specifically implemented as follows: using road observation data from multi-frame observation data, feature matching is performed between adjacent frames to find the correspondence of the same features in different frames, thereby finding the corresponding feature pairs; the motion trajectory is estimated using self-motion estimation data and feature matching results; based on the feature matching results and the estimated motion trajectory, initial three-dimensional features are generated; the initial three-dimensional features are globally optimized according to the Bundle Adjustment algorithm to minimize reprojection error; semantic annotation is performed on the optimized three-dimensional features based on traffic element observation data and road element observation data from the road observation data; the three-dimensional features and semantic annotation results are integrated to generate an online three-dimensional semantic map.
[0098] Optionally, the road observation data includes traffic element observation data and road element observation data. Features can be traffic elements or road elements. Feature matching refers to matching the same element between adjacent frames. Specifically, if the traffic element observation data includes element identifiers for traffic elements, then traffic elements corresponding to the same element identifier between adjacent frames can be considered as feature pairs. Similarly, if the road element observation data includes element identifiers for road elements, then road elements corresponding to the same element identifier between adjacent frames can be considered as feature pairs. The feature matching result includes feature pairs corresponding to multiple frames of observation data.
[0099] The motion trajectory includes motion parameters at each moment within the sliding window, such as position, attitude, velocity, and acceleration. Optionally, the motion trajectory can be represented as time-series data: the position, attitude, velocity, and acceleration information at different time points are represented in time series form. Optionally, the operation of estimating the motion trajectory using self-motion estimation data and feature matching results can be implemented using filters (such as Kalman filters) or optimizers (as shown in the figure).
[0100] Taking a filter as an example, firstly, the filter's state vector is defined, which includes position, velocity, and attitude. During prediction, self-motion estimation data is used to predict the state at the next moment. This specifically involves integrating the acceleration to obtain predicted values for position and velocity, and using the rotation matrix update formula to predict attitude changes. During updating, feature matching results are used to update the state estimate. This specifically involves calculating the reprojection error of the features and using the reprojection error to update the state vector. For each frame of observation data and the corresponding self-motion estimation data, the prediction and update steps are repeated. In this way, the filter can gradually fuse self-motion estimation data and feature matching results to obtain more accurate estimation results. In three-dimensional space, the rotation matrix is a 3x3 orthogonal matrix used to describe the rotational state of the object. The update formula for the rotation matrix depends on the specific rotation representation method and application scenario; for example, it can be an update based on Rodrigues's Formula, a quaternion-based update, etc.
[0101] Optionally, generating initial 3D features based on feature matching results and estimated motion trajectories is specifically implemented as follows: based on feature matching results and estimated motion trajectories, the features in the 2D image are back-projected into 3D space to obtain initial 3D features. The features in the 2D image include traffic elements and road elements.
[0102] In practical applications, online 3D semantic maps are generated online and need to be updated periodically to ensure accuracy. For example, a sliding window technique can be used to update the online 3D semantic map in real time using the latest observation data and self-motion estimation data.
[0103] In this embodiment, multi-frame observation data provides rich perspective and detail information. Combined with self-motion estimation data, it can significantly improve the accuracy and completeness of 3D reconstruction. The SfM algorithm can effectively recover the 3D structure of the scene and generate a high-precision 3D model through multi-view geometric analysis. Through semantic map reconstruction, not only can the geometric information of the scene be obtained, but also semantic tags (such as roads, buildings, pedestrians, etc.) can be combined to provide a richer environmental understanding.
[0104] In this embodiment, after step S205, step S104 is executed, which involves associating online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs. In some embodiments, step S104 is specifically implemented as the following step S206.
[0105] Step S206: Using a graph matching algorithm, online features in the online 3D semantic map are associated with offline features in the second offline 3D semantic map to obtain a set of associated data pairs.
[0106] Graph matching is used to solve for the matching relationships between features and features in different data sets. Optionally, step S206 is specifically implemented as follows: extracting feature identifiers, categories, and location information of offline features from the second offline 3D semantic map to construct an offline feature set; extracting feature identifiers, categories, and location information of online features from the online 3D semantic map to construct an online feature set; using a graph matching algorithm to match the features of online features in the online feature set with the features of offline features in the offline feature set to obtain matching results; and generating a set of associated data pairs based on the matching results. Each associated data pair contains the feature identifier, category, and location information of one online feature and one offline feature. Optionally, the graph matching algorithm includes, but is not limited to: distance-based matching, feature descriptor-based matching (such as SIFT, SURF, etc.), and graph structure-based matching.
[0107] In this embodiment, the graph matching algorithm can utilize graph information for feature matching, which can more accurately identify and associate corresponding features in online and offline maps compared to using only geometric or attribute information. The graph matching algorithm can handle noise and incomplete data in the map. For example, if some online or offline features are missing or have errors, the graph matching algorithm can still find the optimal match through the overall graph structure, thereby improving the robustness of the system.
[0108] Step S207: Based on the associated data, perform nonlinear optimization on the self-motion estimation data in the set and multi-frame observation data to determine the positioning result.
[0109] In some embodiments, the prediction and positioning results are achieved by constructing an error function and solving the error function nonlinearly. Step S207 is specifically implemented as follows: Steps S2071-S2072:
[0110] Step S2071: Using the positioning results at each time point within the sliding window as the state variables to be estimated, construct an error function based on the state variables to be estimated, the associated data set, and the self-motion estimation data in the multi-frame observation data.
[0111] The objective of nonlinear optimization is to minimize the error function.
[0112] In some embodiments, the set of associated data pairs includes multiple associated data pairs, each including location information corresponding to associated online and offline elements. Accordingly, step S2071 includes: for each moment within the sliding window, using the difference between the self-motion estimation data at that moment and the previous moment as the self-motion estimation constraint, and determining the difference between the state variable to be estimated at that moment and the previous moment and the self-motion estimation constraint as the motion error term at that moment; using the location information corresponding to the associated online and offline elements in the associated data pairs as the observation constraint, for each moment within the sliding window, using the positioning result at that moment as the state variable to be estimated, and determining the difference between the predicted position of the target element predicted based on the state variable to be estimated and the location information as the observation error term at that moment; the target element is the same element referred to by the online and offline elements; and determining the sum of the motion error term and the observation error term at each moment within the sliding window as the error function.
[0113] Here, the sliding window refers to the sliding window corresponding to multiple frames of observation data. Specifically, the multiple frames of observation data include the observation data corresponding to each time point within the sliding window.
[0114] The state variables to be estimated are the pose set X = {x0,…,xi,…,xn} within the sliding window, where xi = (R,t)∈SE(3), xi is the representation of the i-th frame pose in three-dimensional space, t represents the position, R represents the pose, i∈[0,n], and n is the length of the sliding window. SE(3) is a special Euclidean group that includes rotation and translation. The observation z has the increment ΔT = (ΔR,Δt) of the self-motion estimation data and the observation zk of various semantic map elements mi, where ΔT represents the relative motion of the mobile platform from xi to xi+1, and zk represents the observation of the k-th element at pose xi. These observations are related to the perception of the mobile platform at a certain pose xi. Specifically, the relative motion can be calculated based on the self-motion estimation data to determine the correlation.
[0115] All the state variables to be estimated and their associated observations within the sliding window can be constructed into a factor graph structure, as shown in Figure 4. x_n represents the state variable to be estimated, the lines represent the associated observations (also called edges), represents the error function between the observations and the state variable to be estimated, the identifier constraint refers to the observation constraint, and the Ego_motion constraint refers to the self-motion estimation constraint. For an error function that follows a Gaussian distribution, the factor can be expressed as follows: (Formula 1)
[0116] Formula 1:
[0117] Where, f(x) k) represents the error function; exp() is the representation of the natural exponential function, used to calculate the negative exponent of e; |||| represents a certain norm, such as the Euclidean norm (i.e., the length of the vector), and ||||2 represents the square of the norm; err(xk,z) is an error vector representing a certain difference between the attitude xk and the observation z, ||err(xk,z)||2∑i represents the weighted sum of squares of the error vector, where ∑i represents the measurement uncertainty and represents the weight of the error vector, which is a fixed value set according to the error distribution of the observation.
[0118] Z = {z0,…,zi} represents all observations within the sliding window. Therefore, maximizing the posterior probability of the state variable X = {x0,…,xi} to be estimated can be expressed as follows: Formula 2:
[0119] Formula 2: X * =argmaxp(X|Z)
[0120] Formula 2 represents the search for the value of X that maximizes the conditional probability p(X|Z) given the observation Z, i.e., X. * p(X|Z) represents the probability of X occurring given Z; argmax() is a mathematical operator used to find the value of () that maximizes the function. According to Bayes' theorem, the joint probability distribution can be decomposed into a combination of prior information and conditional probability, as shown in Formula 3 below:
[0121] Formula 3:
[0122] Each term can be represented by an independent factor, as shown in Formula 4 below:
[0123] Formula 4:
[0124] The problem of finding the maximum value in Equation 1 can be transformed into a problem of finding the minimum value by taking the negative logarithm. Therefore, this maximum a posteriori estimation problem can be equivalently transformed into a nonlinear least squares problem, as shown in Equation 5 below:
[0125] Formula 5:
[0126] The state variable to be estimated is X = {x0,…,xi}, and there are multiple observations. Therefore, the error function can be expressed as Equation 6 below:
[0127] Formula 6:
[0128] Here, E(X) represents the error function, which combines two different types of error terms, and h(xi,xi+1) represents the difference between the two states. This represents a special subtraction for SE(3), where h(xk,mi,zk) represents the observation constraints of the semantic map, and ΔT ii+1 This represents the relative motion of the mobile platform from xi to xi+1.
[0129] By using the difference between the self-motion estimation data at each time step and the previous time step as the self-motion estimation constraint, and determining the difference between the state variable to be estimated at each time step and the self-motion estimation constraint as the motion error term, the self-motion estimation data can be effectively used to constrain the positioning results and reduce accumulated errors. By using the location information corresponding to the online and offline elements already associated in the associated data pair as the observation constraint, and determining the difference between the predicted location of the target element based on the state variable to be estimated and the location information as the observation error term, the location information of the associated data pair can be fully utilized. By comprehensively considering the motion error term and the observation error term, the error function can more comprehensively reflect the actual situation, reduce the impact of single data source error on the positioning results, and thus obtain more accurate positioning results.
[0130] Step S2072: Perform nonlinear optimization on the error function to obtain the positioning result at the current time.
[0131] Alternatively, the nonlinear optimization solution can be obtained using Gauss-Newton's method, gradient descent, or other nonlinear optimization methods.
[0132] By employing nonlinear optimization, the self-motion estimation data from the associated data set and multi-frame observation data can be fully utilized to construct and optimize the error function, significantly improving the accuracy of the positioning results. The sliding window method allows the system to dynamically update the positioning results while processing multi-frame observation data, adapting in real-time to environmental changes and sensor data updates, ensuring the timeliness and accuracy of the positioning results. By constructing the error function, the self-motion estimation data from the associated data set and multi-frame observation data can be comprehensively utilized, fully mining information from different data sources and improving the reliability of the positioning results. The nonlinear optimization process aims to minimize the error function; through iterative optimization, the optimal solution can be gradually approximated, resulting in more accurate positioning results. The sliding window method can limit the range of error accumulation by periodically updating the state variables within the window, reducing the impact of accumulated errors on the positioning results during long-term operation.
[0133] The following uses a mobile platform, specifically a vehicle, to illustrate the positioning method provided in this application. Figure 5 is a schematic diagram of a positioning process according to another exemplary embodiment. This process includes: acquiring images captured by an onboard camera; performing image recognition and feature extraction on the images to obtain traffic element observation data and road element observation data; acquiring vehicle Ego_motion data using an IMU and wheel speed sensors, and acquiring GNSS data and an offline semantic map respectively; synchronizing the traffic element observation data, road element observation data, vehicle Ego_motion data, and GNSS data to obtain multi-frame observation data; initialization: performing range indexing in the offline semantic map based on the GNSS data to obtain an offline semantic map within a certain range; performing real-time 3D reconstruction based on the traffic element observation data, road element observation data, and vehicle Ego_motion data to obtain an online semantic map; associating the online semantic map and the offline semantic map within a certain range with data elements; fusion optimization: transforming the error distribution of various associated data within the sliding window, the adjacent time Ego_motion, and the positioning state to be estimated (positioning state refers to the positioning result) into a maximum a posteriori estimation problem, solving it using nonlinear optimization to obtain the global positioning result of the vehicle.
[0134] The purpose of this application is to design a lightweight, robust, and high-precision global positioning method. Based on this objective, this application proposes a global positioning method based on traffic and road elements. This application fully utilizes information from various traffic and road elements, reducing reliance on single elements. For example, even in areas lacking lane lines, accurate positioning can still be achieved if other elements are available. This application has weak dependence on GNSS data; only a rough global positioning is needed initially, and subsequent positioning stages can be completely independent of GNSS data, thus being largely unaffected by GNSS signals. It also does not rely on RTK technology, offering a better cost advantage. Compared to point cloud map matching schemes, this application is more lightweight, less susceptible to dynamic occlusion, and more robust. The sliding window optimization method in this application allows each state to have multiple associations within the window and performs multiple optimizations, resulting in higher accuracy, greater robustness, and smoother operation.
[0135] Figure 6 is a schematic diagram of a positioning device according to an exemplary embodiment. As shown in Figure 6, in this embodiment, the positioning device 30 can be disposed in an electronic device, and the positioning device 30 includes:
[0136] The first acquisition module 301 is used to acquire a first offline three-dimensional semantic map and multi-frame observation data; the observation data includes road observation data, self-motion estimation data and global navigation satellite system (GNSS) data corresponding to the same time.
[0137] The second acquisition module 302 is used to acquire a second offline three-dimensional semantic map within the target area from the first offline three-dimensional semantic map; the target area is centered on the location indicated by the GNSS data;
[0138] The generation module 303 is used to perform semantic map 3D reconstruction based on road observation data and self-motion estimation data in multi-frame observation data, so as to generate an online 3D semantic map.
[0139] The association module 304 is used to associate online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs;
[0140] The determination module 305 is used to perform nonlinear optimization on the self-motion estimation data in the set and multi-frame observation data based on the associated data, so as to determine the positioning result.
[0141] In some embodiments, the first acquisition module 301, when acquiring multi-frame observation data, is used for:
[0142] Acquire road observation data, self-motion estimation data, and GNSS data corresponding to each time point within the sliding window;
[0143] The road observation data, self-motion estimation data, and GNSS data corresponding to the same moment are determined as one frame of observation data to obtain multiple frames of observation data.
[0144] In some embodiments, the determining module 305 is configured to:
[0145] Using the positioning results at each time point within the sliding window as the state variables to be estimated, an error function is constructed based on the state variables to be estimated, the associated data set, and the self-motion estimation data in the multi-frame observation data.
[0146] The error function is solved by nonlinear optimization to obtain the positioning result at the current time.
[0147] In some embodiments, the set of associated data pairs includes multiple associated data pairs, and each associated data pair includes location information corresponding to the associated online and offline elements, respectively.
[0148] The determination module 305, when constructing an error function based on the self-motion estimation data in the set of associated data and multi-frame observation data, using the positioning results at each time point within the sliding window as the state variables to be estimated, is used for:
[0149] For each moment within the sliding window, the difference between the self-motion estimation data at that moment and the previous moment is used as the self-motion estimation constraint, and the difference between the state variable to be estimated at that moment and the previous moment and the self-motion estimation constraint is determined as the motion error term at that moment.
[0150] Using the location information corresponding to the online and offline elements already associated in the associated data pair as observation constraints, for each time point within the sliding window, the positioning result at that time point is taken as the state variable to be estimated. The difference between the predicted position of the target element based on the state variable to be estimated and the location information is determined as the observation error term at that time point; the target element is the same element referred to by the online and offline elements.
[0151] The sum of the motion error term and the observation error term at each moment within the sliding window is determined as the error function.
[0152] In some embodiments, the generation module 303 is configured to:
[0153] Road observation data and self-motion estimation data from multiple frames of observation data are input into the motion reconstruction structure algorithm. The motion reconstruction structure algorithm is then used to perform semantic map 3D reconstruction to generate an online 3D semantic map.
[0154] In some embodiments, the association module 304 is configured to:
[0155] A graph matching algorithm is used to associate online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs.
[0156] The positioning device 30 provided in this embodiment can execute the technical solution of the corresponding method embodiment. Its implementation principle and technical effect are similar to those of the corresponding method embodiment, and will not be described in detail here.
[0157] This application also provides an electronic device. FIG7 is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. As shown in FIG7, the electronic device 40 includes: a processor 401 and a memory 402 communicatively connected to the processor 401.
[0158] The memory 402 stores computer-executable instructions; the processor 401 executes the computer-executable instructions stored in the memory 402 to implement the positioning method provided in this application.
[0159] In this embodiment, the memory 402 and the processor 401 are connected via a bus. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be categorized as an address bus, a data bus, a control bus, etc.
[0160] The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein. The various components are interconnected via different buses and can be mounted on a common motherboard or otherwise as required.
[0161] In an exemplary embodiment, a mobile platform is also provided, which includes electronic devices. Exemplarily, the mobile platform may be a vehicle, a robotic platform (such as a service robot, an exploratory robot, a scientific research robot, etc.), a drone, or other equipment, but is not limited thereto.
[0162] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores computer-executable instructions that, when executed by a processor, are used to implement the positioning method provided in this application.
[0163] In an exemplary embodiment, a computer program product is also provided, including a computer program, which, when executed by a processor, is used to implement the positioning method provided in this application.
[0164] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0165] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0166] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0167] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0168] When an integrated unit / module is implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor, or other electronic component. Unless otherwise specified, memory can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as USB flash drives, random-access memory (RAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), enhanced dynamic random-access memory (EDRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), resistive random access memory (RRAM), high-bandwidth memory (HBM), and hybrid memory cube (HMC). Cube, magnetic storage, flash storage, disk, optical disk, portable hard drive or magnetic disk, and other media that can store program code.
[0169] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause an electronic device to execute all or part of the steps of the methods of the various embodiments of this application.
[0170] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0171] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0172] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A positioning method, characterized in that, include: Acquire the first offline 3D semantic map and multi-frame observation data; The observation data includes road observation data, self-motion estimation data, and Global Navigation Satellite System (GNSS) data corresponding to the same time. Obtain a second offline 3D semantic map within the target area from the first offline 3D semantic map; the target area is centered on the location indicated by the GNSS data; Semantic map 3D reconstruction is performed based on road observation data and self-motion estimation data from multiple frames of observation data to generate an online 3D semantic map; The online elements in the online 3D semantic map are associated with the offline elements in the second offline 3D semantic map to obtain a set of associated data pairs; Based on the associated data, nonlinear optimization is performed on the self-motion estimation data in the set and multiple frames of the observation data to determine the positioning result.
2. The method according to claim 1, characterized in that, Acquire multi-frame observation data, including: Acquire road observation data, self-motion estimation data, and GNSS data corresponding to each time point within the sliding window; The road observation data, self-motion estimation data, and GNSS data corresponding to the same moment are determined as one frame of observation data to obtain multiple frames of observation data.
3. The method according to claim 1, characterized in that, The step of performing nonlinear optimization on the self-motion estimation data in the set and multiple frames of the observation data based on the associated data to determine the positioning result includes: Using the positioning results at each time point within the sliding window as the state variables to be estimated, an error function is constructed based on the state variables to be estimated, the associated data set, and the self-motion estimation data in the observation data of multiple frames. The error function is solved by nonlinear optimization to obtain the positioning result at the current time.
4. The method according to claim 3, characterized in that, The set of associated data pairs includes multiple associated data pairs, and each associated data pair includes the location information corresponding to the associated online and offline elements respectively; The step of using the positioning results at each time point within the sliding window as the state variable to be estimated, and constructing an error function based on the state variable to be estimated, the associated data set, and the self-motion estimation data in multiple frames of the observation data, includes: For each moment within the sliding window, the difference between the self-motion estimation data of the moment and the previous moment is used as the self-motion estimation constraint, and the difference between the state variable to be estimated at the moment and the previous moment and the self-motion estimation constraint is determined as the motion error term at the moment. Using the location information corresponding to the online and offline elements already associated in the associated data pair as observation constraints, for each time point within the sliding window, the positioning result at that time point is used as the state variable to be estimated. The difference between the predicted position of the target element based on the state variable to be estimated and the location information is determined as the observation error term at that time point; the target element is the same element referred to by the online element and the offline element. The sum of the motion error term and the observation error term at each moment within the sliding window is determined as the error function.
5. The method according to claim 1, characterized in that, The step of performing semantic map 3D reconstruction based on road observation data and self-motion estimation data from multiple frames of observation data to generate an online 3D semantic map includes: The road observation data and self-motion estimation data from multiple frames of the observation data are input into the motion recovery structure algorithm, and the motion recovery structure algorithm is used to perform semantic map 3D reconstruction to generate an online 3D semantic map.
6. The method according to claim 1, characterized in that, The step of associating online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs includes: A graph matching algorithm is used to associate online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs.
7. The method according to claim 1, characterized in that, The process of obtaining the first offline 3D semantic map includes: Use multiple sensor devices to collect three-dimensional data of the target environment; Data from different sensor devices is integrated to obtain 3D data; The three-dimensional data is preprocessed; features are extracted from the preprocessed three-dimensional data, and semantic segmentation is performed on the preprocessed three-dimensional data using a deep learning algorithm to obtain semantic segmentation results; Based on the three-dimensional data and the semantic segmentation results, a three-dimensional model of the target environment is constructed using a three-dimensional reconstruction algorithm; The semantic segmentation results are combined with the 3D model to generate a 3D semantic map with semantic annotations.
8. The method according to claim 2, characterized in that, The step of determining the road observation data, self-motion estimation data, and GNSS data corresponding to the same moment as a single frame of observation data includes: The road observation data, self-motion estimation data, and GNSS data are time-synchronized to align timestamps, and then the road observation data, self-motion estimation data, and GNSS data corresponding to the same timestamp are determined as a single frame of observation data.
9. The method according to claim 5, characterized in that, The step of inputting road observation data and self-motion estimation data from multiple frames of observation data into the motion reconstruction structure algorithm, and using the motion reconstruction structure algorithm to perform semantic map 3D reconstruction to generate an online 3D semantic map, includes: Using road observation data from multiple frames of observation data, feature matching is performed between adjacent frames to find the correspondence of the same features in different frames, thereby finding the corresponding feature pairs as the feature matching results; The motion trajectory is estimated using the self-motion estimation data and the feature matching results; Based on the feature matching results and the estimated motion trajectory, an initial three-dimensional feature is generated; Based on the bundling adjustment algorithm, the initial 3D features are globally optimized to minimize the reprojection error, resulting in optimized 3D features. Based on the traffic element observation data and road element observation data in the road observation data, the optimized three-dimensional features are semantically annotated, and the three-dimensional features and semantic annotation results are integrated to generate an online three-dimensional semantic map.
10. The method according to claim 6, characterized in that, The graph matching algorithm is used to associate online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs, including: Extract the feature identifiers, categories, and location information of offline features from the second offline 3D semantic map to construct an offline feature set; Extract the element identifiers, categories, and location information of online elements from the online 3D semantic map to construct an online element set; Using a graph matching algorithm, the features of online elements in the online feature set and the features of offline elements in the offline feature set are matched to obtain the matching result; Based on the matching results, the set of associated data pairs is generated.
11. The method according to claim 1, characterized in that, The step of obtaining a second offline 3D semantic map within the target area from the first offline 3D semantic map includes: Based on the GNSS data, a range index is performed in the first offline semantic map to obtain the second offline semantic map within the target range.
12. An electronic device, characterized in that, include: A processor and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the positioning method as described in any one of claims 1 to 11.
13. A mobile platform, characterized in that, Including the electronic device as described in claim 12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the positioning method as described in any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the positioning method as described in any one of claims 1 to 11.
16. A positioning device, characterized in that, include: The first acquisition module is used to acquire a first offline 3D semantic map and multi-frame observation data; the observation data includes road observation data, self-motion estimation data and Global Navigation Satellite System (GNSS) data corresponding to the same time. The second acquisition module is used to acquire a second offline 3D semantic map within the target area from the first offline 3D semantic map; the target area is centered on the location indicated by the GNSS data; The generation module is used to perform semantic map 3D reconstruction based on road observation data and self-motion estimation data from multi-frame observation data, so as to generate online 3D semantic maps; The association module is used to associate online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs; The determination module is used to perform nonlinear optimization on the self-motion estimation data in the set and multi-frame observation data based on the associated data, so as to determine the positioning result.
17. The apparatus according to claim 16, characterized in that, The first acquisition module, when acquiring multi-frame observation data, is used for: Acquire road observation data, self-motion estimation data, and GNSS data corresponding to each time point within the sliding window; The road observation data, self-motion estimation data, and GNSS data corresponding to the same moment are determined as one frame of observation data to obtain multiple frames of observation data.
18. The apparatus according to claim 17, characterized in that, The determining module 305 is used for: Using the positioning results at each time point within the sliding window as the state variables to be estimated, an error function is constructed based on the state variables to be estimated, the associated data set, and the self-motion estimation data in the multi-frame observation data. The error function is solved by nonlinear optimization to obtain the positioning result at the current time.
19. The apparatus according to claim 17, characterized in that, The set of associated data pairs includes multiple associated data pairs, and each associated data pair includes the location information corresponding to the associated online and offline elements respectively; The determining module, when constructing an error function based on the positioning results at each moment within the sliding window as the state variable to be estimated, the associated data set, and the self-motion estimation data in the multi-frame observation data, is used to: For each moment within the sliding window, the difference between the self-motion estimation data at that moment and the previous moment is used as the self-motion estimation constraint, and the difference between the state variable to be estimated at that moment and the previous moment and the self-motion estimation constraint is determined as the motion error term at that moment. Using the location information corresponding to the online and offline elements already associated in the associated data pair as observation constraints, for each time point within the sliding window, the positioning result at that time point is taken as the state variable to be estimated. The difference between the predicted position of the target element based on the state variable to be estimated and the location information is determined as the observation error term at that time point; the target element is the same element referred to by the online and offline elements. The sum of the motion error term and the observation error term at each moment within the sliding window is determined as the error function.
20. The apparatus according to claim 16, characterized in that, The generation module is used for: Road observation data and self-motion estimation data from multiple frames of observation data are input into the motion reconstruction structure algorithm. The motion reconstruction structure algorithm is then used to perform semantic map 3D reconstruction to generate an online 3D semantic map.
21. The apparatus according to claim 16, characterized in that, The associated module is used for: A graph matching algorithm is used to associate online features in the online 3D semantic map with offline features in the second offline 3D semantic map to obtain a set of associated data pairs.
Citation Information
Patent Citations
Positioning method and device, equipment and storage medium
CN111511017A
Positioning method and device, computer storage medium and processor
CN114440860A
Pose optimization method and device, electronic equipment and storage medium
CN114842080A
Multi-source matching positioning method and system, electronic equipment and storage medium
CN116026314A
Visual surveying and mapping navigation positioning method based on lightweight point cloud map
CN116839600A