Multi-sensor fusion mobile robot positioning method and system and electronic equipment
By using a multi-sensor fusion method, local map features are constructed using visual data and LiDAR data, and global optimization is performed. This solves the accuracy and stability problems of mobile robot positioning systems in complex environments, and achieves high-precision positioning results.
Patent Information
- Application Number
- CN202511087370.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
Existing mobile robot positioning systems suffer from insufficient positioning accuracy and poor stability in complex environments, especially issues such as wheel encoder slippage, IMU zero-bias drift, lidar positioning failure, and visual sensor performance degradation under varying lighting conditions.
By deeply fusing visual and LiDAR data, local map features are constructed, and global map optimization is performed using visual and LiDAR constraint factors. Sensor weights are dynamically adjusted to ensure positioning accuracy and robustness.
Maintaining high positioning accuracy and precision in complex environments enhances the stability and robustness of the positioning system, especially in feature degradation scenarios where adaptive weight adjustment ensures positioning accuracy.
Smart Images

Figure CN120991827A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mobile robot positioning, and in particular to a mobile robot positioning method and system based on multi-sensor fusion and an electronic device. BACKGROUND
[0002] In a mobile robot positioning system, traditional laser positioning relies on a wheel encoder and an IMU (Inertial Measurement Unit) for pose estimation. However, there are significant limitations. The wheel encoder is prone to slipping on bumpy or low-friction road surfaces, and accumulates large errors over a long period of operation. Although the IMU can provide high-frequency attitude information, low-cost IMUs are prone to zero drift problems during long-term operation, affecting positioning accuracy and making it difficult to meet real-time correction requirements.
[0003] Although the laser radar provides accurate distance measurement, it is prone to positioning failure in environments such as long corridors with few features. At the same time, laser radar is sensitive to dynamic objects and is prone to incorrect map updates in crowded environments. Visual sensors can provide rich environmental information, but traditional visual SLAM performance declines in the presence of light changes or motion blur. Existing multi-sensor fusion positioning systems, although they fuse data from visual sensors and laser radars, to some extent, avoid the problems of positioning drift and map overlay that can occur with single laser radar positioning, but the visual sensor is only used to assist in repositioning and does not fully utilize its pose constraint function in feature-rich scenarios, resulting in obvious deficiencies in system stability and positioning accuracy.
[0004] Therefore, there is an urgent need to develop a high-precision positioning system that can deeply fuse multi-sensor information and fully utilize the advantages of each sensor to meet positioning requirements in complex scenarios. SUMMARY
[0005] The present application provides a multi-sensor fusion mobile robot positioning method, system and electronic device. The reference pose of the robot in the environment is determined by the visual data and laser radar data, and then a local map is constructed based on the motion distance and angle, and the local map subgraph features are extracted. Then, the local area subgraph features are used to construct visual and laser radar limiting factors for global map optimization, effectively solving the cumulative errors that may occur in the local map, and ensuring that the robot can maintain high positioning accuracy and accuracy in complex environments.
[0006] According to an aspect of the present application, a multi-sensor fusion mobile robot positioning method is provided, comprising:
[0007] obtaining laser radar data and visual data of a mobile robot;
[0008] determine a reference pose of the mobile robot according to the visual data and / or the lidar data;
[0009] determine a local map according to a distance value and / or an angle value of motion of the mobile robot;
[0010] construct a visual constraint factor and a lidar constraint factor based on the local map subgraph features, perform global map optimization, and determine a mobile robot positioning result.
[0011] According to another aspect of the present application, a multi-sensor fusion mobile robot positioning system is provided, comprising:
[0012] a data acquisition module configured to acquire lidar data and visual data of a mobile robot;
[0013] a local map construction module configured to determine a local map according to a distance value and / or an angle value of motion of the mobile robot;
[0014] a feature extraction module configured to determine local map subgraph features according to the reference pose and the local map;
[0015] a global optimization module configured to construct a visual constraint factor and a lidar constraint factor based on the local map subgraph features, perform global map optimization, and determine a mobile robot positioning result.
[0016] According to another aspect of the present application, an electronic device is provided, comprising:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein
[0019] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the multi-sensor fusion mobile robot positioning method according to any one of the embodiments of the present application.
[0020] The technical solution of the embodiments of the present application acquires lidar data and visual data of a mobile robot, determines a reference pose of the mobile robot according to the visual data and / or the lidar data, determines a local map according to a distance value and / or an angle value of motion of the mobile robot, determines local map subgraph features according to the reference pose and the local map, constructs a visual constraint factor and a lidar constraint factor based on the local map subgraph features, performs global map optimization, and determines a mobile robot positioning result.
[0021] The technical solution uses visual data instead of the wheel odometer and the IMU in the prior art as the input of the pose estimator, deeply fuses the visual data and the laser radar data to realize positioning and mapping, ensures that the positioning system is not affected by the slippage of the robot movement, and the publication frequency of the visual sensor is higher, so that the robot pose can be corrected in time, the cumulative error that may occur in the local map is effectively solved, the positioning accuracy of the robot in the complex environment is ensured to be higher, and the positioning precision and the robustness are improved.
[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0024] Figure 1 is a flowchart of a multi-sensor fusion mobile robot positioning method according to the first embodiment of the present application;
[0025] Figure 2 is a flowchart of a subgraph construction method suitable for the embodiments of the present application;
[0026] Figure 3 is a flowchart of a multi-sensor fusion mobile robot positioning method according to the second embodiment of the present application;
[0027] Figure 4 is a schematic diagram of forming a common view relationship by observing a map point between key frames suitable for the embodiments of the present application;
[0028] Figure 5 is a schematic diagram of a laser radar limiting factor suitable for the embodiments of the present application;
[0029] Figure 6 is a block diagram of a positioning system suitable for multi-sensor fusion suitable for the embodiments of the present application;
[0030] Figure 7 is a schematic diagram of a multi-sensor fusion positioning system structure according to the third embodiment of the present application;
[0031] Figure 8 is a schematic diagram of an electronic device for implementing a multi-sensor fusion mobile robot positioning method according to the embodiments of the present application. DETAILED DESCRIPTION
[0032] In order for those skilled in the technical field to better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0033] It should be noted that the terms "first", "second", "target" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0034] Embodiment one
[0035] Figure 1 A flow chart of a multi-sensor fusion mobile robot positioning method is provided for the embodiments of the present application. The present embodiment can be applied to mobile robot positioning scenarios, and through deep fusion of visual sensor (the visual sensor used in the present technical scheme is RGB-D) and laser radar data, the positioning accuracy and stability problems in complex motion, feature degradation and other environments are solved. Especially suitable for in feature-rich scenes (such as multi-object, multi-texture environment), relying on visual information can fully play the role of pose constraint, and reduce the error of laser positioning alone. In the present embodiment, the working frequency of the RGB-D camera is set to 30Hz, and the scanning frequency of the laser radar is 10Hz. In order to ensure the time synchronization of the data, the system uses a hardware timestamp mechanism, and all sensor data are marked with accurate timestamps.
[0036] The technical scheme disclosed in the present embodiment can be executed by a mobile robot positioning device, which can be realized in the form of hardware and / or software. The mobile robot positioning device can be configured in any electronic device with network communication function, which can be a self-moving robot, an AGV trolley or the like target device. As shown in the figure, the multi-sensor fusion mobile robot positioning method comprises: Figure 1
[0037] S110 obtains laser radar data and visual data of the mobile robot and pre-processes
[0038] The mobile robot in the technical solution is provided with an RGB-D camera and a laser radar sensor, wherein the RGB-D camera acquires visual data containing color information and depth information at a preset frequency, and the laser radar acquires distance information of objects in the environment.
[0039] Specifically, the depth information collected by the visual sensor such as the RGB-D camera is often affected by factors such as sensor noise, environmental conditions, and unevenness of object surfaces, resulting in noise or errors in the depth data in the image. To improve the quality of the depth data, the embodiment uses bilateral filtering to denoise the depth image. For a pixel, bilateral filtering denoising considers both the spatial relationship with the neighborhood pixels and the similarity in color or grayscale of the neighborhood pixels.
[0040] Before bilateral filtering processing, the depth image data is first quality evaluated. For each depth pixel, the depth difference variance of the pixel and the neighborhood pixels is calculated. If the variance exceeds a preset threshold, the pixel is marked as a low-quality pixel, and its weight is reduced in subsequent processing. After the quality evaluation, the spatial distance between the current pixel and the neighborhood pixels is first calculated, that is, the spatial distance weight. The pixels with closer distance are given higher weight. Then, the pixel value difference between the current pixel and the neighborhood pixels is calculated, that is, the pixel value similarity weight. The neighborhood pixels with similar pixel values are given higher weight. Finally, the weighted average value is calculated, and the values of these weighted neighborhood pixels are used to replace the value of the current pixel. For a pixel I(x) in the image, the output I'(x) of the bilateral filtering is: (x)
[0041]
[0042] wherein N (x) is a set of neighborhood pixels of the pixel x; is a spatial Gaussian weight function, which measures the spatial distance between pixels; is a Gaussian weight function of pixel value difference, which measures the similarity of pixel values.
[0043] The distortion correction processing is to correct the image according to the intrinsic matrix of the RGB-D camera. The distortion correction can eliminate the image distortion caused by the optical characteristics of the lens (such as the radial distortion and the tangential distortion mentioned above), and ensure the accuracy of the correspondence between the pixel position and the real scene geometry, so as to eliminate the influence of the lens distortion on the subsequent feature extraction. The technical solution provided in the embodiment can use the calculation method of the existing technology to correct the radial distortion and the tangential distortion of the RGB-D camera, which is not limited in the embodiment.
[0044] S120 extracts visual features from the preprocessed visual data
[0045] In this embodiment, visual features are extracted by a deep learning model. The RGB visual data can be processed by a convolutional neural network to extract point features and line features. Specifically, a pre-trained convolutional neural network can be loaded. The RGB image is input into the model, and the output feature maps of the shallow layer (e.g., after the first convolutional block) and the slightly deeper layer (e.g., after the second convolutional block) are calculated and obtained for feature calculation. When extracting point features, the local response maximum point of each position in the feature map space is calculated as a candidate point feature position. The pixel coordinates of each point feature are recorded to form a point feature set P_curr.
[0046] The feature maps are fused in multiple channels (e.g., taking the average or maximum value) to construct an edge intensity map, and an edge detection algorithm (e.g., Sobel or Canny) is applied to extract edge pixel points, which are then connected to form initial line segments. The initial line segments are merged (line segments with an angle < 5° and an endpoint distance < 10 pixels are merged into long line segments) and filtered (short line segments with a length less than 30 pixels are removed), and each line feature is recorded to form a line feature set L_curr.
[0047] S130, visual odometry calculation based on the extracted visual features to obtain robot pose transformation information
[0048] Point features and line features of the current frame are extracted from the preprocessed visual data; based on an adaptive point-line tracking pose optimization model, the current frame is matched with the previous frame; and the robot pose transformation information is determined according to the following optimization model:
[0049]
[0050] where T is the pose transformation matrix to be solved, e pi (T) is the re-projection error model of the i-th feature point; e l1j (T), e l2j (T) is the re-projection error model of the two endpoints of the j-th feature line segment; w pi , w lj is an adaptive weight coefficient.
[0051] Point features and line features are extracted from the preprocessed current frame image. In specific implementation, the Harris corner detection method can be used to extract point features with good distinguishability, and the line segment detection algorithm can be used to extract structured edge information in the environment. When matching the point features and line features of the current frame and the previous frame, the weight contribution of the point features and line features in the pose optimization can be dynamically adjusted according to the environmental feature distribution.
[0052] A point-line combined visual odometry pose optimization model is established, and an optimization objective function is defined as:
[0053]
[0054] wherein T is a to-be-solved pose transformation matrix, e pi (T) is a re-projection error model of the i th feature point; e l1j (T) is a re-projection error model of the i th feature point; e l2j (T) is a re-projection error model of the j th feature line segment; w pi , w lj is an adaptive weight coefficient.
[0055] The calculation result of the above visual odometry provides relative motion information of the mobile robot, and the solution can obtain the pose change of the mobile robot from the last frame to the current frame.
[0056] S140 determines a robot reference pose according to visual data and / or lidar data
[0057] Based on the robot pose transformation information calculated by the visual odometry and the lidar data, a mobile robot reference pose is constructed. Specifically, the pose transformation obtained by the visual odometry is used as prior information, and the ranging information of the lidar is combined for pose correction, so that a more accurate reference pose of the mobile robot can be obtained.
[0058] S150 determines a local map based on a motion distance value and / or an angle value of the mobile robot
[0059] Based on the motion state information of the mobile robot, a key frame is determined by setting a distance and / or angle threshold, and a local map containing important visual feature information is constructed.
[0060] In combination with Figure 2 As shown in the figure, specifically, the system starts the initialization phase to take the first frame image as the initial key frame, performs feature point detection on the frame, saves the detected feature points as to-be-determined local map points, and uses them as the initial basis for subsequent frame matching and map construction. The cumulative motion distance value and the cumulative rotation angle of the mobile robot relative to the last key frame are monitored, and when the motion distance value exceeds the preset distance threshold or the angle value exceeds the preset angle threshold, the current frame is set as a new key frame. When the new key frame is determined, the feature information of the key frame is saved to the local map. For the existing local map points, if the current key frame observes the point, the observation count is increased. For the new feature points that are not matched to the existing map points in the current key frame, they are added to the local map, marked as new to-be-determined local map points, and the map size is expanded.
[0061] When the number of key frames reaches a preset threshold frame (such as 20 frames), the pending local map is subjected to a pruning operation. All map points in the current local map are traversed, and the local map points with an observation count less than a preset observation threshold are filtered out, and the high-quality map points are retained to determine the final local map, thereby ensuring the effectiveness and robustness of the local map and reducing the interference of redundant data on subsequent positioning and mapping.
[0062] The new input frame data is continuously processed until a subgraph construction termination condition is met (such as the robot completing the specified area traversal, the map coverage rate meeting the standard, etc.), and the subgraph construction is completed.
[0063] S160 determining local map subgraph features according to the reference pose and the local map
[0064] According to the reference pose and the local map, local map subgraph features are extracted, including the three-dimensional coordinates of feature points, the endpoint coordinates of feature lines, and the co-visibility relationship between key frames and other information.
[0065] S170 constructing visual and lidar limiting factors based on the local map subgraph features, performing global map optimization, and determining the mobile robot positioning result.
[0066] According to the local map subgraph features obtained in the previous steps, visual and lidar limiting factors are constructed, and global map optimization is performed through multi-sensor information fusion to finally determine the accurate positioning result of the mobile robot. Among them, the global optimization constraint model is constructed by fusing the lidar limiting factor and the visual limiting factor as follows:
[0067]
[0068] wherein e lidar is the lidar limiting factor, e vision is the visual limiting factor, and λ is the weight factor. The weight factor λ is dynamically adjusted, and the weight is adaptively allocated according to the reliability of different sensor data.
[0069] The adaptive weight adjustment mechanism proposed in the embodiment of the present application can dynamically adjust the weight allocation of vision and lidar according to the environmental characteristics and sensor data quality. The weight allocation of vision and lidar can be adjusted according to the feature richness index, such as the number of feature points in the current visual frame, the number of effective ranging points of the lidar, and the quality of data collected by each sensor. For example, in an environment with rich visual features, λ can take a value of 0.3-0.5; in an environment with sparse features, λ can take a value of 0.1-0.2; and in a degraded environment, λ can take a value of 0.4-0.6.
[0070] The technical scheme of the present application can detect feature degradation scenarios and automatically adjust the positioning method in the degradation scenario. To avoid positioning errors, when the mobile robot detects that it is running in a long corridor or similar scenario with high similarity, the value of the weight factor λ is automatically adjusted. The technical scheme disclosed in the present application can calculate the geometric feature distribution of the laser radar scan data. When the distance measurement rate of the front and rear is less than the threshold value, the system determines that it is a long corridor scenario. The system ensures the accuracy of positioning in similar feature degradation scenarios by increasing the weight of visual data.
[0071] The optimal estimation value of all pose nodes is obtained by the above global optimization constraint model, and the pose node corresponding to the current time is the final positioning result of the mobile robot. The positioning result fuses the multi-sensor information of vision and laser radar, has high precision and robustness, and can meet the navigation and positioning needs of the mobile robot in complex environments.
[0072] The technical scheme disclosed in the present embodiment fuses visual data and laser radar data in a unified optimization framework to determine the reference pose of the robot in the environment, and then constructs a local map based on the motion distance and angle and extracts local map subgraph features. Then, the local area subgraph features are used to construct a visual limiting factor and a laser radar limiting factor for global optimization, the sensor weight is dynamically adjusted according to the environmental features and data quality, and the maximum advantage of each sensor is ensured in different scenarios. Effectively solve the cumulative error that may occur in the local map, and ensure that the robot can still maintain high positioning accuracy and accuracy in complex environments. For feature degradation scenarios such as long corridors, the scheme disclosed in the present embodiment designs an automatic detection and processing mechanism, which significantly improves the robustness of the positioning system.
[0073] Embodiment two
[0074] Figure 3 A multi-sensor fusion mobile robot positioning method flowchart is provided for the second embodiment of the present application. The present embodiment is optimized based on the above-mentioned embodiment. The main optimization is to construct a visual limiting factor and a laser radar limiting factor based on local map subgraph features for global map optimization, which includes: constructing a visual limiting factor based on local map subgraph features, the visual limiting factor including visual feature point constraints, visual feature line constraints and co-view constraints; constructing a laser radar limiting factor based on local map subgraph features, the laser radar limiting factor including intra-subgraph constraints, inter-subgraph constraints, node constraints and loop closure constraints; and fusing the visual limiting factor and the laser radar limiting factor for global map optimization. As shown in Figure 3 The method for constructing a visual limiting factor based on local map subgraph features specifically includes the following steps:
[0075] S210 adopts a point-line distance matching method to construct a visual feature point constraint
[0076] Specifically, the current frame feature points are projected into the local map, the two map points closest to each feature point of the current frame are searched in the local map, and a space segment is constructed by the two map points; a point-line distance matching method is used to construct a visual feature point constraint.
[0077]
[0078] wherein p is a robust kernel function, p is a current frame feature point, and a and b are the two closest map points.
[0079] In the specific implementation, if only one closest map point satisfying the condition can be found, a distance constraint between the two points can be constructed:
[0080] e plicp = p (‖p-a‖)
[0081] The point-line distance matching method is used to construct a visual feature point constraint, which can ensure that the current frame feature point and the corresponding area set structure in the local map remain consistent, and better preserve the straight line and curved surface information in the geometric shape.
[0082] S220 constructing a visual feature line constraint based on visual data
[0083] Specifically, the line feature endpoints in the current visual data frame are l1 and l2, and the corresponding line feature breakpoints in the local map are m1 and m2. The spatial distance error between the two pairs of endpoints is determined respectively, and the consistency of the line segment direction is calculated.
[0084] First endpoint constraint: e l1 = ||l1-m1||
[0085] Second endpoint constraint: e l2 =‖l2-m2‖
[0086] Line segment direction constraint:
[0087]
[0088] By constructing the endpoint constraint and the line segment direction consistency constraint, the endpoint position and direction of the line feature between different observation frames can be ensured to remain consistent, which can enhance the adaptability of the positioning system to complex environments.
[0089] S230 establishing a co-view constraint between key frames
[0090] In the visual positioning and map construction technical solution, in the process of sub-map construction and map optimization using key frame determination, map point observation counting, etc., the co-view relationship between key frames is formed through map point observation, Figure 4This is a schematic diagram illustrating the co-view relationship formed between keyframes in this invention through map point observation:
[0091] If the same point on the map is observed simultaneously by keyframes Ki and Kj, then a co-view constraint is established between these two keyframes. Figure 4 As shown, for map point P, if keyframe KF successfully observes map point P, and keyframes KF1 and KF2 also observe P, then KF1, KF2, and KF form a first-level co-observation relationship. When KF3 and KF4 observe the set of map points observed by KF1 and KF2, KF3 and KF4 form second-level co-observation relationships with KF1 and KF2, respectively. By associating keyframe networks through co-observation relationships, the dynamic construction and optimization of co-observation visual map points are achieved. This scheme employs a co-observation view combined with an essential graph optimization strategy to improve positioning accuracy and... Figure 1 Coherence and co-visual constraints can improve the accuracy and robustness of pose estimation, especially in situations with few visual features or occlusion, providing technical support for visual localization in applications such as mobile robots.
[0092] S240 constructs lidar constraint factors based on local map subgraph features. The lidar constraint factors include intra-subgraph constraints, inter-subgraph constraints, inter-node constraints, and loop closure constraints.
[0093] Combination Figure 5 As shown, the lidar constraint factor includes two sub-factors. Figure 1 Kazuko Figure 2 and associated nodes.
[0094] Specifically, intra-subgraph constraints characterize the relative pose relationships between LiDAR scan frames within a single subgraph. Pose constraints between adjacent nodes within the same subgraph are generated using a point-line distance algorithm to ensure local consistency within the subgraph. Inter-subgraph constraints characterize the constraints constructed by scanning and matching the current frame with adjacent subgraphs when the robot moves to a new subgraph. Pose associations between different subgraphs are established through overlapping regions, such as sub-... Figure 1 Nodes and children Figure 2 Node matching is used to eliminate map stitching errors and accumulated errors at the front end when connecting different subgraphs. Inter-node constraints characterize the difference between the relative pose transformations of two nodes in the global coordinate system and the relative pose transformations in the local coordinate system, constraining the relative pose relationships between adjacent nodes. Geometric constraints between discontinuous nodes can be generated after RANSAC removes mismatches. Loop closure constraints are used to correct drift errors in long-term operation, establishing strong correlation constraints between historical nodes and current nodes, which can significantly reduce accumulated errors. In this technical solution, the hierarchical constraint network, especially in long corridor scenarios, maintains pose continuity through inter-subgraph constraints.
[0095] The technical solution disclosed in the embodiment fuses visual data and laser radar data in a unified optimization framework, dynamically adjusts the weights of each sensor according to environmental characteristics and data quality, and ensures that each sensor can exert the maximum advantage in different scenes. The embodiment constructs a multi-level optimization model including visual feature point constraints, feature line constraints, and co-view and laser radar constraints, thereby improving the positioning accuracy and global consistency.
[0096] Embodiment three
[0097] The multi-sensor fusion-based mobile robot positioning system architecture provided by the application is shown in Figure 6 The system includes a sensor input layer, a data processing layer, a local map construction layer, and a global map optimization layer. Specifically, the sensor input layer accesses a laser radar and an RGB-D camera, wherein the laser radar is used to collect environmental point cloud data, and the RGB-D camera collects RGB color image and depth image data, providing multi-sensing information for positioning; the data preprocessing layer reduces noise and regularizes the original laser radar data through voxel filtering, adaptive voxel filtering, and removal of motion distortion, thereby ensuring the accuracy of subsequent processing; the local map construction layer, on one hand, completes visual pose estimation by combining pose estimator through point feature and line feature extraction and matching, and on the other hand, achieves motion estimation by relying on motion sensors and scan matching, thereby cooperatively constructing a local subgraph and returning the result; finally, the global map optimization layer integrates multi-source data to complete global map optimization, outputs a global map and a global pose, and realizes accurate positioning. Each module interacts through data flow, forming a complete positioning system of perception-preprocessing-local map construction-global optimization.
[0098] Figure 7 The structure diagram of the multi-sensor fusion-based mobile robot positioning system provided by the application is shown in Figure 7 The system includes:
[0099] The data acquisition module 310 is configured to acquire laser radar data and visual data of the mobile robot.
[0100] The pose estimation module 320 is configured to determine a reference pose of the mobile robot according to the visual data and / or the laser radar data.
[0101] The local map construction module 330 is configured to determine a local map based on the motion distance value and / or the angle value of the mobile robot.
[0102] The feature extraction module 340 is configured to determine local map subgraph features according to the reference pose and the local map.
[0103] The global optimization module 350 is configured to construct a visual constraint factor and a lidar constraint factor based on the local map subgraph features, perform global map optimization, and determine a mobile robot positioning result.
[0104] The pose estimation module further includes:
[0105] The preprocessing unit is configured to pre-process the visual data, including bilateral filtering denoising and distortion removal processing.
[0106] The visual feature extraction unit is configured to extract visual features from the pre-processed visual data.
[0107] The visual odometry calculation unit is configured to perform visual odometry calculation based on the visual features to obtain robot pose transformation information.
[0108] The pose construction unit is configured to construct a mobile robot reference pose based on the robot pose transformation information and the lidar data.
[0109] The visual odometry calculation module includes:
[0110] The feature extraction subunit is configured to extract point features and line features of a current frame from the pre-processed visual data.
[0111] The feature matching subunit is configured to perform feature matching between the current frame and a previous frame based on an adaptive point-line tracking pose optimization model.
[0112] The pose optimization subunit is configured to determine robot pose transformation information according to the following optimization model:
[0113]
[0114] wherein T is a pose transformation matrix to be solved, e pi (T) is a re-projection error model of the i-th feature point; e l1j (T), e l2j (T) is a re-projection error model of the two end points of the j-th feature line segment; w pi , w lj is an adaptive weight coefficient.
[0115] The local map construction module further includes:
[0116] The initialization unit is configured to set a first frame image as a key frame and save detected feature points as to-be-determined local map points.
[0117] The key frame determination unit is configured to set a current frame as a key frame when a motion distance value and / or an angle value exceeds a set threshold.
[0118] The feature saving unit is configured to save feature information of the key frame into a local map.
[0119] The map pruning unit is configured to prune the local map to be determined when the number of key frames reaches a preset threshold.
[0120] The feature saving module is configured to:
[0121] increment an observation count for an existing local map point;
[0122] add the new feature point to the local map;
[0123] The map pruning unit is specifically configured to:
[0124] discard a local map point with an observation count less than a preset observation threshold.
[0125] The global optimization module includes:
[0126] The visual constraint factor construction unit is configured to construct a visual constraint factor based on the local map subgraph features, the visual constraint factor including a visual feature point constraint, a visual feature line constraint, and a co-view constraint;
[0127] The laser radar constraint factor construction unit is configured to construct a laser radar constraint factor based on the local map subgraph features, the laser radar constraint factor including an intra-subgraph constraint, an inter-subgraph constraint, an inter-node constraint, and a loop closure constraint;
[0128] The fusion optimization module is configured to fuse the visual constraint factor and the laser radar constraint factor for global map optimization.
[0129] The fusion optimization module is further configured to:
[0130] fuse the laser radar constraint factor and the visual constraint factor to construct a global optimization constraint, with a constraint expression being:
[0131]
[0132] wherein e lidar is the laser radar constraint factor, e vision is the visual constraint factor, and λ is a weight factor.
[0133] The visual feature point constraint construction module in the visual constraint factor construction module is configured to:
[0134] project a feature point of a current frame into a local map;
[0135] determine two map points closest to each feature point of the current frame to construct a spatial line segment;
[0136] The point-line distance matching method is used to construct the visual feature point constraint by the following formula:
[0137]
[0138] wherein p is a robust kernel function, p is a current frame feature point, and a and b are the nearest two map points.
[0139] Embodiment Four
[0140] Figure 8 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the applications described and / or claimed in this document.
[0141] As shown in Figure 8 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected to the at least one processor 11 in communication, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0142] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0143] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the mobile robot localization method based on multi-sensor fusion.
[0144] In some embodiments, the mobile robot localization method based on multi-sensor fusion can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the mobile robot localization method based on multi-sensor fusion described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured as the mobile robot localization method based on multi-sensor fusion by any other suitable means, such as by means of firmware.
[0145] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0146] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0147] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0148] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0149] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), blockchain network, and the Internet.
[0150] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0151] It should be understood that the various forms of flow shown above can be reordered, added to, or have steps deleted. For example, the steps described in the present application can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which are not limited herein.
[0152] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A mobile robot localization method based on multi-sensor fusion, characterized in that, include: Acquire LiDAR and visual data from the mobile robot; The reference pose of the mobile robot is determined based on the visual data and / or lidar data. A local map is determined based on the motion distance and / or angle values of the mobile robot; sub-map features of the local map are determined based on the reference pose and the local map. Based on the features of the local map sub-map, visual constraint factors and LiDAR constraint factors are constructed to optimize the global map and determine the positioning result of the mobile robot.
2. The method according to claim 1, characterized in that, The visual data is data acquired by an RGB-D camera. Determining the mobile robot's reference pose based on the visual data and / or LiDAR data includes: The visual data is preprocessed, including bilateral filtering for noise reduction and distortion correction. Visual features are extracted from the preprocessed visual data; Visual odometry is performed based on the aforementioned visual features to obtain robot pose transformation information. Based on the robot pose transformation information and the lidar data, a reference pose for the mobile robot is constructed.
3. The method according to claim 2, characterized in that, The step of extracting visual features from the preprocessed visual data and performing visual odometry calculations based on the visual features to obtain robot pose transformation information includes: Extract point and line features of the current frame from the preprocessed visual data; Based on the adaptive point and line tracking pose optimization model, feature matching is performed between the current frame and the previous frame; The robot pose transformation information is determined based on the following optimization model: Where T is the pose transformation matrix to be solved, e pi (T) represents the reprojection error model for the i-th feature point; e l1j (T), e l2j (T) represents the reprojection error model of the two endpoints of the j-th feature line segment; w pi w lj These are adaptive weighting coefficients.
4. The method according to claim 1, characterized in that, A local map is determined based on the motion distance and / or angle values of the mobile robot, and sub-map features of the local map are determined based on the reference pose and the local map, including: The first frame image is used as a keyframe, and the detected feature points are saved as undetermined local map points. If the motion distance and / or angle value exceeds the set threshold, the current frame will be set as a keyframe; Save the feature information of the keyframes to the local map; When the number of keyframes reaches a preset threshold, a cropping operation is performed on the local map to be determined to define the local map.
5. The method according to claim 4, characterized in that, The feature information of the keyframes is saved to the local map, including: For existing local map points, increase the observation count; For newly added feature points, add them to the local map; The process of performing a cropping operation on the local map to be determined includes: Local map points with observation counts lower than a preset observation threshold will be removed.
6. The method according to claim 1, characterized in that, The process of constructing visual constraint factors and LiDAR constraint factors based on the features of the local map sub-map, and then optimizing the global map, includes: Visual constraint factors are constructed based on the features of the local map subgraph, and the visual constraint factors include visual feature point constraints, visual feature line constraints, and co-view constraints. Based on the features of the local map subgraph, a lidar constraint factor is constructed, which includes intra-subgraph constraints, inter-subgraph constraints, inter-node constraints, and loop closure constraints. The visual limiting factor and the lidar limiting factor are fused together for global map optimization.
7. The method according to claim 6, characterized in that, The process of fusing the visual limiting factor and the LiDAR limiting factor for global map optimization includes: By fusing lidar and visual constraint factors, a global optimization constraint is constructed, the expression of which is: Among them, e lidar For lidar, e vision λ is the visual constraint factor, and λ is the weighting factor.
8. The method according to claim 6, characterized in that, The visual feature point constraints in the visual constraint factor include: Project the feature points of the current frame onto the local map; Determine the two map points closest to each feature point in the current frame, and construct spatial line segments; The visual feature point constraints are constructed using the point-to-line distance matching method with the following formula: Where ρ is the robust kernel function, p is the feature point of the current frame, and a and b are the two nearest map points.
9. A mobile robot positioning system based on multi-sensor fusion, characterized in that, include: The data acquisition module is used to acquire LiDAR data and visual data of the mobile robot; The pose estimation module is used to determine the reference pose of the mobile robot based on the visual data and / or lidar data. A local map construction module is used to determine a local map based on the movement distance and / or angle values of the mobile robot. The feature extraction module is used to determine the features of the local map sub-map based on the reference pose and the local map; The global optimization module is used to construct visual constraint factors and lidar constraint factors based on the features of the local map sub-map, perform global map optimization, and determine the positioning result of the mobile robot.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multi-sensor fusion mobile robot localization method according to any one of claims 1-8.