SLAM (Simultaneous Localization and Mapping) positioning system integrating monocular vision and novel wheel type odometer
By using three omnidirectional wheel sensors instead of IMU in the vision-IMU-wheel speed gauge positioning system and designing a tightly coupled vision-wheel odometer fusion SLAM positioning system, the existing system's problems in state dimension and initialization difficulty are solved, and positioning accuracy and robustness are improved, especially when uneven surfaces or wheels are slipping.
Patent Information
- Application Number
- CN202510322912.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-06
AI Technical Summary
The existing vision-IMU-wheel speedometer positioning systems are challenging in state dimensions and initialization difficulty, and low-cost IMUs may lead to greater positioning errors and poor positioning accuracy and robustness when uneven surfaces or wheels are slipped.
Three omnidirectional wheel sensors are used instead of IMU to design a tightly coupled monocular vision and wheel odometer fusion SLAM positioning system. Through the tightly coupled nonlinear sliding window optimization of vision-wheel odometer, combined with local geometric features and dynamic plane constraint adaptive method of meta-learning, the positioning accuracy and robustness are improved.
In small-scale planar motion scenarios, positioning accuracy and system operation efficiency are improved, system burden is reduced, state dimension and initialization difficulty problems caused by IMU fusion are avoided, and positioning accuracy is maintained when uneven surfaces or wheels are slipping.
Smart Images

Figure CN120101773A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of positioning of a planar robot, and specifically designs a wheel odometer and monocular vision fusion SLAM positioning system added with three driven omnidirectional wheel sensors. Background Art
[0002] The robot's automatic navigation first needs to perceive the environment and determine its own position, which is SLAM (Simultaneous Localization and Mapping) technology. After years of development and accumulation, many excellent solutions have emerged in the SLAM field, especially visual SLAM technology, which has formed a basic theoretical framework.
[0003] Among them, V-SLAM technology based on visual sensors has been widely studied due to its low-cost sensor configuration and rich measurement information. From the form of feature point observation error, V-SLAM can be divided into two schemes based on indirect method and direct method. The indirect method has more stable feature point matching by extracting feature points and minimizing the geometric error of feature point reprojection, but the process of extracting and matching feature points is time-consuming. The direct method does not extract feature points, but estimates the motion of the camera by minimizing the photometric error of pixels. It has higher operating efficiency and retains more information in the environment, which can be used for three-dimensional reconstruction of the environment. V-SLAM schemes usually rely only on cameras as measurement information sources. However, in actual operation, when the scene environment is single or the camera moves quickly, it is easy to cause V-SLAM to fail in tracking and positioning, and its robustness is also poor. Therefore, the current research hotspot is the fusion SLAM scheme based on vision and other sensors (such as GPS, IMU, lidar, wheel speed meter, etc.). Among them, VINS-Mono has become one of the best positioning algorithms in V-SLAM with its fast and accurate initialization process, visual-inertial tightly coupled nonlinear sliding window optimization estimator, and four-degree-of-freedom pose graph optimization, and has been successfully applied to accurate pose trajectory estimation on unmanned aerial vehicles.
[0004] For wheeled robots, wheel odometers can effectively replace IMUs and provide high-frequency body state measurements. Under certain conditions, wheel odometers will not introduce degradation problems. For ground mobile robots with wheel speed sensors, the camera, inertial sensor and wheel speed sensor can be fused to solve the robustness problem in the above difficult scenarios. Based on EKF, the wheel speed inertial odometer is fused with the monocular visual odometer. Assuming that the robot runs on an ideal plane, 3-DOF pose estimation is performed. The wheel speed inertial odometer uses the fusion of wheel speed measurement and angular velocity measurement for dead reckoning. The wheel speed inertial odometer is used for EKF state prediction, and the visual odometer method is used for EKF measurement update. In the above loosely coupled method based on EKF, when the visual odometer cannot accurately calculate the pose due to insufficient visual features, the output of the visual odometer in the filter will have a lower weight, resulting in the effective visual observations being discarded together, resulting in a decrease in accuracy. In order to fully utilize the constraints of sensor measurement on pose estimation and improve the accuracy of pose estimation, a visual wheel speed SLAM system is constructed based on tightly coupled nonlinear optimization. The wheel speed sensor and the visual odometer are fused in a tightly coupled manner to solve the scale uncertainty of monocular vision, and the least squares problem corresponding to state estimation is solved using an optimized method. In the optimization process, the robot's pose is constrained to an ideal plane with 3 degrees of freedom, which improves the accuracy of pose estimation. This algorithm does not take into account the unreliability of wheel speed measurement. When the robot moves on an uneven surface or the wheel slips, incorrect wheel speed measurement will seriously affect the scale accuracy and even cause positioning failure.
[0005] In order to balance the system operation efficiency and positioning accuracy, the study found that in indoor and outdoor small-scale planar motion scenarios, the vision-driven wheel speedometer positioning system can reduce the additional system burden compared to the vision-IMU-wheel speedometer positioning system. In particular, the fusion of IMU will significantly increase the state dimension and initialization difficulty of the system, and does not significantly improve the positioning accuracy. In fact, for low-cost IMUs, it may cause greater positioning errors. Therefore, the present invention proposes a tightly coupled vision and wheel speedometer positioning solution to improve the positioning accuracy of the system while improving the system's operating efficiency. Summary of the invention
[0006] In view of the existing problems, the present invention provides a monocular vision-wheel odometer positioning system using a driven wheel sensor instead of an IMU, and its originality is:
[0007] A typical wheel odometer adopts a differential drive model, and the position and posture of the vehicle body are calculated by the linear velocity of the driving wheel. A common driven wheel odometer is a wheel odometer sensor with two omnidirectional wheels distributed at 90 degrees, which can only calculate the speed of the vehicle body posture and needs to work with the heading angle measured by the gyroscope. The present invention designs a sensor, as shown in the attached figure.Figure 3 As shown, a driven wheel odometer sensor with three omnidirectional wheels spaced 120° apart is used. This three-wheel module can calculate the speed and angular velocity of the vehicle body posture, thereby replacing the work of the gyroscope, pre-integrating the posture data between two key frames, and applying it to the positioning system that integrates monocular vision and wheel odometer.
[0008] The technical solution adopted by the present invention is as follows:
[0009] A positioning system integrating monocular vision and a novel wheel odometer sensor comprises the following steps:
[0010] S1. First, the camera and wheel odometer are calibrated internally, and the wheel odometer sensor wheel spacing and installation angle data are verified. Then, the wheel odometer sensor is used to calculate the body posture, and the wheel speed data is converted into the relative motion between two key frames by pre-integration.
[0011] S2. Synchronize the timestamps between sensors. Since the exposure time of the camera is uncertain, the synchronization effect of the timestamp cannot be guaranteed after the timestamp hardware synchronization. Therefore, the timestamp software synchronization is added. After the timestamp synchronization operation is completed, the wheel odometer and the monocular camera are jointly calibrated and initialized.
[0012] S3. Perform nonlinear sliding window optimization of tight coupling between vision and wheel odometer, combine observation data and motion information through maximum a posteriori estimation to form a comprehensive state estimation model. Solving the model can obtain the precise positioning of the robot in the actual environment. A dynamic plane constraint adaptive method based on local geometric features and meta-learning is added to constrain the model. For tight coupling, a hybrid strategy based on motion component decomposition and dynamic adjustment of residual entropy is designed. This method can avoid translation and rotation error coupling, allocate weights in a targeted manner, and perceive sensor degradation through real-time residual distribution;
[0013] S4, perform backend optimization and detect whether there is a loop to optimize the trajectory. The backend optimization module integrates the constraints provided by vision and wheel odometers and uses nonlinear optimization methods to accurately optimize the robot trajectory. Loop detection provides effective global pose constraints, thereby effectively reducing the cumulative error. The entire process optimizes the robot's positioning accuracy and robustness in large-scale environments by reducing errors.
[0014] Compared with the prior art, the present invention has the following beneficial effects:
[0015] (1) In the vision-IMU-wheel speedometer positioning system, the fusion of IMU will significantly increase the state dimension and initialization difficulty of the system, and will not significantly improve the positioning accuracy. In fact, for low-cost IMUs, it may bring greater positioning errors. Using the present invention to replace IMU can effectively improve the positioning accuracy of the system and improve the operating efficiency of the system at a certain cost.
[0016] (2) When a collision, abduction, etc. occurs, causing the wheels to slip, the driving wheels are the ones that slip, which will cause the calculated vehicle body posture to be wrong. However, the driven wheels used in the present invention are not prone to slip. When a collision, abduction, etc. occurs, the real-time posture of the vehicle body can be calculated more accurately without being affected by the slip of the driving wheels, and then the task can be continued.
[0017] (3) Pre-integration of wheel odometer posture data can combine posture changes over multiple time periods into increments, thus avoiding error accumulation caused by direct integration. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solution of the present invention, the following drawings are provided:
[0019] Figure 1 This is the structural diagram of the SLAM positioning system that integrates monocular vision and wheel odometer sensors;
[0020] Figure 2 This is a schematic diagram of the three-wheel sensor module model;
[0021] Figure 3 A three-dimensional schematic diagram of a single wheel of a three-wheel sensor module and its components;
[0022] Figure 4 This is a diagram of the timestamp hardware synchronization system;
[0023] Figure 5 It is the flowchart of timestamp hardware synchronization;
[0024] Figure 6 Schematic diagram of the timestamp software synchronization interpolation method. DETAILED DESCRIPTION
[0025] The specific implementation steps of the present invention are further described in detail below with reference to the accompanying drawings:
[0026] Step S1: Calibrate the intrinsic parameters of the camera and wheel odometer respectively. The camera's intrinsic parameter calibration uses the traditional camera calibration method, using the Kalibr calibration tool. For the intrinsic parameter calibration of the wheel odometer, the wheel spacing and installation angle of the wheel odometer sensor are calculated and verified by controlling the robot's linear motion and rotational motion.
[0027] The conversion relationship between wheel odometer data and robot posture is as follows:
[0028] As attached Figure 2 , define the three-wheel positioning module pose vector in the world coordinate system , then the composite motion speed of the module is ,Will Decompose into world coordinate system axes and Direction, get the velocity component , , using a three-dimensional vector Represents the movement of the module in the world coordinate system, where For wheel odometer around the center angular velocity.
[0029] Will In the coordinate system and Velocity component in direction , Decompose the kinematic equations to obtain is the distance between the three omnidirectional wheels and the center of the module, , , They are the linear speeds of the three omnidirectional wheels:
[0030]
[0031] There is a conversion relationship between the positioning module coordinate system and the world coordinate system :
[0032]
[0033] The kinematic model of the three-wheel omnidirectional wheel alignment system is:
[0034]
[0035] Substituting into the above formula we get:
[0036]
[0037] Then we can get the inverse kinematics model of the module, and solve the velocity and angular velocity of the robot:
[0038]
[0039] The sampling time is set to In the control module, it can be assumed that during the sampling time, the driven wheels A, B, and C are all in uniform motion state, so the posture calculation formula can be obtained:
[0040]
[0041] Mode middle, , , , Respectively represent and Within a sampling period Position coordinate parameters in the coordinate system, is the sampling period. , indicating the initial position, that is and The two coordinate systems coincide.
[0042] The above is the conversion relationship between the wheel odometer data and the robot posture. After obtaining the observation data, the result of pre-integration of the wheel odometer is as follows:
[0043]
[0044]
[0045] The formula represents the robot rotation transformation within a given time period, that is, between two key frames, where It is The angular velocity observation value of the frame can be obtained by the formula; the formula represents the robot position change between two key frames within a given time period, where Indicates Frame to The robot rotation transformation between frames, i.e. the content of the formula, Indicates The linear velocity observation value of the frame can be obtained by the formula.
[0046] Step S2: For the timestamp hardware synchronization solution, as shown in the attached Figure 4 , 5As shown. The frequency of the wheel odometer is set to 100Hz, and the camera frequency is 25Hz. Therefore, it is required to trigger the camera approximately once every 4 frames of wheel odometer data. The stm32 series of single-chip microcomputers are used for sub-control and sampling of the wheel odometer data. The specific method is to add a flag bit on the basis of the odometer data to determine whether to trigger the camera when a frame of data is received, which is the timestamp hardware synchronization. Considering that the transmission time from sensor data acquisition to reception by the host computer is affected by the amount of data, USB bandwidth and system scheduling, the timestamp of the system receiving the data is later than the real timestamp. This period of time is also called timestamp offset. The timestamp offsets of different sensors are different. Although the hardware layer has realized the timestamp synchronization of sensors, the image timestamp will change with the exposure time. Therefore, timestamp software synchronization is still required for subsequent external parameter calibration. For the timestamp software synchronization solution, that is, to achieve time consistency between sensors by software means, the present invention adopts a time synchronization interpolation method, as shown in the attached Figure 6 As shown, the camera’s sampling timeline is selected as a reference because its sampling frequency is lower and more stable. The wheel odometer data is interpolated to align its timestamp with the camera’s timestamp. , calculate the timestamp after , in the wheel odometer data, find the adjacent time and ,satisfy , use Lie group interpolation (Slerp for SO(3), linear interpolation translation) to calculate the interpolated pose:
[0047]
[0048] in:
[0049] After the timestamp synchronization is successful, the two sensors are jointly calibrated using the CamOdomCalibTool calibration tool. The coordinate transformation relationship between the camera and the wheel odometer is optimized using wheel speed measurement, external sensor data, and the PnP algorithm.
[0050] Then, the displacement and rotation information provided by the wheel odometer is used to assist initialization, ensuring that the visual system can obtain relatively accurate position and rotation estimates at the initial moment. This method first estimates the initial pose of the robot through the wheel odometer and the visual system, and determines the initial position of the robot in the global coordinate system based on the relative displacement and rotation information provided by the wheel odometer. The displacement estimated by the wheel odometer is matched with the relative position of the feature points provided by the visual system to achieve preliminary visual-wheel odometer joint initialization. Based on these preliminary estimates, we can obtain a relatively reliable initial pose.
[0051] Step S3: Combine the observation data and motion information to form a comprehensive state estimation model.
[0052]
[0053] The above formula is the estimated value obtained by maximizing the objective function. Substituting the observations of the camera and wheel odometer into it, we can get:
[0054]
[0055] in Represents the visual observation model, describing the Status at a moment Corresponding feature points probability. Represents the wheel odometer observation model, describing the motion information between two consecutive positions. By solving these models, the robot can be accurately positioned in the actual environment.
[0056] At the same time, once the system is initialized, a nonlinear sliding window optimization of the tightly coupled vision-wheel odometry is performed to perform a BA optimization on the poses of the image frames in the window and the positions of the landmarks they observe. If the number of frames in the window reaches the maximum limit, one frame needs to be discarded. Referring to the marginalization strategy in VIN-Mono, the system performs operations by determining whether the next-newest frame is a keyframe. If the next-newest frame is a keyframe, the system marginalizes the oldest keyframe in the window, and when marginalizing, the landmarks observed by the oldest keyframe are also processed to ensure the sparsity of subsequent BA optimization. On the contrary, if the next-newest frame is not a keyframe, it means that it is very close to the nearest keyframe. The system will directly discard the next-newest frame and its feature point observation information, but will retain the wheel speedometer measurement information of the next-newest frame. The specific strategy is to integrate the wheel odometry measurement between the current frame and the next-newest frame into the wheel speedometer pre-integration information of the next-newest frame.
[0057] The core framework of the dynamic plane constraint adaptation method based on local geometric features and meta-learning is as follows:
[0058] Drawing on the idea of adaptive smoothing terms in stereo matching, the constraint strength is dynamically adjusted by analyzing geometric features such as local curvature and gradient direction of the plane. The implementation method is as follows: design a local feature classifier based on the Gaussian mixture model, calculate the regional geometric complexity in real time, and map it to a constraint weight function to ensure that the constraint strength matches the local structure. For meta-learning-guided online parameter optimization, an environmental adaptive learning mechanism is introduced, and a lightweight neural network is pre-trained using meta-learning to quickly predict the optimal constraint parameters in new scenarios. The advantage of this method is that it solves the problem of traditional methods relying on manual parameter adjustment and improves the generalization ability for unknown plane scenes. Then, combined with the adaptive constraint strategy in nonlinear system control, a dual-threshold relaxation mechanism is designed:
[0059] (1) Main constraint threshold: ensuring safety margin;
[0060] (2) Sub-constraint threshold: Dynamically adjusted according to real-time error feedback, allowing local breakthroughs of constraints within a safe range to optimize global performance.
[0061] Based on the hybrid strategy of motion component decomposition and dynamic adjustment of residual entropy, this method can avoid the coupling of translation and rotation errors, allocate weights in a targeted manner, and perceive sensor degradation through real-time residual distribution, effectively solving the complementarity problem between wheel odometer and vision in structured or unstructured scenes. The specific scheme is as follows:
[0062] Experimentally decompose the motion state into translational components and the rotational component , respectively assign basic weights:
[0063] Translation weight: Wheel odometry is more robust to the translation component and is given a higher base weight:
[0064]
[0065] Rotation weight: Visual feature points are more sensitive to rotation, giving higher visual weight:
[0066]
[0067] Then adjust the dynamic residual entropy and the visual residual The chi-square norm of the feature point reprojection error is expressed as: , is the feature point in the world coordinate system, is the feature point in the image coordinate system; the wheel odometry residual is calculated by the difference between the motion prediction and the wheel odometry measurement: , statistically analyze the residual distribution within the sliding window and calculate its information entropy to quantify the uncertainty:
[0068]
[0069] in is the normalized residual histogram probability. The larger the entropy value, the lower the current confidence of the sensor.
[0070] Then calculate the entropy adjustment factor for the translation and rotation components separately:
[0071]
[0072] in To adjust the parameters, ensure that the weight decays exponentially when the entropy increases. The final weight is obtained by dynamically combining the basic weight and the entropy factor, and the translation weight is: ; Rotation weight:
[0073] After assigning the weights, the objective function can be obtained:
[0074]
[0075] in is the robustness function, weight Dynamically updated according to the above hybrid strategy.
[0076] Step S4: In the sliding window, the state vector is represented as:
[0077]
[0078] in Represents the state of the wheel odometer, Indicates The inverse depth of the first observed feature point. is the number of all feature points in the sliding window.
[0079] The objective of optimizing the state estimation is to minimize the sum of squares of all measurement residuals within the sliding window:
[0080]
[0081] in, represents the residual of marginalized prior information, represents the wheel odometer measurement residual, Represents the visual reprojection residual.
[0082] For the visual-wheel odometry system, the wheel odometry provides the robot's displacement, heading angle, and rotation information, but when using the wheel odometry, it cannot directly provide accurate acceleration and angular velocity information like the IMU. Therefore, the back-end optimization needs to rely on the fusion of the visual system and the wheel odometry to optimize the robot's posture. By removing the initial errors of the wheel odometry and the visual system from the initialization estimate, the back-end optimizer only needs to optimize the 3 degrees of freedom in the plane.
[0083] When loop closure successfully matches the historical keyframe and the current image frame, the feature points of the historical keyframe will be sent to the pose estimator. The pose estimator estimates the pose of the historical keyframe in the current sliding window based on the associated feature points. This estimation result is passed to the backend optimization module as a constraint.
Claims
1. A SLAM positioning system integrating monocular vision and a novel wheel odometer is specifically designed. A positioning system integrating a wheel odometer with three driven omnidirectional wheel sensors and monocular vision is specifically designed. The system is characterized by: S1. Designed a three-wheeled wheel odometer sensor composed of three independently usable omnidirectional wheel sensors; S2, added a dynamic plane constraint adaptive method based on local geometric features and meta-learning to constrain the model; S3. A hybrid strategy based on motion component decomposition and dynamic adjustment of residual entropy is designed to cope with sudden interference in the environment.
2. The three-wheeled odometer sensor in S1 according to claim 1, characterized in that: The driven wheel sensor is composed of three minimum units, each of which can be used independently and consists of omnidirectional wheel fixing frames 5 and 4, where 1 is a schematic diagram of the omnidirectional wheel. An axis passes through the middle of the omnidirectional wheel and is fixed to the fixing frame through a bearing. A bipolar magnet is fixed to the end of the axis, and the wheel speed information is read by a magnetic encoding sensor. The slide rail fixing frame 4 is connected to the slide rail, and the slider 3 cooperates with the slide rail and is fixed to the connecting frame 2. The connecting frame 2 is fixed to the robot chassis by bolts. There is a spring on the omnidirectional wheel fixing frame 4 to provide tension, which presses the omnidirectional wheel part against the ground and provides sufficient friction.
3. The method for dynamic plane constraint adaptation based on local geometric features and meta-learning in S2 of claim 1, characterized in that: Drawing on the idea of adaptive smoothing terms in stereo matching, the constraint strength is dynamically adjusted by analyzing geometric features such as local curvature and gradient direction of the plane. The implementation method is as follows: a local feature classifier based on a Gaussian mixture model is designed to calculate the regional geometric complexity in real time and map it to a constraint weight function to ensure that the constraint strength matches the local structure. For online parameter optimization guided by meta-learning, an environmental adaptive learning mechanism is introduced, and a lightweight neural network is pre-trained using meta-learning to quickly predict the optimal constraint parameters in new scenarios. The advantage of this method is that it solves the problem of traditional methods relying on manual parameter adjustment and improves the generalization ability of unknown plane scenes. Then, combined with the adaptive constraint strategy in nonlinear system control, a dual threshold relaxation mechanism is designed: (1) Main constraint threshold: ensuring safety margin; (2) Sub-constraint threshold: Dynamically adjusted according to real-time error feedback, allowing local breakthroughs of constraints within a safe range to optimize global performance.
4. The hybrid strategy based on motion component decomposition and dynamic adjustment of residual entropy in S3 of claim 1, characterized in that: This method can avoid the coupling of translation and rotation errors, allocate weights in a targeted manner, and perceive sensor degradation through real-time residual distribution, effectively solving the complementarity problem between wheel odometer and vision in structured or unstructured scenes. The specific solution is as follows: Experimentally decompose the motion state into translational components and the rotational component , respectively assign basic weights: Translation weight: Wheel odometry is more robust to the translation component and is given a higher base weight: ; Rotation weight: Visual feature points are more sensitive to rotation, giving higher visual weight: ; Then adjust the dynamic residual entropy and the visual residual The chi-square norm of the feature point reprojection error is expressed as: , is the feature point in the world coordinate system, is the feature point in the image coordinate system; the wheel odometry residual is calculated by the difference between the motion prediction and the wheel odometry measurement: , statistically analyze the residual distribution within the sliding window and calculate its information entropy to quantify the uncertainty: ; in is the normalized residual histogram probability. The larger the entropy value, the lower the current confidence of the sensor. The entropy adjustment factor is calculated for the translation and rotation components respectively: ; in To adjust the parameters and ensure that the weight decays exponentially when the entropy increases, the final weight is obtained by dynamically combining the basic weight and the entropy factor, and the translation weight is: ; Rotation weight: ; After assigning the weights, the objective function can be obtained: ; in is the robustness function, weight Dynamically updated according to the above hybrid strategy.
Citation Information
Cited By
Automatic labeling method and device, electronic equipment and storage medium
CN121330425A
Multi-sensor fusion positioning method and system based on feature enhancement and dynamic weighting
CN121612275A