Calculation method of mobile robot reinforcement learning high-precision speedometer
By training lightweight neural networks in virtual environments, the problems of high hardware dependence, high cost and poor environmental adaptability of the mobile robot odometer solution are solved, and high-precision, low-cost, real-time multimodal robot positioning is achieved, which improves the stability of autonomous motion control and positioning reliability in complex environments.
Patent Information
- Application Number
- CN202510504954.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-12
AI Technical Summary
The existing mobile robot odometer scheme relies on high-precision hardware, is expensive, and is not robust in complex environments, has high computational complexity, and is single environmental adaptability, making it difficult to adapt to multimodal mobile robots and dynamic motion scenarios.
Using reinforcement learning and physical simulation methods, we train lightweight neural networks in a virtual environment, and use multimodal input data to build an end-to-end motion state mapping model, replacing the traditional multi-sensor fusion solution, and achieving high-precision and low-cost odometer calculations.
It reduces hardware costs to 1/10 of the traditional solution, improves positioning accuracy and robustness in unstructured environments, adapts to multimodal robots, reduces computing complexity, and supports high-frequency real-time output.
Smart Images

Figure CN120470892A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mobile robot technology, and in particular to a method for calculating a high-precision odometer for a mobile robot through reinforcement learning, which is used to implement high-precision odometer calculation for a mobile robot under a reinforcement learning framework, and provide technical support for precise positioning and navigation of the mobile robot. Background Art
[0002] In the field of mobile robotics, the odometry system is a core component of positioning and navigation, and its accuracy directly affects the robot's path planning, obstacle avoidance control, and task execution capabilities. The current mainstream odometry solutions for wheeled robots, legged robots, and aircraft are all based on multi-sensor fusion technology, achieving data resolution through the combination of high-precision hardware devices and classical filtering algorithms. Specifically, such solutions generally rely on encoders (such as incremental photoelectric encoders) to obtain basic displacement information of the robot's motion, using inertial measurement units (IMUs) to perceive attitude angles and accelerations in real time, supplemented by laser radar (LiDAR) or visual sensors to build environmental feature maps. Finally, multi-source data is fused through algorithms such as extended Kalman filters (EKFs), unscented Kalman filters (UKFs), or particle filters (PFs), outputting odometry data (such as position, velocity, attitude angles, etc.) at different control frequencies, providing real-time motion state feedback for the robot control system. In the prior art, Chinese patent CN115479598A discloses a multi-sensor fusion-based positioning and mapping method and tightly coupled system, which is applied to the field of real-time positioning and mapping for mobile robots. The method involves acquiring lidar data, IMU inertial measurement data, wheel odometry data, and GPS data, then constructing a tightly coupled laser inertial odometry to fuse the relative and absolute measurement data from the different sensors. The fusion results are then subjected to sliding window optimization, loop closure detection, and global pose optimization to complete the positioning and map construction.
[0003] Although the current solution achieves high positioning accuracy in structured environments, it still has the following problems in practical applications: (1) High hardware dependence and high cost: To meet high-precision requirements, high-precision encoders (such as photoelectric encoders with more than 2000 lines), high-precision IMUs (zero-bias stability <5° / h) and high-frame-rate lidars (such as mechanical LiDARs with more than 16 lines) are required. As a result, the cost of a single odometer system generally exceeds 20,000 yuan, making it difficult to adapt to low-cost commercial robot scenarios. (2) Insufficient robustness in complex environments: In unstructured environments (such as uneven roads, soft sand, strong light or low-texture scenes), the lidar point cloud matching error increases significantly, the IMU cumulative drift problem is aggravated, and the encoder is affected by wheel slippage, legged robot foot slippage or aircraft airflow disturbance, resulting in reduced Kalman filter convergence and positioning drift or even divergence. (3) The contradiction between computational complexity and real-time performance: The traditional Kalman filter algorithm needs to maintain a high-dimensional state space (such as a 15-dimensional state vector containing position, velocity, and attitude angle). In high-frequency control scenarios (such as above 200 Hz), the processor computing power consumption exceeds the carrying capacity of the embedded platform (such as ARM chip), resulting in data processing delays and affecting the robot control accuracy. (4) Single environmental adaptability: Existing solutions are designed for specific types of robots (e.g., wheeled robots assume pure wheel rolling, and legged robots rely on gait models). They are difficult to be compatible with multimodal mobile robots (e.g., wheel-legged hybrid robots) or dynamic motion scenarios (e.g., aircraft take-off and landing, and legged robots switching gaits when overcoming obstacles), and lack the ability to generalize complex motion patterns.
[0004] In summary, existing odometry calculation schemes rely too much on high-precision hardware and specific environmental assumptions, and are deficient in cost, robustness, real-time performance, and generalization capabilities. There is an urgent need for a high-precision odometry calculation method that does not rely on ultra-precision sensors and can adapt to complex dynamic environments. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned shortcomings and provide a calculation method for a high-precision odometer of a mobile robot through reinforcement learning, so as to overcome the problems of large cumulative error and low positioning accuracy of traditional mobile robot odometers caused by sensor noise, motion model errors and environmental interference in complex dynamic environments, while avoiding reliance on high-cost external positioning equipment or complex multi-sensor fusion solutions, and realizing a high-precision, adaptive, low-cost real-time odometer calculation method based on reinforcement learning, thereby improving the stability of autonomous motion control and long-term positioning reliability of the robot in dynamic unknown scenarios.
[0006] The object of the present invention is achieved like this: A method for calculating high-precision odometer for mobile robot reinforcement learning includes the following: S1, multi-modal input data, building a physical simulation environment; The multimodal input data includes 32-dimensional robot body state data, 24-dimensional environmental contact force data, and 8-dimensional historical odometry data; Build virtual models of various types of robots based on physics engines such as Genesis and ISSACGYM. These robots include wheeled robots, legged robots, and aircraft. Accurately simulate the robot's dynamic characteristics and environmental interactions to create a physical simulation environment. Robot dynamic characteristics include mass distribution, joint friction, and contact mechanics. Environmental interaction data includes ground friction coefficient, airflow disturbances, and terrain undulations. S2, reinforcement learning training mechanism; A two-stage training strategy is adopted, which is mainly based on supervised learning driven by simulation data and supplemented by reinforcement learning: S2.1, supervised learning pre-training; Millions of training samples are generated in a structured / unstructured virtual environment. The input is the 64-dimensional feature vector mentioned above, and the output is the odometer data actually solved by the simulation engine. The loss function uses mean square error + regularization: ; Among them, f is the MLP network, x i is the input feature, y i is the real odometer data, θ is the network parameter, and λ is the regularization coefficient; S2.2, reinforcement learning fine-tuning; Introducing a proximal policy optimization algorithm to optimize policies under extreme working conditions, including wheel slippage, legged robots hanging in the air, and aircraft subjected to strong wind disturbances; S2.3. Through two-stage training, the network automatically learns the contact force-motion state mapping rules under different environments, solving the dependence of traditional Kalman filtering on model assumptions; S3, lightweight network architecture; S3.1. Dimensionality Compression: A 64-dimensional input layer is used to integrate multimodal features, and a 32-dimensional hidden layer is used to extract core interaction features to avoid overfitting caused by high-dimensional state space. S3.2. Activation function selection: The hidden layer uses the Swish activation function , while maintaining the nonlinear fitting capability, the computational complexity is reduced; S3.3. Output layer design: Dynamically adjust the output dimension according to the robot type, and use a single network to be compatible with multiple robot platforms to solve the problem of poor modal adaptability of traditional solutions.
[0007] Furthermore, the robot body state data in step S1 includes the following parameters: Kinematic parameters: angular velocity and angular displacement of each driving joint, including the driving wheel speed of wheeled robots, the hip joint angle of legged robots, and the motor speed of aircraft; Dynamic parameters: center of mass linear velocity, angular velocity, and torque feedback of each joint; Attitude parameters: real-time attitude angles based on quaternions.
[0008] Furthermore, the environmental contact force data in step S1 includes the following data: Wheeled robots: contact force vectors of the left and right drive wheels, forces and torques in the x / y / z directions, reflecting ground adhesion and slippage; Legged robots: Contact force tensor for each foot, 6 dimensions per foot, including normal force, tangential friction force, and contact torque, representing the foot-ground interaction state; Aircraft: fuselage aerodynamic drag / lift vector and propeller anti-torque.
[0009] Furthermore, the historical odometer data in step S1 includes the following data: the position (x, y, z), speed (vx, vy, vz), and attitude angle (yaw) at the previous moment, and a time series dependency relationship is constructed to achieve dynamic time series modeling through a sliding window.
[0010] Furthermore, the label error of the odometer data actually solved by the simulation engine is output in step S2.1 .
[0011] Furthermore, the strategy optimization in step S2.2 includes the following: State space: current network output odometry error; Action space: adjust the activation function parameters of the network hidden layer; Reward function: With error convergence speed and solution delay as optimization goals, the network is made adaptive to different motion modes, such as the transition from translation to climbing for wheeled robots and from quadrupedal gait to bipedal stance for legged robots.
[0012] Furthermore, in step S3.3, the output dimension is dynamically adjusted according to the robot type: 6-dimensional pose for wheeled or legged robots, and 12-dimensional state space for aircraft.
[0013] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a method for calculating high-precision odometers for mobile robots through reinforcement learning, constructs an end-to-end motion state mapping model, and replaces traditional multi-sensor fusion solutions that rely on high-precision hardware by training lightweight neural networks in a virtual simulation environment. (1) This invention breaks through the bottleneck of traditional odometer technology from multiple dimensions by building a lightweight data-driven model through the fusion of reinforcement learning and physical simulation: it relies on the robot state and contact force data solved in the simulation environment to replace high-precision encoders, IMUs and lidars, and only requires low-cost force sensors or kinematic inversion input features. The hardware cost is reduced to 1 / 10 of the traditional solution (less than 2,000 yuan), and no complex calibration is required, which significantly improves the feasibility of large-scale deployment of commercial robots; (2) The reinforcement learning training of the present invention covers extreme working conditions such as wheel slippage, foot suspension, and strong wind disturbance. By fitting the nonlinear mapping relationship between contact force and motion state, the positioning error in unstructured environments is reduced by more than 60%, and the divergence problem of the Kalman filter's dependent model assumptions is completely solved; the two-layer MLP network takes less than 5μs for single inference, supports high-frequency output above 200Hz, and has a computing power requirement so low that it can be supported by embedded platforms, making it suitable for power-sensitive scenarios such as drones; (3) The present invention unifies the 64-dimensional input features and the dynamic output layer design, making a single network compatible with multimodal robots such as wheeled, legged, and aerial vehicles. The domain randomization technology improves the robustness of the real environment by 40% without the need for targeted training. The model trained based on simulated real data directly learns the "force-motion" causal relationship, avoiding the drift accumulation of IMU integration and lidar matching. The attitude angle error is stable within 0.3° and the position error is less than 3cm in long-term missions, providing reliable support for high-precision navigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flow chart of the calculation method of the mobile robot reinforcement learning high-precision odometer of the present invention. DETAILED DESCRIPTION
[0015] To better understand the technical solution of the present invention, the following detailed description is provided with reference to the relevant illustrations. It should be understood that the following specific embodiments are not intended to limit the specific implementation of the technical solution of the present invention; they are merely examples of possible implementations of the technical solution of the present invention. It should be noted that references herein to the positional relationships of various components, such as component A being located above component B, are based on the relative positions of the components in the illustrations and are not intended to limit the actual positional relationships of the components.
[0016] See also Figure 1 , Figure 1 A flow chart of a method for calculating a high-precision odometer using reinforcement learning for a mobile robot according to Example 1 is drawn. As shown in the figure, a method for calculating a high-precision odometer using reinforcement learning for a mobile robot according to Example 1 includes the following contents: S1, multi-modal input data, building a physical simulation environment; The network input consists of three core data types to construct a complete feature space for robot-environment interaction: S1.1. Robot state data (32 dimensions) Kinematic parameters: angular velocity / angular displacement of each driving joint (wheeled robot driving wheel speed, legged robot hip joint angle, aircraft motor speed); Dynamic parameters: center of mass linear velocity / angular velocity, torque feedback of each joint (obtained in real time through the simulation engine); Attitude parameters: real-time attitude angle based on quaternion (replacing traditional IMU integral calculation to avoid drift accumulation); S1.2. Environmental contact force data (24 dimensions) Wheeled robots: contact force vectors of the left and right drive wheels (forces and torques in the x / y / z directions, reflecting ground adhesion and slippage); Legged robots: contact force tensor for each foot (6 dimensions / foot, including normal force, tangential friction force, and contact torque, representing the foot-ground interaction state); Aircraft: fuselage aerodynamic drag / lift vector and propeller anti-torque (replacing lidar environmental perception and adapting to low-texture scenes); S1.3. Historical odometer data (8 dimensions) The previous moment's position (x, y, z), velocity (vx, vy, vz), and attitude angle (yaw) are used to build time series dependencies (dynamic time series modeling is achieved through sliding windows); S1.4. Build virtual models of various robot types (wheeled, legged, and aerial) based on physics engines such as Genesis and ISSAC GYM, accurately simulating the robot's dynamic characteristics (mass distribution, joint friction, contact mechanics) and environmental interactions (ground friction coefficient, airflow disturbances, and terrain undulations). Through the above-mentioned multimodal input data, the external sensor measurements that rely on encoders, IMUs, and lidars in traditional solutions are converted into internal force characteristics of robot-environment interactions that can be accurately obtained in a simulation environment, thus eliminating the dependence on high-precision hardware in principle.
[0017] S2, reinforcement learning training mechanism; A two-stage training strategy is adopted, which is mainly based on supervised learning driven by simulation data and supplemented by reinforcement learning: S2.1. Supervised Learning Pre-training Generate millions of training samples in a structured / unstructured virtual environment, with the input being the above 64-dimensional feature vector and the output being the actual odometer data calculated by the simulation engine (label error ); The loss function uses mean square error (MSE) + regularization: ; Among them, f is the MLP network, x i is the input feature, y i is the real odometer data, θ is the network parameter, and λ is the regularization coefficient; S2.2. Reinforcement Learning Fine-tuning The Proximal Policy Optimization (PPO) algorithm is introduced to optimize policies under extreme working conditions (such as wheel slippage, a legged robot hanging in the air, and strong wind disturbances on an aircraft): State space: current network output odometry error; Action space: Adjust the parameters of the network hidden layer activation function (such as ReLU slope); Reward function: Optimizes error convergence speed and solution latency to enable the network to adapt to different motion modes (e.g., from translation to climbing for wheeled robots, from quadrupedal gait to bipedal stance for legged robots). S2.3. Through two-stage training, the network can automatically learn the contact force-motion state mapping rules in different environments, solve the traditional Kalman filter's dependence on model assumptions (such as pure wheel rolling and small-angle posture changes), and improve the robustness in unstructured environments.
[0018] S3, lightweight network architecture; A two-layer MLP architecture (64→32→output dimensions) is used, and its design logic is as follows: S3.1. Dimensionality Compression: A 64-dimensional input layer is used to integrate multimodal features, and a 32-dimensional hidden layer is used to extract core interaction features (such as the coupling relationship between contact force and joint torque) to avoid overfitting caused by high-dimensional state space. S3.2. Activation function selection: The hidden layer uses the Swish activation function , while maintaining nonlinear fitting capabilities, reducing computational complexity (compared to ReLU, the gradient vanishing problem is better, and compared to Sigmoid, the gradient calculation is simpler); S3.3. Output layer design: Dynamically adjust the output dimension according to the robot type (6-dimensional posture for wheeled / legged robots, 12-dimensional state space for aircraft), and use a single network to be compatible with multiple robot platforms to solve the problem of poor modal adaptability of traditional solutions.
[0019] Specific implementation cases: 1. Simulation modeling: Build a 3D model of the target robot in the physics engine and define the mass, inertia tensor, and joint kinematic parameters; 2. Data Acquisition: Drive the robot to move randomly in a virtual environment (e.g., S-shaped paths for wheeled robots, random gaits for legged robots, and tumbling maneuvers for aircraft), while simultaneously recording 64-dimensional input features and real-world odometry data. 3. Network training: Use PyTorch / TensorFlow to train the MLP, with a batch size of 128 for 500 epochs in the supervised learning phase and 2048 steps / update cycle in the reinforcement learning phase; 4. Online deployment: Convert the trained model to ONNX format, embed it into the robot controller, input real-time state and contact force data (obtained in the real environment through low-cost force sensors or kinematic inversion), and output high-frequency odometry data (supporting solution frequencies above 200Hz). Through the above steps, "simulation training instead of hardware calibration" and "data-driven replacement of model assumptions" are realized, providing a new technical path for low-cost, highly robust mobile robot odometry.
[0020] Working principle: This invention provides an odometer calculation method based on reinforcement learning and physical simulation. The overall process is as follows: (1) Construction of physical simulation environment: Building virtual models of various types of robots (wheeled / legged / aircraft) based on physics engines such as Genesis and ISSAC GYM, accurately simulating the robot's dynamic characteristics (mass distribution, joint friction, contact mechanics) and the interaction with the environment (ground friction coefficient, airflow disturbance, terrain undulation); (2) Reinforcement learning training framework: A training strategy combining supervised learning and reinforcement learning is used, with real odometry data (position, velocity, and attitude angle) output from the simulation environment as labels to drive the neural network to learn the nonlinear mapping relationship from the robot state space to the odometry output space; (3) Lightweight network design: Construct a two-layer fully connected MLP (multi-layer perceptron), with an input layer dimension of 64 (integrating kinematic state and contact force characteristics), a hidden layer dimension of 32, and an output layer corresponding to the odometer data dimension (such as 6 dimensions for wheeled robots: x / y / z position, roll / pitch / yaw attitude angles), to achieve real-time solution with low computing power consumption.
[0021] Several core principles involved in the present invention are as follows: The causal relationship between contact mechanics and motion state: During robot motion, contact forces (such as foot support and wheel friction) are key variables connecting the body's motion with environmental constraints. Physical simulation can accurately model the "force-motion" mapping relationship, while traditional solutions rely on the "position-velocity" differential relationship (which is susceptible to noise interference).
[0022] Data-driven nonlinear modeling: The MLP network is trained with simulation data and can fit complex nonlinear functions in high-dimensional space (such as the relationship between contact force mutation and odometry error during slip), replacing the linearization assumptions of traditional Kalman filtering (such as the Taylor expansion approximation of the state transition matrix by the EKF). Environmental generalization of transfer learning: Networks trained in virtual environments can automatically adapt to parameter differences in the real environment (such as ground friction coefficient deviation and sensor noise) through domain randomization technology, without the need to recalibrate hardware parameters. The technical problems solved by the solution provided by the present invention are as follows: 1. To address the existing issues of high hardware dependency and high cost, the present invention uses input data derived from built-in physical parameters in the simulation environment. This eliminates the need for high-precision encoders / IMUs and only requires low-cost force sensors (or simulated force data mapping). 2. To address the problem of insufficient robustness in complex environments in existing technologies, the present invention uses reinforcement learning training to cover extreme working conditions such as wheel slip, foot sliding, and airflow disturbances. The network automatically learns the mapping rules under abnormal conditions. 3. To solve the problem of high computational complexity in the prior art, the two-layer MP network of the present invention has low computing power requirements (a single inference process only requires about , floating-point operations), adapted to embedded platforms such as ARM / Movidius; 4. To solve the problem of single environmental adaptability in the existing technology, the present invention unifies the input features (state + contact force) to be compatible with wheeled / legged / aircraft vehicles, and realizes the universality of multimodal robots by adjusting the output layer dimension.
[0023] The above are only specific application examples of the present invention and do not constitute any limitation on the scope of protection of the present invention. Any technical solutions formed by equivalent transformation or equivalent replacement shall fall within the scope of protection of the present invention.
Claims
1. A method for calculating high-precision odometer using reinforcement learning for a mobile robot, characterized in that: Includes the following: S1, multi-modal input data, building a physical simulation environment; The multimodal input data includes 32-dimensional robot body state data, 24-dimensional environmental contact force data, and 8-dimensional historical odometry data; Build virtual models of various types of robots based on physics engines such as Genesis and ISSAC GYM. These robots include wheeled robots, legged robots, and aircraft. Accurately simulate the robot's dynamic characteristics and environmental interactions to create a physical simulation environment. Robot dynamic characteristics include mass distribution, joint friction, and contact mechanics. Environmental interaction data includes ground friction coefficient, airflow disturbances, and terrain undulations. S2, reinforcement learning training mechanism; A two-stage training strategy is adopted, which is mainly based on supervised learning driven by simulation data and supplemented by reinforcement learning: S2.1, supervised learning pre-training; Millions of training samples are generated in a structured / unstructured virtual environment. The input is the 64-dimensional feature vector mentioned above, and the output is the odometer data actually solved by the simulation engine. The loss function uses mean square error + regularization: ; Among them, f is the MLP network, x i is the input feature, y i is the real odometer data, θ is the network parameter, and λ is the regularization coefficient; S2.2, reinforcement learning fine-tuning; Introducing a proximal policy optimization algorithm to optimize policies under extreme working conditions, including wheel slippage, legged robots hanging in the air, and aircraft subjected to strong wind disturbances; S2.
3. Through two-stage training, the network automatically learns the contact force-motion state mapping rules under different environments, solving the dependence of traditional Kalman filtering on model assumptions; S3, lightweight network architecture; S3.
1. Dimensionality Compression: A 64-dimensional input layer is used to integrate multimodal features, and a 32-dimensional hidden layer is used to extract core interaction features to avoid overfitting caused by high-dimensional state space. S3.
2. Activation function selection: The hidden layer uses the Swish activation function , while maintaining the nonlinear fitting capability, the computational complexity is reduced; S3.
3. Output layer design: Dynamically adjust the output dimension according to the robot type, and use a single network to be compatible with multiple robot platforms to solve the problem of poor modal adaptability of traditional solutions.
2. The method for calculating high-precision odometer for mobile robot reinforcement learning according to claim 1, characterized in that: The robot body state data in step S1 includes the following parameters: Kinematic parameters: angular velocity and angular displacement of each driving joint, including the driving wheel speed of wheeled robots, the hip joint angle of legged robots, and the motor speed of aircraft; Dynamic parameters: center of mass linear velocity, angular velocity, and torque feedback of each joint; Attitude parameters: real-time attitude angles based on quaternions.
3. The method for calculating a high-precision odometer for a mobile robot using reinforcement learning according to claim 1, characterized in that: The environmental contact force data in step S1 includes the following data: Wheeled robots: contact force vectors of the left and right drive wheels, forces and torques in the x / y / z directions, reflecting ground adhesion and slippage; Legged robots: Contact force tensor for each foot, 6 dimensions per foot, including normal force, tangential friction force, and contact torque, representing the foot-ground interaction state; Aircraft: fuselage aerodynamic drag / lift vector and propeller anti-torque.
4. The method for calculating a high-precision odometer for a mobile robot using reinforcement learning according to claim 1, characterized in that: The historical odometer data in step S1 includes the following data: the position (x, y, z), speed (vx, vy, vz), and attitude angle (yaw) at the previous moment. Time series dependencies are constructed and dynamic time series modeling is achieved through sliding windows.
5. The method for calculating a high-precision odometer for a mobile robot through reinforcement learning according to claim 1, characterized in that: Output the label error of the odometer data actually solved by the simulation engine in step S2.1 .
6. The method for calculating high-precision odometer for mobile robot reinforcement learning according to claim 1, characterized in that: The strategy optimization in step S2.2 includes the following: State space: current network output odometry error; Action space: adjust the activation function parameters of the network hidden layer; Reward function: With error convergence speed and solution delay as optimization goals, the network is made adaptive to different motion modes, such as the transition from translation to climbing for wheeled robots and from quadrupedal gait to bipedal stance for legged robots.
7. The method for calculating high-precision odometer for mobile robot reinforcement learning according to claim 1, characterized in that: In step S3.3, the output dimension is dynamically adjusted according to the robot type: 6-dimensional pose for wheeled or legged robots, and 12-dimensional state space for aircraft.
Citation Information
Patent Citations
Positioning and mapping method based on multi-sensor fusion and tight coupling system
CN115479598A
Cited By
Robot model training method, device, equipment and medium
CN121303202A