A mobile robot trajectory tracking control method and system
Through multi-sensor data fusion and deep reinforcement learning optimization sliding mode control methods, the problem of insufficient parameters fixed and disturbance processing in mobile robot trajectory tracking control is solved, and high-precision and robust trajectory tracking control is achieved, which improves the dynamic response performance and robustness of the system.
Patent Information
- Application Number
- CN202510542243.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing mobile robot trajectory tracking control methods have problems such as fixed parameters, poor real-time adaptability, and insufficient processing of complex system disturbances (such as body friction, drag chain coupling and external environment interference), making it difficult to achieve high-precision and robust trajectory tracking in dynamic and complex environments.
Through multiple sensors, the motion state data of the robot body and drag part are collected, and the total disturbance value of the system is calculated after preprocessing, a trajectory tracking control model is constructed, and the control strategy is adjusted in real time to compensate for tracking errors and drag swings.
It significantly improves the dynamic response speed and control robustness of the system, reduces high-frequency vibration, and improves overall control accuracy and stability.
Smart Images

Figure CN120065757B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile robot trajectory tracking control, and in particular to a mobile robot trajectory tracking control method and system. Background Art
[0002] In recent years, with the rapid development of industrial automation and intelligent manufacturing, mobile robotics has been widely used in logistics, warehousing, security, and other fields. Leveraging cutting-edge technologies such as multi-sensor fusion, advanced control algorithms, and deep learning, robots' environmental perception, path planning, and motion control capabilities have continued to improve. In particular, control schemes combining sliding mode control, model predictive control, and deep reinforcement learning have become a research hotspot. These technologies have achieved considerable success both theoretically and experimentally, providing strong support for achieving high-precision trajectory tracking.
[0003] However, the existing technology still has obvious shortcomings. First, traditional mobile robot trajectory tracking methods mostly rely on sliding mode or PID control with fixed parameters, which is difficult to adapt to dynamically changing environmental interference and system nonlinearity, resulting in insufficient tracking accuracy in complex scenarios. Secondly, although some studies have introduced deep learning algorithms, they usually only target a single compensation strategy and fail to fully utilize multi-type sensor data to achieve comprehensive calculation of disturbances, making it difficult to effectively suppress the combined disturbances caused by motor friction, trailer coupling and environmental interference. In addition, existing methods fail to deeply model the dynamic coupling characteristics of the trailer system and cannot simultaneously consider the dynamic effects of load changes and flexible deformation of the trailer chain, resulting in poor system stability and robustness. Finally, many control methods are insufficiently designed in terms of real-time control signal calculation and actuator response compensation, resulting in slow error convergence and prominent high-frequency jitter phenomena. These issues limit the ability of existing solutions to achieve high-precision and stable trajectory tracking in complex dynamic environments. The present invention addresses these shortcomings by integrating multi-sensor data fusion, disturbance calculation, multi-level sliding mode control, and online parameter adjustment of a deep reinforcement learning model to achieve coordinated compensation for trajectory tracking errors and trailer sway, thereby significantly improving the dynamic response performance and robustness of the system. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is: the existing mobile robot trajectory tracking control method has the problems of fixed parameters, poor real-time adaptability, and insufficient processing of complex system disturbances (such as body friction, tow chain coupling and external environmental interference), as well as the problem of how to achieve high-precision and robust trajectory tracking control in a dynamic and complex environment.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides a trajectory tracking control method for a mobile robot, comprising collecting raw motion state data of a robot body and a towed part through multiple types of sensors, preprocessing the raw motion state data, and generating preprocessed motion state data;
[0008] Calculating a total disturbance value of the system based on the preprocessed motion state data;
[0009] Constructing a trajectory tracking control model, inputting the pre-processed motion state data and the total disturbance value of the system, optimizing the sliding mode control parameters, and generating a control model after dynamic parameter adjustment;
[0010] According to the control model after dynamic parameter adjustment, the control signal of the main robot is calculated in real time, and the dynamic response compensation of the actuator is performed on the control signal of the main robot;
[0011] The trajectory tracking error and the trailer swing amplitude error after the execution of the main robot control signal are monitored, the control strategy is adjusted according to the error results, and the feedback is fed back to the trajectory tracking control model for adaptive optimization.
[0012] As a preferred embodiment of the mobile robot trajectory tracking control method of the present invention, the method of collecting original motion state data of the robot body and the trailer using multiple types of sensors includes collecting the motion state data of the robot body using a body sensor group, wherein the body sensor group includes an inertial measurement unit, a wheel speed encoder, and a steering motor current sensor;
[0013] The trailer sensor group collects the dynamic state data of the trailer part, and the trailer sensor group includes a trailer chain tension sensor, a trailer swing angle gyroscope, and a six-dimensional torque sensor of the hitch point;
[0014] External disturbance feature data is collected through an environment perception sensor group, which includes a binocular vision camera, a solid-state laser radar and a millimeter wave radar array.
[0015] As a preferred embodiment of the mobile robot trajectory tracking control method of the present invention, the preprocessing of the raw motion state data includes performing multi-source fusion filtering on the motion state data of the robot body collected by the body sensor group to extract the state vector of the robot body, wherein the state vector of the robot body includes position, velocity, heading angle, and steering motor current;
[0016] The tow chain tension and swing angle data collected by the tow sensor group are denoised, and the tow chain coupling disturbance characteristic matrix is constructed based on the six-dimensional torque data of the hitch point.
[0017] Perform spatiotemporal registration of images, laser point clouds, and radar data collected by the environmental perception sensor group, and generate environment-robot interaction feature tensors through spatial transformation;
[0018] The state vector of the robot body, the tow chain coupling disturbance characteristic matrix and the environment-robot interaction characteristic tensor are uniformly normalized, and pre-processed motion state data is output.
[0019] As a preferred solution of the mobile robot trajectory tracking control method of the present invention, wherein: the calculating the total disturbance value of the system based on the preprocessed motion state data includes constructing a multi-level linear extended state observer, wherein the multi-level linear extended state observer includes a first-level observer, a second-level observer, and a third-level observer;
[0020] The pre-processed state vector of the robot body is input into the first-level observer, and the body dynamic disturbance value is calculated in real time through the linear extended state observation equation. The body dynamic disturbance value includes the motor nonlinear friction and wheel-ground contact slip disturbance;
[0021] The pre-processed tow chain coupling disturbance characteristic matrix is input into the second-level observer, and the coupling disturbance value caused by the trailer inertia force and the flexible deformation of the coupling point is calculated in real time through the linear expansion state observation equation;
[0022] The preprocessed environment-robot interaction feature tensor is input into the third-level observer, which extracts key environmental disturbance features through the attention mechanism and calculates the environmental disturbance value.
[0023] The body dynamics disturbance value, coupling disturbance value and environmental disturbance value are integrated to generate the total disturbance value of the system.
[0024] As a preferred solution of the mobile robot trajectory tracking control method of the present invention, wherein: the constructing of the trajectory tracking control model includes establishing a sliding mode control framework and designing a sliding mode surface, wherein the sliding mode surface is composed of a trajectory tracking error integral term and an adaptive compensation term based on a total disturbance value of the system;
[0025] Constructing a sliding mode control law, wherein the sliding mode control law includes an equivalent control term and a switching control term, wherein the equivalent control term is generated based on a nominal dynamic model of the system and a total disturbance value of the system;
[0026] Integrate a deep reinforcement learning model to dynamically adjust sliding mode control parameters through a policy network. The sliding mode control parameters include switching gain, boundary layer thickness, and sliding surface weight coefficient.
[0027] A disturbance compensation mechanism is designed to embed the total disturbance value of the system into the equivalent control term.
[0028] As a preferred embodiment of the mobile robot trajectory tracking control method of the present invention, generating a control model after dynamic parameter adjustment includes inputting the trajectory tracking error, the trailer swing amplitude, and the total system disturbance value into a deep reinforcement learning model to construct a multidimensional state space;
[0029] Based on the double-delayed deep deterministic policy gradient algorithm, the adjustment amount of the sliding mode control parameters is output, including the switching gain attenuation rate and the boundary layer thickness proportional coefficient;
[0030] Dynamically calculate the weight coefficient of the integral term in the sliding surface based on the current movement speed and the mass of the towed load;
[0031] The equivalent control items, switching control items and parameter adjustment results are integrated to generate a control model after dynamic parameter adjustment, and the convergence of the control model after dynamic parameter adjustment is verified by the Lyapunov stability criterion.
[0032] As a preferred solution of the mobile robot trajectory tracking control method described in the present invention, the actuator dynamic response compensation includes motor torque limitation and steering angle constraint.
[0033] As a preferred embodiment of the mobile robot trajectory tracking control method of the present invention, the adaptive optimization includes collecting the actual trajectory data and the trailer swing angle data after the main robot control signal is executed in real time, and calculating the trajectory tracking error and the trailer swing amplitude error at the current moment;
[0034] Set dynamic error thresholds. When the trajectory tracking error exceeds the first threshold or the trailer swing amplitude error exceeds the second threshold, the online parameter re-optimization process of the deep reinforcement learning module is triggered to update the sliding mode control parameters.
[0035] If the error continues to exceed the limit and does not converge after optimization, it switches to the emergency control mode and uses the preset conservative sliding mode parameters and fixed boundary layer thickness;
[0036] The updated control parameters or emergency control instructions are fed back to the trajectory tracking control model to form a closed-loop adaptive optimization.
[0037] In a second aspect, an embodiment of the present invention provides a mobile robot trajectory tracking control system, comprising:
[0038] Data acquisition module: collects the original motion state data of the robot body and the towing part through multiple types of sensors, pre-processes the original motion state data, and generates pre-processed motion state data;
[0039] Disturbance calculation module: calculates the total disturbance value of the system based on the pre-processed motion state data;
[0040] Trajectory tracking control model construction module: constructs a trajectory tracking control model, inputs the pre-processed motion state data and the total disturbance value of the system, optimizes the sliding mode control parameters, and generates a control model after dynamic parameter adjustment;
[0041] Real-time control signal generation module: calculates the control signal of the main robot in real time according to the control model after dynamic parameter adjustment, and performs actuator dynamic response compensation on the control signal of the main robot;
[0042] Adaptive optimization feedback module: monitors the trajectory tracking error and the trailer swing amplitude error after the execution of the main robot control signal, adjusts the control strategy according to the error results, and feeds back to the trajectory tracking control model for adaptive optimization.
[0043] Beneficial effects of the present invention: The present invention obtains fine state information through multi-sensor data fusion, and uses a multi-level extended state observer to calculate the total disturbance value of the system. Combined with the sliding mode control method based on online parameter adjustment of the deep reinforcement learning model, the present invention realizes the coordinated compensation of trajectory tracking error and trailer swing, thereby significantly improving the dynamic response speed and control robustness of the system, reducing high-frequency jitter, and improving the overall control accuracy and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:
[0045] Figure 1 This is an overall flow chart of a mobile robot trajectory tracking control method provided by the first embodiment of the present invention. DETAILED DESCRIPTION
[0046] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0047] Example 1, with reference to Figure 1 , as one embodiment of the present invention, provides a mobile robot trajectory tracking control method, comprising:
[0048] S1: collecting original motion state data of the robot body and the towing part through multiple types of sensors, preprocessing the original motion state data, and generating preprocessed motion state data.
[0049] S11: collecting motion state data of the robot body through a body sensor group, wherein the body sensor group includes an inertial measurement unit, a wheel speed encoder, and a steering motor current sensor;
[0050] The trailer sensor group collects the dynamic state data of the trailer part, and the trailer sensor group includes a trailer chain tension sensor, a trailer swing angle gyroscope, and a six-dimensional torque sensor of the hitch point;
[0051] External disturbance feature data is collected through an environment perception sensor group, which includes a binocular vision camera, a solid-state laser radar and a millimeter wave radar array.
[0052] In this embodiment of the present invention, the "property sensor set" includes an inertial measurement unit (IMU), a wheel speed encoder, and a steering motor current sensor. The IMU (Inertial Measurement Unit) measures the robot's acceleration and angular velocity in real time, providing high-frequency, high-precision motion information. The wheel speed encoder accurately reflects wheel speed through pulse output, enabling calculation of the robot's travel distance and speed. The steering motor current sensor reflects changes in motor load, indirectly detecting steering quality. The output parameters of these three devices (IMU, encoder, and current sensor) are fused to form the robot's property state vector, providing data support for subsequent error calculation and controller design.
[0053] For the towing system, the present invention utilizes a towing sensor assembly, whose core sensors include a towing chain tension sensor, a trailer yaw gyroscope, and a six-dimensional torque sensor at the hitch point. The towing chain tension sensor captures load fluctuations in the drive chain, reflecting the stress conditions of the towing system; the trailer yaw gyroscope accurately measures the trailer's lateral swing angle, directly detecting dynamic instability; and the six-dimensional torque sensor at the hitch point decomposes forces and torques in various directions, providing comprehensive mechanical coupling information. These parameters, forming the towing state parameters, not only complete the robot's overall motion state but also make the system more sensitive and responsive to complex disturbances caused by the flexible connection.
[0054] The environmental perception sensor suite consists of a binocular vision camera, a solid-state lidar, and a millimeter-wave radar array, each offering unique advantages. The binocular vision camera captures a stereoscopic image of the scene, enabling depth calculation of obstacles and the surrounding environment. The solid-state lidar provides high-resolution 2D or 3D point cloud data, facilitating real-time environmental modeling. The millimeter-wave radar array operates reliably even in adverse weather conditions, compensating for the limitations of vision and lasers in rain and fog. After multi-sensor fusion, the resulting environmental disturbance signature data enables the system to perceive dynamic changes in the external environment in real time, enabling intelligent obstacle avoidance and interference compensation.
[0055] S12: performing multi-source fusion filtering on the motion state data of the robot body collected by the body sensor group to extract the state vector of the robot body, where the state vector of the robot body includes position, velocity, heading angle, and normalized steering motor current;
[0056] The tow chain tension and swing angle data collected by the tow sensor group are denoised, and the tow chain coupling disturbance characteristic matrix is constructed based on the six-dimensional torque data of the hitch point.
[0057] Perform spatiotemporal registration of images, laser point clouds, and radar data collected by the environmental perception sensor group, and generate environment-robot interaction feature tensors through spatial transformation;
[0058] The state vector of the robot body, the tow chain coupling disturbance characteristic matrix and the environment-robot interaction characteristic tensor are uniformly normalized to output preprocessed motion state data; the preprocessed motion state data includes the preprocessed state vector of the robot body, the preprocessed tow chain coupling disturbance characteristic matrix and the preprocessed environment-robot interaction characteristic tensor.
[0059] In this embodiment of the present invention, the robot's main sensor group collects data, including an inertial measurement unit (IMU), wheel speed encoders, and steering motor current sensors. These data digitally reflect the robot's acceleration, angular velocity, travel speed, travel distance, and motor load changes. Applying a multi-source fusion filtering algorithm (such as a Kalman filter or extended Kalman filter) to this data effectively suppresses noise and interference, extracting a more realistic and accurate robot main body state. The extracted robot main body state vector includes position, velocity, heading angle, and normalized steering motor current. These parameters can be directly quantified from the sensor signals. After normalization, the data from each channel are numerically aligned, improving the stability and robustness of subsequent data fusion and controller design. This step, through fusion filtering, generates a stable and low-noise state vector, providing a solid data foundation for subsequent system disturbance calculation and control law computation, thereby achieving higher-precision motion control in harsh dynamic environments.
[0060] Secondly, the trailer sensor system collects trailer chain tension data and trailer swing angle data to reflect the forces acting on the trailer and its dynamic posture changes. Furthermore, the multidimensional torque information provided by the six-dimensional torque sensor at the hitch point reveals the flexible deformation and inertial effects of the trailer chain in its mechanical coupling. To address this issue, the tension and swing angle data are first subjected to time-domain or frequency-domain denoising (such as wavelet denoising) to remove high-frequency interference. The torque data are then used to construct a trailer chain coupling perturbation characteristic matrix, which reflects the dynamic coupling relationships between the sensor data. The key to constructing this perturbation characteristic matrix is to uniformly map the dispersed sensor signals into a high-dimensional feature space, allowing the dynamic uncertainty of the trailer chain to be quantitatively expressed in matrix form. The output of this matrix helps to accurately capture the nonlinear coupling effects of the trailer in subsequent perturbation calculations using a multi-level extended state observer, thereby better implementing trailer dynamic compensation and overcoming the limitations of traditional methods that rely solely on a single signal.
[0061] Third, the images, laser point clouds, and radar data collected by the environmental perception sensor group each have their own advantages. For example, image data contains rich color and texture information, while laser radar and millimeter-wave radar provide precise information on distance and shape. By performing spatiotemporal registration on these data, the three-dimensional point cloud and image data are integrated using spatial transformations (such as Lie group SE(3) transformations) to generate a high-dimensional environment-robot interaction feature tensor. This tensor not only quantifies the spatial position of obstacles and the dynamic environment, but also captures the time-varying characteristics of interference factors in the environment, providing a key reference for the disturbance calculation module. Normalization ensures that the preprocessed motion state data is consistent in scale, facilitating subsequent data analysis and robust control algorithm application by the controller.
[0062] S2: Calculating a total disturbance value of the system based on the preprocessed motion state data.
[0063] Constructing a multi-level linear extended state observer, wherein the multi-level linear extended state observer includes a first-level observer, a second-level observer, and a third-level observer;
[0064] The pre-processed state vector of the robot body is input into the first-level observer, and the body dynamic disturbance value is calculated in real time through the linear extended state observation equation. The body dynamic disturbance value includes the motor nonlinear friction and wheel-ground contact slip disturbance;
[0065] The pre-processed tow chain coupling disturbance characteristic matrix is input into the second-level observer, and the coupling disturbance value caused by the trailer inertia force and the flexible deformation of the coupling point is calculated in real time through the linear expansion state observation equation;
[0066] The preprocessed environment-robot interaction feature tensor is input into the third-level observer, which extracts key environmental disturbance features through the attention mechanism and calculates the environmental disturbance value.
[0067] The body dynamics disturbance value, coupling disturbance value and environmental disturbance value are integrated to generate the total disturbance value of the system.
[0068] In step S2, the pre-processed motion state data is used as input, and the multi-level linear extended state observer (MLESO) is used to calculate the total disturbance value of the system in real time. First, the first-level observer takes the pre-processed state vector of the robot body as input. for:
[0069] ;
[0070] in, Indicates location, Indicates speed, represents the heading angle, represents the normalized steering motor current, Represents a transpose operation.
[0071] The expanded state vector Defined as:
[0072] ;
[0073] in, is the body dynamics disturbance value at the previous moment stored inside the observer;
[0074] Using the linear extended state observation equation, calculate the time derivative of the extended state vector , that is, the instantaneous rate of change of system state and disturbance:
[0075] ;
[0076] in, is the state transfer matrix, with the same dimension as The matching is obtained by linearizing the robot body dynamics model, describing the "natural" evolution of the extended state when there is no control input and no measurement correction; is the input matrix, which controls the input vector Mapping to the extended state space to reflect the impact of control actions on state and disturbance estimation; To control the input vector, such as motor torque command or drive current; is the observer gain matrix, which is used to convert the measurement error Feedback into extended state updates to correct model prediction bias; The actual measurement output vector usually contains physical quantities that can be directly obtained by the robot (such as ); As the output matrix, the expanded state vector Mapped to the measurement space.
[0077] It should be noted here that is the extended state vector The instantaneous rate of change at the current moment. In each sampling period, the observer will "accumulate" this rate of change - Continuously add to the current To generate the next moment .
[0078] This "calculate the change first, then accumulate it to the state" mechanism is like calculating the velocity first and then updating the position in numerical integration, which makes It can track the evolution of system states and disturbances in real time and adaptively.
[0079] The updated, latest body dynamics disturbance value is obtained :
[0080] ;
[0081] in, For the next moment (Updated ), is the mapping matrix.
[0082] This step effectively captures disturbance information such as motor nonlinear friction and wheel-ground slip. Its beneficial effect is that it provides an accurate compensation basis for subsequent towing and environmental disturbance calculations, solving the deficiency of traditional single sensors that cannot independently distinguish disturbance components.
[0083] Secondly, the second-level observer takes the pre-processed tow chain coupling disturbance characteristic matrix collected by the tow sensor group as input and decomposes the tow chain dynamics based on the virtual pendulum model. For each segment, an extended state vector is established. :
[0084] ;
[0085] in, For the The swing angle of the segmented tow chain, For the The rate of change of the swing angle of the segment of the drag chain, For the previous moment Estimated value of segment coupling disturbance;
[0086] The extended state vector is obtained by linearly expanding the state observation equation The derivative of :
[0087] ;
[0088] in, For the The state transfer matrix obtained after linearization of the segmented virtual pendulum; Correction gain matrix for the observer, used to feed the measurement error back into the state update; For the The measurement vector of the segment-drag chain coupling disturbance characteristic matrix, which contains the chain pendulum dynamic characteristics (such as tension, torque, etc.); is the output matrix, which maps the extended state to the measurement space;
[0089] The perturbation value of each segment is given by the mapping matrix Get, that is
[0090] ;
[0091] in, is the perturbation mapping matrix, used to Extract the coupled disturbance component from For the next moment (Updated );
[0092] All After the segmented disturbances are superimposed, the total disturbance value of the trailer system is obtained. :
[0093] ;
[0094] in, The total number of tow chain segments.
[0095] This design enables a quantitative description of the nonlinear coupling of trailer dynamics through model decomposition, significantly improving the calculation accuracy of the trailer inertia force and the flexible deformation disturbance of the hitch point, and solving the problem that previous methods only rely on single signal compensation.
[0096] Secondly, the third-level observer integrates the images, laser point clouds and radar data collected by the environmental perception sensor group to generate the pre-processed environment-robot interaction feature tensor As input, the attention mechanism is used for feature extraction. This process is done by calculating:
[0097] ;
[0098] in, Assign the attention mechanism to the Weight of environmental characteristics; is the attention weight matrix, which is used to perform a linear transformation on each feature vector; For the Column environment feature vector; is the Softmax function;
[0099] Calculate environmental disturbance values :
[0100] ;
[0101] in, is the environmental perturbation mapping matrix, which is used to map each column feature to a scalar perturbation estimate; For the Prediction of the disturbance contribution of each feature; is the final environmental disturbance value; To represent the environment-robot interaction feature tensor The number of columns, that is, the total number of environmental features (the number of candidate feature vectors).
[0102] This method uses softmax and attention mechanisms to prioritize the extraction of key disturbance information in the environment, solving the problems of large noise interference and data redundancy in traditional environmental disturbance calculations.
[0103] Calculate the total disturbance value of the system , the formula is:
[0104] ;
[0105] in, 、 、 They represent the body dynamic disturbance (such as motor friction, wheel-ground slip), the tow chain coupling disturbance (such as inertial force, flexible deformation) and the external environment disturbance (such as crosswind, ground undulation). 、 、 are the disturbance weight coefficients, which are based on the load quality, speed And the environmental sensitivity coefficient is dynamically adjusted to ensure accurate compensation for anti-motion.
[0106] S3: Construct a trajectory tracking control model, input the pre-processed motion state data and the total disturbance value of the system, optimize the sliding mode control parameters, and generate a control model after dynamic parameter adjustment.
[0107] Establishing a sliding mode control framework and designing a sliding surface, which is composed of a trajectory tracking error integral term and an adaptive compensation term based on the total disturbance value of the system;
[0108] Constructing a sliding mode control law, wherein the sliding mode control law includes an equivalent control term and a switching control term, wherein the equivalent control term is generated based on a nominal dynamic model of the system and a total disturbance value of the system;
[0109] A deep reinforcement learning model is integrated to dynamically adjust the sliding mode control parameters through a policy network. The sliding mode control parameters include switching gain, boundary layer thickness, and sliding surface weight coefficient.
[0110] It should be noted that the sliding surface It is a key component in trajectory tracking control and is composed of the error integral term and the adaptive disturbance compensation term. In the present invention, the sliding surface design takes into account the long-term accumulation of trajectory tracking error and the influence of the total disturbance of the system, thereby achieving compensation for external disturbances and effective suppression of tracking error. Its mathematical expression is:
[0111] ;
[0112] The trajectory tracking error integral term is:
[0113] ;
[0114] The adaptive compensation term is:
[0115] ;
[0116] in, is the trajectory tracking error, which represents the deviation between the target state and the actual state; is the time derivative of the error, which represents the rate of change of the error; is the current moment (the "outer" time); is the integration variable (the “inner” time); and are the integral weight coefficients, which are dynamically adjusted through deep reinforcement learning to balance the error convergence speed and anti-interference performance; is the disturbance compensation gain, which is linearly related to the towed load mass; is the disturbance sensitivity coefficient, which controls the saturation interval of the hyperbolic tangent function. Its typical value is ; is the total disturbance value of the system, which is output by the multi-level linear extended state observer (MLESO).
[0117] Furthermore, in order to achieve accurate trajectory tracking and effective suppression of anti-disturbance, the system designed a sliding mode control law , which includes equivalent control terms and switch controls Specific control law The formula is as follows:
[0118] ;
[0119] Equivalent Control Items It is used to compensate for the nonlinear problems caused by the nominal dynamics and disturbances of the system. The calculation formula is:
[0120] ;
[0121] in, is the state vector of the system; is the acceleration of the target trajectory; is the control input gain matrix, is the driving wheel radius; are the error feedback gains, designed by pole placement method; is the nominal dynamic model of the system, established by the Lagrangian method, as shown below:
[0122] ;
[0123] in, is the mass matrix, is the Coriolis force matrix, is the gravity vector, is the position vector of the joint or coordinate of the main robot or tow link; is the inverse of the mass matrix, used to map moments or torques to accelerations; Describes the inertial force caused by the velocity coupling term, and the function independent variable is and ; Represents the effects of Coriolis and centrifugal forces on acceleration; is the gravity vector, describing the position of each joint gravity or potential energy gradient under the
[0124] Toggle Control It is used to suppress oscillation caused by errors and external disturbances in real time. The calculation formula is:
[0125] ;
[0126] in, is the boundary layer thickness, which varies with the speed Adaptive expansion, initial value ; It is a saturation function, which is used to limit the amplitude of the switching control term to avoid excessive excitation; is the dynamic switching gain, output by the deep reinforcement learning model (DRL), expressed as:
[0127] ;
[0128] The initial value is , is the gain decay time constant, is the activation function, is the weight matrix in the DRL policy network, is the feature extraction function.
[0129] In order to further improve the adaptive capability of the control system, a deep reinforcement learning (DRL) model is integrated into the system to dynamically adjust the sliding mode control parameters. Through the double-delayed deep deterministic policy gradient (TD3) algorithm, the DRL model can automatically adjust the key parameters in the sliding mode control according to the system state information. The specific control parameters to be adjusted include: switching gain , boundary layer thickness and sliding surface weight coefficient and .
[0130] It should also be noted that using a deep reinforcement learning (DRL) model to automatically adjust key sliding mode control parameters actually leverages the powerful decision-making and adaptability of deep reinforcement learning, taking control system state information as input and dynamically adjusting control parameters based on this information. Specifically, the DRL model interacts with the system environment to learn how to optimize the control strategy based on the system's current state (e.g., trajectory error, trailer sway, speed, and load mass).
[0131] In this paper, the DRL model uses the double-delayed deep deterministic policy gradient (TD3) algorithm, a reinforcement learning algorithm for continuous control tasks. The TD3 algorithm utilizes an actor-critic network structure, combined with delayed updates and a target network mechanism to stabilize the training process and optimize sliding mode control parameters (such as switching gain, boundary layer thickness, and integral weight).
[0132] The TD3 algorithm is a commonly used algorithm in the field of reinforcement learning. It is widely used in control systems, particularly for high-dimensional continuous control tasks. Compared to the traditional Deep Deterministic Policy Gradient (DDPG) algorithm, the TD3 algorithm offers significant advantages in stability and training efficiency. In particular, the TD3 algorithm can reduce policy instability by delaying updates and limiting policy noise in noisy environments.
[0133] Specifically, the TD3 algorithm uses an actor network to output adjustments to sliding mode control parameters (such as the attenuation rate of the switching gain and the proportional coefficient of the boundary layer thickness). These adjustments can respond to system dynamics in real time and improve the control system's adaptability. Furthermore, a critic network evaluates the strategy to further optimize the system's long-term control performance.
[0134] With the help of the DRL model, the control system no longer relies on static parameters. Instead, it can automatically adjust key control parameters based on the real-time operating environment and load conditions to achieve more accurate trajectory tracking and greater robustness. In this way, the system can adaptively respond to various dynamic changes, such as load changes, speed changes, and environmental disturbances, significantly improving the performance of the control system.
[0135] Furthermore, in the present invention, a disturbance compensation mechanism is designed to further enhance the robustness and anti-disturbance performance of the system. Specifically, the disturbance compensation mechanism reduces the total disturbance value of the system to Equivalent control terms embedded in the sliding mode control framework , to achieve real-time compensation for system disturbances.
[0136] The core of the disturbance compensation mechanism is to convert the total disturbance value of the system calculated by the multi-level linear extended state observer (MLESO) into Feedforward injection into the equivalent control term In order to compensate for the anti-motion.
[0137] Step S31 of the present invention proposes a sliding mode control framework and designs a highly adaptable sliding surface to achieve precise control of the mobile robot's trajectory tracking error while effectively compensating for external disturbances to the system. The core of this step is to combine traditional sliding mode control with deep reinforcement learning (DRL) to dynamically adjust control parameters, thereby enhancing the control system's adaptability and robustness.
[0138] The trajectory tracking error, trailer swing amplitude, and total system disturbance value are input into the deep reinforcement learning model to construct a multi-dimensional state space.
[0139] Based on the double-delayed deep deterministic policy gradient algorithm, the adjustment amount of the sliding mode control parameters is output, including the switching gain attenuation rate and the boundary layer thickness proportional coefficient;
[0140] Dynamically calculate the weight coefficient of the integral term in the sliding surface based on the current movement speed and the mass of the towed load;
[0141] The equivalent control items, switching control items and parameter adjustment results are integrated to generate a control model after dynamic parameter adjustment, and the convergence of the control model after dynamic parameter adjustment is verified by the Lyapunov stability criterion.
[0142] In the embodiment of the present invention, the trajectory tracking error , trailer swing amplitude and the total disturbance value Input the deep reinforcement learning (DRL) model to construct a multi-dimensional state space. Specifically, the multi-dimensional state space It consists of the following dynamic features:
[0143] ;
[0144] in, is the trajectory tracking error, which represents the deviation between the target and the actual state; is the rate of change of the error, reflecting the speed at which the error changes; is the trailer swing angle, which indicates the swing amplitude of the trailer part; It is the total disturbance value of the system output by the multi-level linear extended state observer (MLESO), which covers the disturbances of the system itself, the tow chain and the external environment; is the current speed of motion, which affects the dynamic characteristics of the system; It is the mass of the trailer load, which affects the dynamic response of the trailer system.
[0145] By incorporating this critical information into the DRL model, the model can fully understand the system's dynamic state, providing more accurate feedback for adjusting sliding mode control. This multidimensional state space construction enhances the system's adaptability in complex environments and lays the foundation for subsequent control parameter optimization.
[0146] Furthermore, in the deep reinforcement learning (DRL) model, the double-delayed deep deterministic policy gradient (TD3) algorithm is used to dynamically adjust the sliding mode control parameters based on the Actor-Critic network structure. Specifically, the Actor network dynamically adjusts the sliding mode control parameters by The information extracted from the output is used to adjust the sliding mode control parameters. :
[0147] ;
[0148] in, It is the adjustment amount of the switching gain, which affects the smooth transition of the control signal; is the adjustment amount of the boundary layer thickness, which controls the degree of error smoothing; is the adjustment amount of the error integral weight, which balances the error convergence speed and anti-interference performance.
[0149] Through the critic network, the TD3 algorithm evaluates the current policy and, based on the optimization of the value function, further adjusts parameters to minimize long-term cumulative error. This update process utilizes a delayed strategy to ensure convergence of policy optimization by reducing overfitting and improving stability.
[0150] Furthermore, the weight coefficient of the integral term in the sliding surface is dynamically calculated according to the current motion speed. and trailer load mass The present invention dynamically calculates the integral weight coefficient in the sliding surface and These two coefficients determine the weight of the trajectory tracking error and its rate of change, which in turn affects the response speed and robustness of the sliding surface. and The dynamic adjustment mechanism is:
[0151] ;
[0152] in, Current time The moment The integral weight coefficient, is the integral weight value of the next time step, which is the result of dynamic update; Based on the speed of movement and load mass The calculated mapping function is used to dynamically adjust the weight of the integral term. This design allows the control strategy to be adjusted according to the actual operating conditions of the system, ensuring the system has optimal response performance under different speed and load conditions.
[0153] Finally, the system adjusts the parameters of deep reinforcement learning (including ) into the existing calculation to generate a control model after dynamic parameter adjustment , .
[0154] in, is the equivalent control item after dynamic parameter adjustment, It is the switching control item after dynamic parameter adjustment.
[0155] The control model after dynamic parameter adjustment is verified using the Lyapunov stability criterion to ensure that the system can maintain closed-loop stability under all working conditions. That is, under the influence of disturbances and errors, the system can still ensure the convergence of trajectory tracking errors.
[0156] It should be noted that by fusing control terms and optimized parameters, the dynamically tuned control model generated can adapt to various dynamic changes, improving trajectory tracking accuracy. The application of the Lyapunov criterion ensures the stability of the dynamically tuned control model, further enhancing the reliability and safety of the system.
[0157] S4: Calculate the main robot control signal in real time according to the control model after dynamic parameter adjustment, and perform actuator dynamic response compensation on the main robot control signal.
[0158] Actuator dynamic response compensation includes motor torque limitation and steering angle constraint.
[0159] It should be noted that this compensation ensures that the control signal of the main robot can be reasonably and safely executed in the actuator to avoid system overload or instability.
[0160] The motor torque limiter is used to limit the maximum output torque of the motor to prevent damage or performance degradation caused by loads exceeding the motor's tolerance.
[0161] The steering angle constraint is used to limit the steering angle to prevent excessive steering angles from causing unstable robot motion, especially at high speeds or in complex terrain.
[0162] After compensation, the final control signal of the main robot is converted into an actual executable drive instruction and transmitted to the main robot through the actuator, thereby achieving precise trajectory tracking control.
[0163] S5: monitoring the trajectory tracking error and the trailer swing amplitude error after the execution of the main robot control signal, adjusting the control strategy according to the error results, and feeding back to the trajectory tracking control model for adaptive optimization.
[0164] Real-time collection of actual trajectory data and trailer swing angle data after the main machine control signal is executed, and calculation of the trajectory tracking error and trailer swing amplitude error at the current moment;
[0165] Set dynamic error thresholds. When the trajectory tracking error exceeds the first threshold or the trailer swing amplitude error exceeds the second threshold, the online parameter re-optimization process of the deep reinforcement learning model is triggered to update the sliding mode control parameters.
[0166] If the error continues to exceed the limit and does not converge after optimization, it switches to the emergency control mode and uses the preset conservative sliding mode parameters and fixed boundary layer thickness;
[0167] The updated control parameters or emergency control instructions are fed back to the trajectory tracking control model to form a closed-loop adaptive optimization.
[0168] In the embodiment of the present invention, the system calculates the trajectory tracking error and the trailer swing amplitude error at the current moment by collecting the actual trajectory data and trailer swing angle data after the main machine control signal is executed in real time.
[0169] The trajectory tracking error reflects the deviation between the robot's current position and the target trajectory. This error indicates the current trajectory accuracy of the system.
[0170] Trailer swing amplitude error represents the deviation between the trailer system's actual swing angle and the target swing angle. Trailer swing amplitude error affects the stability of the robot and trailer system, especially in complex road conditions.
[0171] The system sets dynamic error thresholds based on real-time calculated tracking error and trailer swing amplitude error. When the error exceeds the first and second thresholds, the system triggers the online parameter re-optimization process of the deep reinforcement learning (DRL) model, dynamically adjusting the control strategy.
[0172] The first threshold is the maximum allowable value of the trajectory tracking error. When this threshold is exceeded, it means that the system deviates too far from the trajectory and the control parameters need to be optimized.
[0173] The second threshold is the maximum allowable value of the trailer swing amplitude. When this threshold is exceeded, it means that the swing amplitude of the trailer system is too large, affecting the system stability and needs to be adjusted in time.
[0174] When the error exceeds a threshold, the DRL model uses the policy network to optimize the sliding mode control parameters online based on the current trajectory tracking error and trailer swing amplitude error. Specifically, the system dynamically adjusts sliding mode control parameters (such as switching gain and boundary layer thickness) to improve the current control effect.
[0175] This optimization process gradually adjusts the control parameters based on the feedback mechanism of the deep reinforcement learning model, enabling the system to adapt to different working environments and avoid further expansion of errors.
[0176] If the error continues to exceed the threshold during the optimization process of the deep reinforcement learning model and the optimization fails to converge, the system will switch to emergency control mode. In emergency mode, the system will use preset conservative sliding mode parameters and fixed boundary layer thickness to ensure that the system can still maintain a certain degree of stability even if the error exceeds the limit.
[0177] The emergency control mode effectively reduces the risk of the system and avoids further instability caused by excessive adjustment of the system.
[0178] In emergency control mode, or based on feedback from deep reinforcement learning model optimization, control commands are fed back to the trajectory tracking control model, forming a closed-loop adaptive optimization. This continuously optimizes control parameters, ensuring efficient and precise operation of the robot and trailer system in dynamic environments. Through continuous adaptive optimization, the system can adjust control strategies in real time as the operating environment changes, maximizing trajectory tracking accuracy and trailer stability.
[0179] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
[0180] Example 2 is the second embodiment of the present invention, which is different from the previous embodiment in that:
[0181] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art or the portion of the current technical solution, can be embodied in the form of a software product. The current computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0182] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0183] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.
[0184] Example 3 is an embodiment of the present invention, which provides a mobile robot trajectory tracking control system, including a data acquisition module, a disturbance calculation module, a trajectory tracking control model construction module, a real-time control signal generation module and an adaptive optimization feedback module.
[0185] Data acquisition module: collects the original motion state data of the robot body and the towing part through multiple types of sensors, pre-processes the original motion state data, and generates pre-processed motion state data;
[0186] Disturbance calculation module: calculates the total disturbance value of the system based on the pre-processed motion state data;
[0187] Trajectory tracking control model construction module: constructs a trajectory tracking control model, inputs the pre-processed motion state data and the total disturbance value of the system, optimizes the sliding mode control parameters, and generates a control model after dynamic parameter adjustment;
[0188] Real-time control signal generation module: calculates the control signal of the main robot in real time according to the control model after dynamic parameter adjustment, and performs actuator dynamic response compensation on the control signal of the main robot;
[0189] Adaptive optimization feedback module: monitors the trajectory tracking error and the trailer swing amplitude error after the execution of the main robot control signal, adjusts the control strategy according to the error results, and feeds back to the trajectory tracking control model for adaptive optimization.
Claims
1. A mobile robot trajectory tracking control method, characterized in that: include: Collecting original motion state data of the robot body and the towing part through multiple types of sensors, preprocessing the original motion state data to generate preprocessed motion state data; Calculating a total disturbance value of the system based on the preprocessed motion state data; Constructing a trajectory tracking control model, inputting the pre-processed motion state data and the total disturbance value of the system, optimizing the sliding mode control parameters, and generating a control model after dynamic parameter adjustment; According to the control model after dynamic parameter adjustment, the control signal of the main robot is calculated in real time, and the dynamic response compensation of the actuator is performed on the control signal of the main robot; Monitor the trajectory tracking error and the trailer swing amplitude error after the main robot control signal is executed, adjust the control strategy according to the error results, and feed them back to the trajectory tracking control model for adaptive optimization; The preprocessing of the raw motion state data includes performing multi-source fusion filtering on the motion state data of the robot body collected by the body sensor group to extract the state vector of the robot body, wherein the state vector of the robot body includes position, velocity, heading angle and normalized steering motor current; The tow chain tension and swing angle data collected by the tow sensor group are denoised, and the tow chain coupling disturbance characteristic matrix is constructed based on the six-dimensional torque data of the hitch point. Perform spatiotemporal registration of images, laser point clouds, and radar data collected by the environmental perception sensor group, and generate environment-robot interaction feature tensors through spatial transformation; The state vector of the robot body, the tow chain coupling disturbance characteristic matrix and the environment-robot interaction characteristic tensor are uniformly normalized, and preprocessed motion state data is output.
2. The mobile robot trajectory tracking control method according to claim 1, wherein: The collecting of original motion state data of the robot body and the trailer part by using multiple types of sensors includes collecting the motion state data of the robot body by using a body sensor group, wherein the body sensor group includes an inertial measurement unit, a wheel speed encoder, and a steering motor current sensor; The trailer sensor group collects the dynamic state data of the trailer part, and the trailer sensor group includes a trailer chain tension sensor, a trailer swing angle gyroscope, and a six-dimensional torque sensor of the hitch point; External disturbance feature data is collected through an environment perception sensor group, which includes a binocular vision camera, a solid-state laser radar and a millimeter wave radar array.
3. The mobile robot trajectory tracking control method according to claim 2, wherein: The calculating of the total disturbance value of the system based on the pre-processed motion state data includes constructing a multi-level linear extended state observer, wherein the multi-level linear extended state observer includes a first-level observer, a second-level observer, and a third-level observer; The pre-processed state vector of the robot body is input into the first-level observer, and the body dynamic disturbance value is calculated in real time through the linear extended state observation equation. The body dynamic disturbance value includes the motor nonlinear friction and wheel-ground contact slip disturbance; The pre-processed tow chain coupling disturbance characteristic matrix is input into the second-level observer, and the coupling disturbance value caused by the trailer inertia force and the flexible deformation of the coupling point is calculated in real time through the linear expansion state observation equation; The preprocessed environment-robot interaction feature tensor is input into the third-level observer, which extracts key environmental disturbance features through the attention mechanism and calculates the environmental disturbance value. The body dynamics disturbance value, coupling disturbance value and environmental disturbance value are integrated to generate the total disturbance value of the system.
4. The mobile robot trajectory tracking control method according to claim 3, wherein: The constructing of the trajectory tracking control model includes establishing a sliding mode control framework and designing a sliding mode surface, wherein the sliding mode surface is composed of a trajectory tracking error integral term and an adaptive compensation term based on a total disturbance value of the system; Constructing a sliding mode control law, wherein the sliding mode control law includes an equivalent control term and a switching control term, wherein the equivalent control term is generated based on a nominal dynamic model of the system and a total disturbance value of the system; A deep reinforcement learning model is integrated to dynamically adjust the sliding mode control parameters through a policy network. The sliding mode control parameters include switching gain, boundary layer thickness, and sliding surface weight coefficient.
5. The mobile robot trajectory tracking control method according to claim 4, characterized in that: Generating the control model after dynamic parameter adjustment includes inputting the trajectory tracking error, the trailer swing amplitude, and the total system disturbance value into the deep reinforcement learning model to construct a multi-dimensional state space; Based on the double-delayed deep deterministic policy gradient algorithm, the adjustment amount of the sliding mode control parameters is output, including the switching gain attenuation rate and the boundary layer thickness proportional coefficient; Dynamically calculate the weight coefficient of the integral term in the sliding surface based on the current movement speed and the mass of the towed load; The equivalent control items, switching control items and parameter adjustment results are integrated to generate a control model after dynamic parameter adjustment, and the convergence of the control model after dynamic parameter adjustment is verified by the Lyapunov stability criterion.
6. The mobile robot trajectory tracking control method according to claim 5, characterized in that: The actuator dynamic response compensation includes motor torque limitation and steering angle constraint.
7. The mobile robot trajectory tracking control method according to claim 6, characterized in that: The adaptive optimization includes collecting the actual trajectory data and trailer swing angle data after the main robot control signal is executed in real time, and calculating the trajectory tracking error and trailer swing amplitude error at the current moment; Set dynamic error thresholds. When the trajectory tracking error exceeds the first threshold or the trailer swing amplitude error exceeds the second threshold, the online parameter re-optimization process of the deep reinforcement learning module is triggered to update the sliding mode control parameters. If the error continues to exceed the limit and does not converge after optimization, it switches to the emergency control mode and uses the preset conservative sliding mode parameters and fixed boundary layer thickness; The updated control parameters or emergency control instructions are fed back to the trajectory tracking control model to form a closed-loop adaptive optimization.
8. A mobile robot trajectory tracking control system, used to implement the mobile robot trajectory tracking control method according to any one of claims 1 to 7, characterized in that: include: Data acquisition module: collects the original motion state data of the robot body and the towing part through multiple types of sensors, pre-processes the original motion state data, and generates pre-processed motion state data; Disturbance calculation module: calculates the total disturbance value of the system based on the pre-processed motion state data; Trajectory tracking control model construction module: constructs a trajectory tracking control model, inputs the pre-processed motion state data and the total disturbance value of the system, optimizes the sliding mode control parameters, and generates a control model after dynamic parameter adjustment; Real-time control signal generation module: calculates the control signal of the main robot in real time according to the control model after dynamic parameter adjustment, and performs actuator dynamic response compensation on the control signal of the main robot; Adaptive optimization feedback module: monitors the trajectory tracking error and the trailer swing amplitude error after the execution of the main robot control signal, adjusts the control strategy according to the error results, and feeds back to the trajectory tracking control model for adaptive optimization.
Citation Information
Patent Citations
Traction type trailer trajectory tracking method based on robust H infinite control
CN111352442A
Rope traction parallel robot based on double-rope model and control method and device thereof
CN117656036A