Mobile robot trajectory tracking control method and system
Through the sliding mode control method of multi-sensor data fusion and deep reinforcement learning model online parameter adjustment, the problem of poor parameters fixed and real-time adaptability in mobile robot trajectory tracking control is solved, and high-precision and robust trajectory tracking control is achieved, which improves the dynamic response and stability of the system.
Patent Information
- Application Number
- CN202510542243.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing mobile robot trajectory tracking control methods have problems such as fixed parameters, poor real-time adaptability, and insufficient handling of complex disturbances of the system, making it difficult to achieve high-precision and robust trajectory tracking control in dynamic and complex environments.
Fine state information is obtained through multi-sensor data fusion, the system total disturbance value is calculated using a multi-stage expansion state observer, and combined with the sliding mode control method based on the deep reinforcement learning model to achieve coordinated compensation of trajectory tracking error and drag swing.
It significantly improves the dynamic response speed and control robustness of the system, reduces high-frequency vibration, and improves overall control accuracy and stability.
Smart Images

Figure CN120065757A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile robot trajectory tracking control, and specifically provides a mobile robot trajectory tracking control method and system. Background Art
[0002] In recent years, with the rapid development of industrial automation and intelligent manufacturing, mobile robot technology has been widely applied in fields such as logistics, warehousing, and security. Based on cutting-edge technologies such as multi-sensor fusion, advanced control algorithms, and deep learning, the environmental perception, path planning, and motion control levels of robots have been continuously improved. In particular, control schemes that combine sliding mode control, model predictive control, and deep reinforcement learning have become research hotspots, and certain achievements have been made both theoretically and experimentally, providing strong support for achieving high-precision trajectory tracking.
[0003] However, there are still obvious deficiencies in the existing technologies. First, traditional mobile robot trajectory tracking methods mostly rely on sliding mode or PID control with fixed parameters, and it is difficult to adapt to dynamic environmental disturbances and system nonlinearities, resulting in insufficient tracking accuracy in complex scenarios. Second, although some studies have introduced deep learning algorithms, they usually only target a single compensation strategy and fail to fully utilize multi-type sensor data to comprehensively calculate disturbances, making it difficult to effectively suppress the combined disturbances caused by motor friction, towing coupling, and environmental interference. In addition, existing methods have not deeply modeled the dynamic coupling characteristics of the towing system, and cannot simultaneously consider the dynamic effects caused by load changes and flexible deformation of the towing chain, resulting in poor system stability and robustness. Finally, many control methods are insufficient in the design of real-time control signal calculation and actuator response compensation, resulting in slow error convergence speed and prominent high-frequency chattering phenomena. These problems limit the ability of existing solutions to achieve high-precision and stable trajectory tracking in complex dynamic environments, and the present invention aims at the above deficiencies and realizes the coordinated compensation of trajectory tracking errors and towing swing through multi-sensor data fusion, disturbance calculation, multi-level sliding mode control, and online tuning of the deep reinforcement learning model, thereby greatly improving the dynamic response performance and robustness of the system. Summary of the Invention
[0004] In view of the above problems, the present invention is proposed.
[0005] Therefore, the technical problems solved by the present invention are: the existing mobile robot trajectory tracking control methods have problems such as fixed parameters, poor real-time adaptability, and insufficient handling of complex system disturbances (such as body friction, towing chain coupling, and external environmental interference), and the problem of how to achieve high-precision and robust trajectory tracking control in a dynamic and complex environment.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, an embodiment of the present invention provides a method for trajectory tracking control of a mobile robot, including collecting original motion state data of the robot body and the towed part through multiple types of sensors, preprocessing the original motion state data to generate preprocessed motion state data; Calculating the total system disturbance value based on the preprocessed motion state data; Constructing a trajectory tracking control model, inputting the preprocessed motion state data and the total system disturbance value, optimizing the sliding mode control parameters, and generating a control model with dynamically adjusted parameters; According to the control model with dynamically adjusted parameters, calculating the control signal of the body robot in real time, and performing actuator dynamic response compensation on the control signal of the body robot; Monitoring the trajectory tracking error and the towed swing amplitude error after the execution of the control signal of the body robot, adjusting the control strategy according to the error result, and feeding it back to the trajectory tracking control model for adaptive optimization.
[0007] As a preferred scheme of the method for trajectory tracking control of the mobile robot according to the present invention, wherein: the collecting of the original motion state data of the robot body and the towed part through multiple types of sensors includes collecting the motion state data of the robot body through the body sensor group, and the body sensor group includes an inertial measurement unit, a wheel speed encoder, and a steering motor current sensor; Collecting the dynamic state data of the towed part through the towed sensor group, and the towed sensor group includes a tow chain tension sensor, a trailer swing angle gyroscope, and a hitch six-dimensional torque sensor; Collecting external disturbance feature data through the environmental perception sensor group, and the environmental perception sensor group includes a binocular vision camera, a solid-state lidar, and a millimeter-wave radar array.
[0008] As a preferred scheme of the method for trajectory tracking control of the mobile robot according to the present invention, wherein: the preprocessing of the original motion state data includes performing multi-source fusion filtering on the motion state data of the robot body collected by the body sensor group, and extracting the state vector of the robot body, and the state vector of the robot body includes position, speed, heading angle, and steering motor current; Performing denoising processing on the tow chain tension data and swing angle data collected by the towed sensor group, and constructing a tow chain coupling disturbance feature matrix based on the hitch six-dimensional torque data; Performing spatio-temporal registration on the images, laser point clouds, and radar data collected by the environmental perception sensor group, and generating an environment-robot interaction feature tensor through spatial transformation; Unifying the normalization processing of the state vector of the robot body, the tow chain coupling disturbance feature matrix, and the environment-robot interaction feature tensor, and outputting the preprocessed motion state data.
[0009] As a preferred solution of the mobile robot trajectory tracking control method described in the present invention, wherein: calculating the total system disturbance value based on the pre-processed motion state data includes constructing a multi-level linear extended state observer, and the multi-level linear extended state observer includes a first-level observer, a second-level observer, and a third-level observer; Input the state vector of the robot body after pre-processing into the first-level observer, and calculate the dynamic disturbance value of the body in real time through the linear extended state observation equation. The dynamic disturbance value of the body includes motor non-linear friction and wheel-ground contact slip disturbance; Input the pre-processed coupling disturbance characteristic matrix of the towing chain into the second-level observer, and calculate the coupling disturbance value caused by the trailer inertia force and the flexible deformation of the hitch point in real time through the linear extended state observation equation; Input the pre-processed environment-robot interaction feature tensor into the third-level observer, extract key environmental disturbance features through the attention mechanism, and calculate the environmental disturbance value; Fuse the dynamic disturbance value of the body, the coupling disturbance value, and the environmental disturbance value to generate the total system disturbance value.
[0010] As a preferred solution of the mobile robot trajectory tracking control method described in the present invention, wherein: constructing the trajectory tracking control model includes establishing a sliding mode control framework and designing a sliding mode surface, and the sliding mode surface is jointly composed of a trajectory tracking error integral term and an adaptive compensation term based on the total system disturbance value; Construct a sliding mode control law, and the sliding mode control law includes an equivalent control term and a switching control term, wherein the equivalent control term is generated based on the system nominal dynamic model and the total system disturbance value; Integrate a deep reinforcement learning model, and dynamically adjust the sliding mode control parameters through a policy network. The sliding mode control parameters include a switching gain, a boundary layer thickness, and a sliding mode surface weight coefficient; Design a disturbance compensation mechanism and embed the total system disturbance value into the equivalent control term.
[0011] As a preferred solution of the mobile robot trajectory tracking control method described in the present invention, wherein: generating the control model with dynamically adjusted parameters includes inputting the trajectory tracking error, the towing swing amplitude, and the total system disturbance value into the deep reinforcement learning model to construct a multi-dimensional state space; Based on the double-delayed deep deterministic policy gradient algorithm, output the adjustment amount of the sliding mode control parameters, including the switching gain attenuation rate and the boundary layer thickness proportional coefficient; Dynamically calculate the weight coefficient of the integral term in the sliding mode surface according to the current motion speed and the towing load mass; Integrate the equivalent control item, switching control item, and parameter adjustment result to generate a control model after dynamic parameter adjustment, and verify the convergence of the control model after dynamic parameter adjustment through the Lyapunov stability criterion.
[0012] As a preferred solution of the mobile robot trajectory tracking control method described in the present invention, wherein: the actuator dynamic response compensation includes motor torque limiting and steering angle constraint.
[0013] As a preferred solution of the mobile robot trajectory tracking control method described in the present invention, wherein: the adaptive optimization includes collecting the actual trajectory data and towed swing angle data after the execution of the control signal of the main body robot in real time, and calculating the trajectory tracking error and towed swing amplitude error at the current moment; Set a dynamic error threshold. When the trajectory tracking error exceeds the first threshold or the towed swing amplitude error exceeds the second threshold, trigger the online parameter re-optimization process of the deep reinforcement learning module to update the sliding mode control parameters; If the error continues to exceed the limit and does not converge after optimization, switch to the emergency control mode and adopt the preset conservative sliding mode parameters and fixed boundary layer thickness; Feed back the updated control parameters or emergency control instructions to the trajectory tracking control model to form a closed-loop adaptive optimization.
[0014] In a second aspect, an embodiment of the present invention provides a mobile robot trajectory tracking control system, including: Data acquisition module: Collect the original motion state data of the robot main body and the towed part through multiple types of sensors, preprocess the original motion state data, and generate preprocessed motion state data; Disturbance calculation module: Calculate the total system disturbance value based on the preprocessed motion state data; Trajectory tracking control model construction module: Construct a trajectory tracking control model, input the preprocessed motion state data and the total system disturbance value, optimize the sliding mode control parameters, and generate a control model after dynamic parameter adjustment; Real-time control signal generation module: According to the control model after dynamic parameter adjustment, calculate the control signal of the main body robot in real time, and perform actuator dynamic response compensation on the control signal of the main body robot; Adaptive optimization feedback module: Monitor the trajectory tracking error and towed swing amplitude error after the execution of the control signal of the main body robot, adjust the control strategy according to the error result, and feedback it to the trajectory tracking control model for adaptive optimization.
[0015] Advantages of the present invention: By fusing multi-sensor data to obtain fine state information, calculating the total system disturbance value using a multi-level extended state observer, and combining a sliding mode control method with online parameter tuning based on a deep reinforcement learning model, the present invention realizes the collaborative compensation for trajectory tracking error and trailer swing, thereby significantly improving the dynamic response speed and control robustness of the system, reducing high-frequency chattering, and improving the overall control accuracy and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where: Figure 1 FIG. is the overall flowchart of a mobile robot trajectory tracking control method provided by the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0018] Example 1, referring to Figure 1 , which is an embodiment of the present invention, provides a mobile robot trajectory tracking control method, including: S1: Collect the original motion state data of the robot body and the trailer part through multi-type sensors, and preprocess the original motion state data to generate preprocessed motion state data.
[0019] S11: Collect the motion state data of the robot body through the body sensor group, and the body sensor group includes an inertial measurement unit, a wheel speed encoder, and a steering motor current sensor; Collect the dynamic state data of the trailer part through the trailer sensor group, and the trailer sensor group includes a trailer chain tension sensor, a trailer swing angle gyroscope, and a hitch point six-dimensional torque sensor; Collect the external disturbance feature data through the environmental perception sensor group, and the environmental perception sensor group includes a binocular vision camera, a solid-state lidar, and a millimeter-wave radar array.
[0020] In the embodiment of the present invention, the "body sensor group" used includes an inertial measurement unit (IMU), a wheel speed encoder, and a steering motor current sensor. Among them, the IMU (Inertial Measurement Unit) measures the acceleration and angular velocity of the robot in real time, providing high-frequency, high-precision motion information; the wheel speed encoder accurately reflects the wheel speed through pulse output, thereby calculating the robot's travel distance and speed; the steering motor current sensor reflects the change in motor load and indirectly detects the quality of steering execution. The output parameters of these three types of equipment (IMU, encoder, current sensor) constitute the robot body state vector after fusion, providing data support for subsequent error calculation and controller design.
[0021] In the towing part, the present invention adopts a towing sensor group, whose core sensors include a towing chain tension sensor, a trailer swing angle gyroscope and a six-dimensional torque sensor at the attachment point. The towing chain tension sensor can capture the load fluctuation in the transmission chain and reflect the stress condition of the towing system; the trailer swing angle gyroscope accurately measures the lateral swing angle of the trailer and directly detects dynamic instability; the six-dimensional torque sensor at the attachment point can decompose the forces and torques in all directions and provide comprehensive mechanical coupling information. The towing state parameters composed of these parameters not only complete the overall motion state of the robot, but also make the system more sensitive and responsive to complex disturbances caused by flexible connections.
[0022] The environmental perception sensor group consists of binocular vision cameras, solid-state laser radars and millimeter-wave radar arrays, each with its own unique advantages: binocular vision cameras obtain stereoscopic images of the scene to achieve depth calculation of obstacles and environment; solid-state laser radars provide high-resolution two-dimensional or three-dimensional point cloud data for real-time modeling of the environment; millimeter-wave radar arrays can still work stably under adverse weather conditions, making up for the shortcomings of vision and laser in rainy and foggy environments. After multi-sensor fusion, the output environmental disturbance feature data enables the system to perceive dynamic changes in the outside world in real time, thereby achieving intelligent obstacle avoidance and interference compensation.
[0023] S12: performing multi-source fusion filtering on the motion state data of the robot body collected by the body sensor group to extract the state vector of the robot body, where the state vector of the robot body includes position, speed, heading angle and normalized steering motor current; The towing chain tension data and swing angle data collected by the towing sensor group are denoised, and the towing chain coupling disturbance characteristic matrix is constructed based on the six-dimensional torque data of the attachment point; Perform spatiotemporal registration of images, laser point clouds, and radar data collected by the environment perception sensor group, and generate environment-robot interaction feature tensors through spatial transformation; Normalize the state vector of the robot body, the coupling perturbation feature matrix of the trailing cable, and the environment-robot interaction feature tensor uniformly, and output the preprocessed motion state data; the preprocessed motion state data includes the preprocessed state vector of the robot body, the preprocessed coupling perturbation feature matrix of the trailing cable, and the preprocessed environment-robot interaction feature tensor.
[0024] In the embodiment of the present invention, first, the data collected by the body sensor group includes an inertial measurement unit (IMU), a wheel speed encoder, and a steering motor current sensor. These data respectively reflect the acceleration, angular velocity, driving speed, driving distance, and motor load change of the robot in digital form. Using a multi-source fusion filtering algorithm (such as Kalman filtering or extended Kalman filtering) for these data can effectively suppress noise and interference, and extract a more real and accurate state of the robot body. The extracted state vector of the robot body includes position, speed, heading angle, and the normalized steering motor current. These parameters can be directly quantified from the sensor signals. After normalization, the data of each channel are in the same numerical order of magnitude, improving the stability and robustness of subsequent data fusion and controller design. This step obtains a stable and low-noise state vector through fusion filtering, providing a solid data basis for subsequent system perturbation calculation and control law calculation, so as to achieve higher-precision motion control in a harsh dynamic environment.
[0025] Secondly, in terms of the trailing sensor group, by collecting the trailing cable tension data and the trailer swing angle data, the force and dynamic attitude changes of the trailing part can be reflected; at the same time, the multi-dimensional torque information provided by the hitch six-dimensional torque sensor can reveal the flexible deformation and inertial effect existing in the mechanical coupling of the trailing cable. For this, first use time-domain or frequency-domain denoising (such as wavelet denoising and other methods) for the tension and swing angle data to remove high-frequency interference, and then use the torque data to construct a coupling perturbation feature matrix of the trailing cable. This matrix reflects the dynamic coupling relationship between the sensor data. The key to constructing this perturbation feature matrix is to uniformly map the scattered sensor signals to a high-dimensional feature space, so that the dynamic uncertainty of the trailing cable can be quantitatively expressed in matrix form. The output of this matrix helps to accurately capture the non-linear coupling effect of the trailing part through a multi-stage extended state observer in subsequent perturbation calculations, so as to better achieve trailing dynamic compensation and overcome the problem that traditional methods only rely on a single signal.
[0026] Third, the images, lidar point clouds, and radar data collected by the environmental perception sensor group each have their own advantages. For example, image data has rich color and texture information, while lidar and millimeter-wave radars provide accurate information on distance and shape. By performing spatio-temporal registration on these data and integrating the three-dimensional point cloud with the image data using a spatial transformation (such as the Lie group SE(3) transformation), a high-dimensional environment-robot interaction feature tensor is generated. This tensor not only quantifies the spatial positions of obstacles and the dynamic environment but also captures the time-varying characteristics of interference factors in the environment, providing a key reference for the perturbation calculation module. Normalization ensures that the preprocessed motion state data is consistent in scale, facilitating subsequent data parsing by the controller and the application of robust control algorithms.
[0027] S2: Calculate the total system perturbation value based on the preprocessed motion state data.
[0028] Construct a multi-level linear extended state observer, which includes a first-level observer, a second-level observer, and a third-level observer; Input the state vector of the preprocessed robot body into the first-level observer, and calculate the dynamic perturbation value of the body in real time through the linear extended state observation equation. The dynamic perturbation value of the body includes motor nonlinear friction and wheel-ground contact slip perturbation; Input the preprocessed coupling perturbation feature matrix of the towing chain into the second-level observer, and calculate the coupling perturbation value caused by the trailer inertia force and the flexible deformation of the hitch point in real time through the linear extended state observation equation; Input the preprocessed environment-robot interaction feature tensor into the third-level observer, extract key environmental perturbation features through the attention mechanism, and calculate the environmental perturbation value; Fuse the dynamic perturbation value of the body, the coupling perturbation value, and the environmental perturbation value to generate the total system perturbation value.
[0029] In step S2, taking the preprocessed motion state data as the input, use a multi-level linear extended state observer (MLESO) to calculate the total system perturbation value in real time. First, the first-level observer takes the state vector of the preprocessed robot body as the input, and the state vector of the preprocessed robot body is: ; where represents the position, represents the velocity, represents the heading angle, represents the normalized steering motor current, represents the transpose operation.
[0030] Define the extended state vector as: ; Among them, is the value of the body dynamics disturbance at the previous moment saved inside the observer; Using the linear extended state observation equation, calculate the time derivative of the extended state vector , that is, the instantaneous change rate of the system state and the disturbance: ; Among them, is the state transition matrix, with the dimension matching that of , obtained by linearizing the robot body dynamics model, and describes the "natural" evolution of the extended state without control input and without measurement correction; is the input matrix, which maps the control input vector to the extended state space and is used to reflect the influence of the control action on the state and disturbance estimation; is the control input vector, such as the motor torque command or the drive current; is the observer gain matrix, which is used to feedback the measurement error into the extended state update to correct the model prediction deviation; is the actual measurement output vector, usually including the physical quantities directly available to the robot (such as ); is the output matrix, which maps the extended state vector to the measurement space.
[0031] It should be noted here that is the instantaneous change rate of the extended state vector at the current moment. In each sampling period, the observer will "accumulate" this change rate - continuously accumulate onto the current to generate the at the next moment.
[0032] This mechanism of "first calculating the change and then accumulating it to the state" is like calculating the velocity first and then updating the position in numerical integration, enabling to track the evolution of the system state and the disturbance in real time and adaptively.
[0033] Thus, the updated and latest value of the body dynamics disturbance is obtained: ; Among them, is the at the next moment (the updated ), is the mapping matrix.
[0034] This step effectively captures disturbance information such as motor non-linear friction and wheel-ground slip. Its beneficial effect is to provide an accurate compensation basis for subsequent towed and environmental disturbance calculations, solving the deficiency that traditional single sensors cannot independently distinguish disturbance components.
[0035] Secondly, the second-level observer takes the preprocessed towed-chain coupling disturbance feature matrix collected by the towed sensor group as input, and decomposes the towed-chain dynamics based on the virtual pendulum model. An extended state vector is established for each segment : ; where is the swing angle of the th segment of the towed chain, is the change rate of the swing angle of the th segment of the towed chain, is the estimated value of the coupling disturbance of the th segment at the previous moment; The derivative of the extended state vector is obtained through the linear extended state observation equation: ; where is the state transition matrix obtained after linearization of the th segment of the virtual pendulum; is the observer correction gain matrix, used to feedback the measurement error into the state update; is the measurement vector of the th segment of the towed-chain coupling disturbance feature matrix, including chain swing dynamics characteristics (such as tension, torque, etc.); is the output matrix, mapping the extended state to the measurement space; The disturbance value of each segment is obtained by the mapping matrix , that is ; where is the disturbance mapping matrix, used to extract the coupling disturbance component from ; is at the next moment (updated ); After superimposing all segment disturbances, the total disturbance value of the towed system is obtained: ; where is the total number of segments of the towed chain.
[0036] This design enables the quantitative description of the non - linear coupling of trailer dynamics through model decomposition, significantly improving the calculation accuracy of trailer inertial forces and the disturbances of the hitch point flexible deformation, and solving the problem that previous methods only rely on single - signal compensation.
[0037] Thirdly, the third - level observer takes the pre - processed environment - robot interaction feature tensor generated by integrating the images, laser point clouds, and radar data collected by the environmental perception sensor group as the input and uses the attention mechanism for feature extraction. This process is calculated as: ; where is the weight assigned to the th environmental feature by the attention mechanism; is the attention weight matrix used for linear transformation of each feature vector; is the th column environmental feature vector; is the Softmax function; Calculate the environmental disturbance value : ; where is the environmental disturbance mapping matrix used to map each column feature to a scalar disturbance estimate value; is the prediction of the disturbance contribution to the th feature; is the finally obtained environmental disturbance value; represents the number of columns of the environmental - robot interaction feature tensor , that is, the total number of environmental features (the number of candidate feature vectors).
[0038] This method uses the softmax and attention mechanisms to prioritize the extraction of key disturbance information in the environment, solving the problems of large noise interference and data redundancy in traditional environmental disturbance calculations.
[0039] Calculate the total system disturbance value , and the formula is: ; where , , respectively represent the body dynamics disturbance (such as motor friction, wheel - ground slip), trailer chain coupling disturbance (such as inertial force, flexible deformation), and external environmental disturbance (such as cross - wind, ground undulation); , , are the disturbance weight coefficients, and these weight coefficients will be based on the current system's load mass, speed and the environmental sensitivity coefficient is dynamically adjusted to ensure accurate compensation for dynamic disturbances.
[0040] S3: Construct a trajectory tracking control model, input the preprocessed motion state data and the total system disturbance value, optimize the sliding mode control parameters, and generate a control model with dynamically adjusted parameters.
[0041] Establish a sliding mode control framework and design a sliding mode surface, which is jointly composed of an integral term of the trajectory tracking error and an adaptive compensation term based on the total system disturbance value; Construct a sliding mode control law, which includes an equivalent control term and a switching control term, where the equivalent control term is generated based on the system nominal dynamic model and the total system disturbance value; Integrate a deep reinforcement learning model to dynamically adjust the sliding mode control parameters through a policy network. The sliding mode control parameters include a switching gain, a boundary layer thickness, and a sliding mode surface weight coefficient.
[0042] It should be noted that the sliding mode surface is a key component in trajectory tracking control and is jointly composed of an error integral term and an adaptive disturbance compensation term. In the present invention, the design of the sliding mode surface takes into account the long-term accumulation of the trajectory tracking error and the influence of the total system disturbance, thereby realizing the compensation for external disturbances and the effective suppression of tracking errors. Its mathematical expression is: ; The integral term of the trajectory tracking error is: ; The adaptive compensation term is: ; Among them, is the trajectory tracking error, representing the deviation between the target state and the actual state; is the time derivative of the error, representing the rate of change of the error; is the current time ("outer layer" time); is the integration variable ("inner layer" time); and are integral weight coefficients, which are dynamically adjusted through deep reinforcement learning to balance the error convergence rate and disturbance rejection ability; is the disturbance compensation gain, which is linearly related to the mass of the towed load; is the disturbance sensitivity coefficient, which controls the saturation interval of the hyperbolic tangent function, and the typical value is ; is the total system disturbance value, which is output by a multi-stage linear extended state observer (MLESO).
[0043] Furthermore, to achieve precise trajectory tracking and effective suppression of disturbances, the system designs a sliding mode control law , which includes an equivalent control term and a switching control term . The specific control law is as follows: ; The equivalent control term is used to compensate for the nominal dynamics of the system and the nonlinear problems caused by disturbances. Its calculation formula is: ; where is the state vector of the system; is the acceleration of the target trajectory; is the control input gain matrix, is the driving wheel radius; are the error feedback gains respectively, designed by the pole placement method; is the nominal dynamics model of the system, established by the Lagrangian method, as shown below: ; where is the mass matrix, is the Coriolis force matrix, is the gravity vector, is the position vector of the joints or coordinates of the main robot or the towed link; is the inverse of the mass matrix, used to map the torque or moment to the acceleration; describes the inertial force caused by the velocity coupling term, and the function arguments are and ; represents the influence of the Coriolis and centrifugal forces on the acceleration; is the gravity vector, describing the gravity or potential energy gradient of each joint at the position .
[0044] The switching control term is used to suppress the oscillations caused by errors and external disturbances in real time. Its calculation formula is: ; where is the boundary layer thickness, which adaptively expands with the motion speed , and the initial value is ; is a saturation function, used to limit the amplitude of the switching control term and avoid excessive excitation; is the dynamic switching gain, output by the deep reinforcement learning model (DRL), expressed as: ; The initial value is , is the gain attenuation time constant, is the activation function, is the weight matrix in the DRL policy network, is the feature extraction function.
[0045] To further improve the adaptive ability of the control system, a deep reinforcement learning (DRL) model is integrated into the system to dynamically adjust the sliding mode control parameters. Through the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, the DRL model can automatically adjust the key parameters in the sliding mode control according to the state information of the system. The specific control parameters to be adjusted include: the switching gain , the boundary layer thickness , and the sliding mode surface weight coefficient and .
[0046] It should also be noted that automatically adjusting the key parameters in the sliding mode control through the deep reinforcement learning (DRL) model is actually using the powerful decision-making ability and adaptability of deep reinforcement learning, taking the state information of the control system as input, and dynamically adjusting the control parameters according to this information. Specifically, the DRL model learns how to optimize the control strategy according to the current state of the system (such as trajectory error, trailer swing amplitude, speed, load mass, etc.) through interaction with the system environment.
[0047] In the present invention, the DRL model adopts the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, which is a reinforcement learning algorithm for continuous control tasks. The TD3 algorithm stabilizes the training process by using the Actor-Critic network structure, combining delayed updates and target network mechanisms, and optimizes the sliding mode control parameters (such as switching gain, boundary layer thickness, integral weight, etc.).
[0048] The TD3 algorithm is a commonly used algorithm in the current field of reinforcement learning. It is widely applied in control systems, especially suitable for high-dimensional continuous control tasks. Compared with the traditional Deep Deterministic Policy Gradient (DDPG) algorithm, the TD3 algorithm has significant advantages in terms of stability and training efficiency. Especially in an environment with large noise, the TD3 algorithm can reduce the instability of the policy by means of delayed updates and limiting policy noise.
[0049] Specifically for the present invention, the TD3 algorithm outputs the adjustment amounts of the sliding mode control parameters (such as the attenuation rate of the switching gain and the proportional coefficient of the boundary layer thickness) through the Actor network. These adjustment amounts can respond to the dynamic changes of the system in real time and improve the adaptive ability of the control system. In addition, the policy is evaluated through the Critic network to further optimize the long-term control performance of the system.
[0050] With the help of the DRL model, the control system no longer depends on static parameters, but can automatically adjust the key control parameters according to the real-time operating environment and load conditions to achieve more accurate trajectory tracking and stronger robustness. In this way, the system can adaptively cope with different dynamic changes, such as load changes, speed changes, and environmental disturbances, etc., significantly improving the performance of the control system.
[0051] Furthermore, in the present invention, a disturbance compensation mechanism is designed to further enhance the robustness and disturbance resistance of the system. Specifically, the disturbance compensation mechanism embeds the total system disturbance value into the equivalent control term in the sliding mode control framework to achieve real-time compensation for system disturbances.
[0052] The core of the disturbance compensation mechanism is to feed forward and inject the total system disturbance value calculated by the multi-stage linear extended state observer (MLESO) into the equivalent control term to compensate for the disturbance.
[0053] Step S31 of the present invention proposes a sliding mode control framework and designs an adaptable sliding mode surface to achieve precise control of the trajectory tracking error of the mobile robot and effectively compensate for external disturbances of the system. The core of this step is to combine traditional sliding mode control with deep reinforcement learning (DRL) to dynamically adjust the control parameters, thereby enhancing the adaptability and robustness of the control system.
[0054] Input the trajectory tracking error, the towing swing amplitude, and the total system disturbance value into the deep reinforcement learning model to construct a multi-dimensional state space; Based on the double-delayed deep deterministic policy gradient algorithm, output the adjustment amounts of the sliding mode control parameters, including the attenuation rate of the switching gain and the proportional coefficient of the boundary layer thickness; Dynamically calculate the weight coefficient of the integral term in the sliding mode surface according to the current motion speed and the towing load mass; Fuse the equivalent control term, the switching control term, and the parameter adjustment results to generate a control model with dynamically adjusted parameters, and verify the convergence of the control model with dynamically adjusted parameters through the Lyapunov stability criterion.
[0055] In the embodiment of the present invention, the trajectory tracking error , the swing amplitude of the trailer and the total disturbance value are input into a deep reinforcement learning (DRL) model to construct a multi-dimensional state space. Specifically, the multi-dimensional state space is composed of the following multiple dynamic characteristics: ; Among them, is the trajectory tracking error, representing the deviation between the target and the actual state; is the change rate of the error, reflecting the speed of error change; is the trailer swing angle, representing the swing amplitude of the trailer part; is the total system disturbance value output by a multi-stage linear extended state observer (MLESO), covering the disturbances of the system body, the trailer chain, and the external environment; is the current motion speed, which affects the dynamic characteristics of the system; is the trailer load mass, which affects the dynamic response of the trailer system.
[0056] By taking these key information as inputs, the DRL model can fully understand the dynamic state of the system, thus providing more accurate feedback for the adjustment of sliding mode control. The construction of this multi-dimensional state space enhances the adaptability of the system in complex environments and lays a foundation for the subsequent optimization of control parameters.
[0057] Furthermore, in the deep reinforcement learning (DRL) model, the twin-delayed deep deterministic policy gradient (TD3) algorithm is adopted to dynamically adjust the sliding mode control parameters based on the Actor-Critic network structure. Specifically, the Actor network outputs the adjustment amount of the sliding mode control parameters by extracting information from the multi-dimensional state space : ; Among them, is the adjustment amount of the switching gain, which affects the smooth transition of the control signal; is the adjustment amount of the boundary layer thickness, which controls the degree of error smoothing; is the adjustment amount of the error integral weight, which balances the error convergence speed and the disturbance rejection ability.
[0058] Through the Critic network, the TD3 algorithm evaluates the current policy and further adjusts the parameters to minimize the long-term cumulative error based on the optimization update of the value function. The specific update process adopts a delayed policy, which ensures the convergence of policy optimization by reducing overfitting and improving stability.
[0059] Furthermore, the weight coefficient of the integral term in the sliding mode surface is dynamically calculated according to the current motion speed and the mass of the towed load . In the present invention, the weight coefficient of the integral term in the sliding mode surface is dynamically calculated and . These two coefficients determine the weights of the trajectory tracking error and its rate of change, and thus affect the response speed and robustness of the sliding mode surface. The sub-item weight coefficients and have the following dynamic adjustment mechanism: ; wherein, is the current time at the th integral weight coefficient, is the integral weight value for the next time step, which is the result of dynamic update; is a mapping function calculated based on the motion speed and the load mass for dynamically adjusting the weight of the integral term. This design can adjust the control strategy according to the actual working conditions of the system to ensure that the system has the best response performance under different speed and load conditions.
[0060] Finally, the system substitutes the parameters adjusted by deep reinforcement learning (including ) into the existing calculation to generate a control model with dynamically adjusted parameters , .
[0061] wherein, is the equivalent control term after dynamic parameter adjustment, is the switching control term after dynamic parameter adjustment.
[0062] The control model with dynamically adjusted parameters is verified by the Lyapunov stability criterion to ensure that the system can maintain closed-loop stability under all working conditions, that is, under the action of disturbances and errors, the system can still ensure the convergence of the trajectory tracking error.
[0063] It should be noted that by fusing the control terms and optimizing the parameters, the generated control model with dynamically adjusted parameters can adapt to various dynamic changes and improve the accuracy of trajectory tracking. The application of the Lyapunov criterion ensures the stability of the control model with dynamically adjusted parameters and further improves the reliability and safety of the system.
[0064] S4: According to the control model with dynamically adjusted parameters, the control signal of the main body robot is calculated in real time, and actuator dynamic response compensation is performed on the control signal of the main body robot.
[0065] Actuator dynamic response compensation includes motor torque limiting and steering angle constraint.
[0066] It should be noted that this compensation ensures that the control signal of the main body robot can be reasonably and safely executed in the actuator, avoiding system overload or instability.
[0067] Motor torque limit is used to limit the maximum output of the motor torque, preventing damage or performance degradation caused by loads beyond the motor's bearing capacity.
[0068] Steering angle constraint is used to limit the steering angle, preventing excessive steering angles from causing unstable robot movement, especially at high speeds or in complex terrains.
[0069] After compensation, the final control signal of the main body robot is converted into an actual executable drive instruction and transmitted to the main body robot through the actuator, thereby achieving precise trajectory tracking control.
[0070] S5: Monitor the trajectory tracking error and the towing swing amplitude error after the execution of the control signal of the main body robot, adjust the control strategy according to the error results, and feedback it to the trajectory tracking control model for adaptive optimization.
[0071] Real-time collect the actual trajectory data and the towing swing angle data after the execution of the control signal of the main body machine, and calculate the trajectory tracking error and the towing swing amplitude error at the current moment; Set dynamic error thresholds. When the trajectory tracking error exceeds the first threshold or the towing swing amplitude error exceeds the second threshold, trigger the online parameter re-optimization process of the deep reinforcement learning model and update the sliding mode control parameters; If the error continues to exceed the limit and does not converge after optimization, switch to the emergency control mode and adopt the preset conservative sliding mode parameters and fixed boundary layer thickness; Feedback the updated control parameters or emergency control instructions to the trajectory tracking control model to form a closed-loop adaptive optimization.
[0072] In the embodiment of the present invention, the system calculates the trajectory tracking error and the towing swing amplitude error at the current moment by real-time collecting the actual trajectory data and the towing swing angle data after the execution of the control signal of the main body machine.
[0073] The trajectory tracking error reflects the deviation between the current position of the robot and the target trajectory. This error represents the current trajectory accuracy of the system.
[0074] The towing swing amplitude error represents the deviation between the actual swing angle of the towing system and the target swing angle. The towing swing amplitude error affects the stability of the robot and the towing system, especially the control performance under complex road conditions.
[0075] The system sets dynamic error thresholds based on the real-time calculated trajectory tracking error and the trailer swing amplitude error. When the error exceeds the set first and second thresholds, it triggers the online parameter re-optimization process of the deep reinforcement learning (DRL) model, thereby dynamically adjusting the control strategy.
[0076] The first threshold is the maximum allowable value of the trajectory tracking error. When it is exceeded, it means that the system deviates too far from the trajectory and the control parameters need to be optimized.
[0077] The second threshold is the maximum allowable value of the trailer swing amplitude. When it is exceeded, it means that the swing amplitude of the trailer system is too large, affecting the system stability and timely adjustment is required.
[0078] After the error exceeds the threshold, the DRL model will perform online optimization of the sliding mode control parameters according to the current trajectory tracking error and the trailer swing amplitude error. Specifically, the system dynamically adjusts the sliding mode control parameters (such as switching gain, boundary layer thickness, etc.) to improve the current control effect.
[0079] This optimization process gradually adjusts the control parameters according to the feedback mechanism of the deep reinforcement learning model, enabling the system to adapt to different working environments and avoid further expansion of the error.
[0080] If during the optimization process of the deep reinforcement learning model, the error continuously exceeds the threshold and fails to converge after optimization, the system will switch to the emergency control mode. In the emergency mode, the system will adopt preset conservative sliding mode parameters and a fixed boundary layer thickness to ensure that the system can still maintain a certain stability when the error exceeds the limit.
[0081] The emergency control mode effectively reduces the risk of the system and avoids further instability caused by excessive system adjustment.
[0082] In the emergency control mode, or based on the feedback result after the optimization of the deep reinforcement learning model, the control command will be fed back to the trajectory tracking control model to form a closed-loop adaptive optimization. This can continuously optimize the control parameters to ensure that the robot and the trailer system maintain efficient and accurate operation in a dynamic environment. Through continuous adaptive optimization, the system can adjust the control strategy in real time as the operating environment changes, thereby maximizing the trajectory tracking accuracy and the trailer stability.
[0083] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
[0084] Embodiment 2 is the second embodiment of the present invention. The difference from the previous embodiment is as follows: If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the current technical solution can be embodied in the form of a software product. The current computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0085] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0086] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or processing in other suitable ways when necessary, and then stored in a computer memory.
[0087] Embodiment 3 is an embodiment of the present invention, which provides a mobile robot trajectory tracking control system, including a data acquisition module, a disturbance calculation module, a trajectory tracking control model construction module, a real-time control signal generation module, and an adaptive optimization feedback module.
[0088] Data acquisition module: Collect the original motion state data of the robot body and the towed part through multiple types of sensors, preprocess the original motion state data, and generate preprocessed motion state data; Disturbance calculation module: Calculate the total system disturbance value based on the preprocessed motion state data; Trajectory tracking control model construction module: Construct a trajectory tracking control model, input the preprocessed motion state data and the total system disturbance value, optimize the sliding mode control parameters, and generate a control model with dynamically adjusted parameters; Real-time control signal generation module: According to the control model with dynamically adjusted parameters, calculate the control signal of the main body robot in real time, and perform actuator dynamic response compensation on the control signal of the main body robot; Adaptive optimization feedback module: Monitor the trajectory tracking error and the towed swing amplitude error after the execution of the control signal of the main body robot, adjust the control strategy according to the error results, and feedback to the trajectory tracking control model for adaptive optimization.
Claims
1. A mobile robot trajectory tracking control method, characterized in that: include: Collecting original motion state data of the robot body and the towing part through multiple types of sensors, preprocessing the original motion state data, and generating preprocessed motion state data; Calculating a total disturbance value of the system based on the preprocessed motion state data; Constructing a trajectory tracking control model, and inputting the preprocessed motion state data and the total disturbance value of the system, optimizing the sliding mode control parameters, and generating a control model after dynamic parameter adjustment; According to the control model after dynamic parameter adjustment, the control signal of the main robot is calculated in real time, and the dynamic response compensation of the actuator is performed on the control signal of the main robot; The trajectory tracking error and the trailer swing amplitude error after the execution of the main robot control signal are monitored, the control strategy is adjusted according to the error result, and the feedback is fed back to the trajectory tracking control model for adaptive optimization.
2. The mobile robot trajectory tracking control method according to claim 1, characterized in that: The collecting of the original motion state data of the robot body and the trailer part by using multiple types of sensors includes collecting the motion state data of the robot body by using a body sensor group, wherein the body sensor group includes an inertial measurement unit, a wheel speed encoder and a steering motor current sensor; The dynamic state data of the trailer part is collected by a trailer sensor group, wherein the trailer sensor group includes a trailer chain tension sensor, a trailer swing angle gyroscope, and a six-dimensional torque sensor of the hitch point; External disturbance feature data is collected through an environment perception sensor group, which includes a binocular vision camera, a solid-state laser radar and a millimeter wave radar array.
3. The mobile robot trajectory tracking control method according to claim 2, characterized in that: The preprocessing of the original motion state data includes performing multi-source fusion filtering on the motion state data of the robot body collected by the body sensor group to extract the state vector of the robot body, wherein the state vector of the robot body includes position, speed, heading angle and normalized steering motor current; The towing chain tension data and swing angle data collected by the towing sensor group are denoised, and the towing chain coupling disturbance characteristic matrix is constructed based on the six-dimensional torque data of the attachment point; Perform spatiotemporal registration of images, laser point clouds, and radar data collected by the environment perception sensor group, and generate environment-robot interaction feature tensors through spatial transformation; The state vector of the robot body, the tow chain coupling disturbance characteristic matrix and the environment-robot interaction characteristic tensor are uniformly normalized, and the preprocessed motion state data is output.
4. The mobile robot trajectory tracking control method according to claim 3, characterized in that: The calculating of the total disturbance value of the system based on the preprocessed motion state data comprises constructing a multi-level linear extended state observer, wherein the multi-level linear extended state observer comprises a first-level observer, a second-level observer and a third-level observer; The state vector of the robot body after preprocessing is input into the first-level observer, and the body dynamic disturbance value is calculated in real time through the linear extended state observation equation, wherein the body dynamic disturbance value includes the motor nonlinear friction and wheel-ground contact slip disturbance; The pre-processed towing chain coupling disturbance characteristic matrix is input into the second-level observer, and the coupling disturbance value caused by the trailer inertia force and the flexible deformation of the coupling point is calculated in real time through the linear expansion state observation equation; The preprocessed environment-robot interaction feature tensor is input into the third-level observer, and the key environmental disturbance features are extracted through the attention mechanism to calculate the environmental disturbance value; The main body dynamics disturbance value, coupling disturbance value and environmental disturbance value are integrated to generate the total disturbance value of the system.
5. The mobile robot trajectory tracking control method according to claim 4, characterized in that: The constructing of the trajectory tracking control model includes establishing a sliding mode control framework and designing a sliding mode surface, wherein the sliding mode surface is composed of a trajectory tracking error integral term and an adaptive compensation term based on a total disturbance value of the system; Constructing a sliding mode control law, wherein the sliding mode control law includes an equivalent control term and a switching control term, wherein the equivalent control term is generated based on a nominal dynamic model of the system and a total disturbance value of the system; A deep reinforcement learning model is integrated to dynamically adjust the sliding mode control parameters through a policy network. The sliding mode control parameters include switching gain, boundary layer thickness and sliding surface weight coefficient.
6. The mobile robot trajectory tracking control method according to claim 5, characterized in that: Generating the control model after dynamic parameter adjustment includes inputting the trajectory tracking error, the trailer swing amplitude and the total disturbance value of the system into the deep reinforcement learning model to construct a multi-dimensional state space; Based on the double-delayed deep deterministic policy gradient algorithm, the adjustment amount of the sliding mode control parameters is output, including the switching gain attenuation rate and the boundary layer thickness proportional coefficient; According to the current movement speed and the mass of the towing load, the weight coefficient of the integral term in the sliding surface is dynamically calculated; The equivalent control items, switching control items and parameter adjustment results are integrated to generate a control model after dynamic parameter adjustment, and the convergence of the control model after dynamic parameter adjustment is verified by the Lyapunov stability criterion.
7. The mobile robot trajectory tracking control method according to claim 6, characterized in that: The actuator dynamic response compensation includes motor torque limitation and steering angle constraint.
8. The mobile robot trajectory tracking control method according to claim 7, characterized in that: The adaptive optimization includes collecting the actual trajectory data and the trailer swing angle data after the control signal of the main robot is executed in real time, and calculating the trajectory tracking error and the trailer swing amplitude error at the current moment; Set a dynamic error threshold. When the trajectory tracking error exceeds the first threshold or the trailer swing amplitude error exceeds the second threshold, trigger the online parameter re-optimization process of the deep reinforcement learning module to update the sliding mode control parameters. If the error continues to exceed the limit and does not converge after optimization, it switches to the emergency control mode and uses the preset conservative sliding mode parameters and fixed boundary layer thickness; The updated control parameters or emergency control instructions are fed back to the trajectory tracking control model to form a closed-loop adaptive optimization.
9. A mobile robot trajectory tracking control system, used to implement the mobile robot trajectory tracking control method according to any one of claims 1 to 8, characterized in that: include: Data acquisition module: collects the original motion state data of the robot body and the towing part through multiple types of sensors, pre-processes the original motion state data, and generates pre-processed motion state data; Disturbance calculation module: calculates the total disturbance value of the system based on the preprocessed motion state data; Trajectory tracking control model building module: building a trajectory tracking control model, and inputting the pre-processed motion state data and the total disturbance value of the system, optimizing the sliding mode control parameters, and generating a control model after dynamic parameter adjustment; Real-time control signal generation module: calculates the control signal of the main robot in real time according to the control model after dynamic parameter adjustment, and performs actuator dynamic response compensation on the control signal of the main robot; Adaptive optimization feedback module: monitors the trajectory tracking error and the trailer swing amplitude error after the execution of the main robot control signal, adjusts the control strategy according to the error result, and feeds back to the trajectory tracking control model for adaptive optimization.
Citation Information
Patent Citations
Hitch assist system
CN110884308A
Traction type trailer trajectory tracking method based on robust H infinite control
CN111352442A
Unmanned helicopter tracking control method considering input saturation
CN114237270A
Unmodeled compensation control method for batch operation vehicles
CN117492453A
Rope traction parallel robot based on double-rope model and control method and device thereof
CN117656036A
Cited By
Error compensation control method for self-stabilizing holder under multi-branch redundancy cooperation
CN120669548A
Anti-interference control method and system for Mecanum wheel omnidirectional robot
CN121043160A
A disturbance rejection control method and system for a Mecanum wheel omnidirectional robot
CN121043160B
Multi-source data-based motion recovery control method and system
CN121276966A
Robot intelligent identification and remote cooperation method and system
CN121386517A