High-order robust control method for mobile robots using weighted double-Q learning algorithm
By combining a weighted double-Q learning algorithm and a nonlinear extended state observer with super-helical sliding mode control, the trajectory tracking problem of mobile robots in complex environments is solved, achieving high-precision and stable trajectory tracking control, reducing chattering, and improving the system's adaptability and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing control methods struggle to achieve high-precision trajectory tracking for mobile robots in complex environments, especially under the influence of factors such as changes in ground friction, load fluctuations, external disturbances, and sensing errors, which leads to a decline in control performance. Furthermore, traditional methods suffer from chattering and response hysteresis issues.
By employing a weighted double-Q learning algorithm combined with a nonlinear extended state observer (NLESO) and super-helical sliding mode control, the observer gain and control parameters are adaptively adjusted to achieve real-time estimation and compensation for external disturbances, thus constructing an adaptive control framework to ensure the stability and accuracy of trajectory tracking.
High-precision trajectory tracking of mobile robots was achieved in complex environments, reducing jitter, improving the system's adaptability and robustness, and ensuring the smoothness and stability of control.
Smart Images

Figure CN121541484B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology for mobile robots, and in particular to a high-order robust control method for mobile robots that combines a weighted double-Q learning algorithm. Background Technology
[0002] With the continuous development of automation and intelligent transportation technologies, intelligent vehicles have been widely used in industrial handling, warehousing and logistics, and unmanned delivery. However, achieving stable and high-precision trajectory tracking in complex environments remains a challenge. During operation, vehicles are susceptible to multiple factors such as changes in ground friction, load fluctuations, external disturbances, and sensor errors, causing their trajectories to deviate from the expected path and consequently reducing control performance. Existing control methods are mostly designed based on ideal operating conditions, making it difficult to effectively handle nonlinear and time-varying characteristics. Traditional PID control is prone to overshoot and lag under strong disturbance conditions, while ordinary sliding mode control, although possessing certain robustness, suffers from chattering due to frequent switching actions, reducing system stability. Among disturbance compensation methods, although the Nonlinear Extended State Observer (NLESO) can improve estimation performance through nonlinear feedback, its gain parameter... Typically, the system remains fixed, lacking adaptive adjustment capabilities. This makes it difficult to simultaneously balance response speed and estimation accuracy under different operating conditions, resulting in inconsistent control performance across various environments. Therefore, designing a trajectory tracking control method that can adaptively adjust the observer gain while maintaining high accuracy and low chattering characteristics has become a critical problem urgently needing to be solved in the field of intelligent vehicle control.
[0003] The paper "Nonlinear Extended State Observer Based Active Disturbance Rejection Control for Mobile Robots" (Y. Li et al., ISA Transactions, 2021) proposes an active disturbance rejection control method based on NLESO for trajectory tracking of wheeled mobile robots under conditions of nonlinear disturbances and parameter uncertainties. This method improves the observation accuracy of traditional ESO by introducing a nonlinear error feedback function and enhances the system's disturbance rejection performance to some extent. However, the gain parameters of this type of NLESO are usually set manually based on experience or with a fixed bandwidth, lacking the ability to adaptively adjust to environmental changes. When the system operates under different conditions (such as time-varying friction, slope disturbances, load fluctuations, etc.), a fixed observer bandwidth is difficult to balance fast response and noise immunity: too large a bandwidth leads to amplified observation noise and increased output oscillations; too small a bandwidth causes lag in disturbance estimation, affecting control accuracy. Therefore, traditional NLESO still suffers from problems such as fixed parameters, hysteresis, and noise sensitivity in complex dynamic environments, limiting its control robustness and adaptive performance under multiple conditions.
[0004] The paper "Trajectory Tracking of Wheeled Mobile Robots Using Conventional Sliding Mode Control" (R. Olfati-Saber, IEEE Conference on Decision and Control, 2017) proposes a nonlinear control method based on low-order sliding mode control for trajectory tracking of wheeled intelligent robots. This method achieves finite-time convergence of errors by designing a sliding surface, exhibiting good disturbance rejection capabilities. However, the switching term in low-order sliding mode control is a sign function, which is prone to high-frequency chattering, leading to actuator vibration, energy loss, and attitude fluctuations. Especially in high-precision trajectory tracking, chattering can cause discontinuities in the control signal, affecting stability. To improve this issue, subsequent research proposed a super-helical sliding mode control strategy, replacing the sign switching term with a continuous power function, enabling the system to converge smoothly near the sliding surface, significantly reducing chattering and improving control stability.
[0005] How to solve the above-mentioned technical problems is the challenge facing this invention. Summary of the Invention
[0006] The purpose of this invention is to provide a high-order robust control method for mobile robots combined with a weighted double-Q learning algorithm, solving the problem of decreased trajectory tracking accuracy caused by the difficulty of mobile robots to adapt to complex and abruptly changing environments in real time. Due to factors such as changes in ground friction coefficient, load fluctuations, and sensor noise, the vehicle often exhibits increased system uncertainty and discontinuous controller output during actual operation, making it difficult to ensure stable and reliable motion control. To address this, a nonlinear extended state observer is designed to uniformly model and estimate the total disturbance, enabling the system to accurately compensate for external disturbances under different operating conditions. Combined with the double-Q learning algorithm, by constructing state, action, and reward mechanisms, adaptive adjustment of the observer gain and optimization of control parameters are achieved, thereby improving the system's self-learning and self-adjustment capabilities. At the control layer, a super-spiral sliding mode algorithm is employed, utilizing a continuous switching law to reduce high-frequency chattering and maintain smooth control output. This design, by forming a closed-loop structure in the observation, learning, and control stages, achieves accurate tracking and stable control of the vehicle's trajectory under uncertain dynamic environments.
[0007] The inventive concept of this invention is to construct an adaptive control framework that integrates observation, learning, and sliding mode control. A nonlinear dynamic model of an intelligent vehicle is established in a simulation platform, and random disturbances and noise are introduced to reproduce complex operating conditions. Real-time estimation of the total system disturbance is achieved through NLESO, and a gain adjustment mechanism is designed in conjunction with a dual-Q learning algorithm, enabling the system to autonomously select appropriate observer parameters based on state changes during operation. The control section employs a superspiral sliding mode structure to ensure continuous, smooth, and robust control signals. The entire system achieves dynamic optimization of trajectory errors and adaptive adjustment of control performance through closed-loop interaction.
[0008] To achieve the above-mentioned objectives, the present invention adopts a technical solution that includes the following steps:
[0009] Step 1: Construct a motion scenario for the intelligent vehicle in the MATLAB / Simulink simulation platform (including the Simscape Multibody module), which includes complex uncertainties such as time-varying friction, variable slope, random external disturbances, and sensor noise.
[0010] Step 2: Design a Weighted Double Q-Learning Adaptive Nonlinear Extended State Observer (WDQ-A-NLESO). This observer employs a nonlinear error feedback function. A variable bandwidth structure is also incorporated to enhance the system's ability to estimate external disturbances. Simultaneously, a weighted double-Q learning algorithm is introduced to adaptively adjust the observer gain parameters based on changes in system state error and disturbance. This enables rapid and accurate estimation and dynamic compensation of disturbances under different operating conditions, thereby realizing an adaptive and stable trajectory tracking strategy for intelligent vehicles.
[0011] Step 3: Based on Step 2, a reward function is constructed to address the trajectory tracking error, disturbance estimation error, and stationarity constraints caused by complex environmental changes in the mobile robot system. Through continuous strategy exploration and parameter updates, the intelligent vehicle can autonomously learn the optimal trajectory tracking under different disturbance conditions.
[0012] Step 4: Design an improved superspiral sliding mode control law based on WDQ-A-NLESO to generate low-level control inputs. The control law form is as follows: The control law is based on the sliding surface. Based on this, the sliding mode parameters are updated online adaptively through the double Q learning algorithm, and the control law is corrected in real time based on the latest disturbance estimation results of the observer, thereby achieving smooth and continuous high-precision trajectory tracking and effectively suppressing chattering.
[0013] Step 5: Establish multiple sets of working condition simulation comparison schemes, conduct simulation experiments to verify the results, and conduct comparative analysis with the traditional sliding mode control method based on extended state observer (ESO-SMC) in the simulation environment.
[0014] Furthermore, the simulation scenario built in step 1 includes several training condition sets, which include at least 5 typical conditions: flat road surface, low friction road surface, high friction road surface, uphill and downhill. New combined perturbation scenarios are added to the test set to test the generalization ability of the algorithm. The perturbation sources in the simulation include sinusoidal time-varying perturbation, random impulse perturbation and Gaussian noise superposition.
[0015] Furthermore, the weighted double-Q learning adaptive nonlinear extended state observer in step 2 consists of three parts: a master observer module, an extended state compensation channel, and an adaptive gain adjustment unit. The master observer module uses the system input and output signals as its core and employs a nonlinear feedback structure to estimate the main state variables of the system. The extended compensation channel is used to identify and compensate for external disturbances and unmodeled dynamics. The adaptive adjustment unit, combined with a weighted double-Q learning strategy, adjusts the observer gain parameters in real time based on the state error and disturbance estimation error, thereby achieving a balance between observation accuracy and response speed under dynamic conditions.
[0016] For the trajectory tracking control system of a wheeled mobile robot, to achieve nonlinear dynamic compensation based on an extended state observer, it is necessary to first establish an error system that matches the observer structure. During robot operation, due to friction variations, load disturbances, and model uncertainties, a deviation will occur between the actual trajectory and the desired trajectory. To describe this deviation, the desired trajectory is mapped in the robot's body coordinate system, and the attitude error vector is defined as follows:
[0017] ;
[0018] In the formula, This represents the linear position error along the direction of travel. This indicates the lateral offset error. This represents the heading angle error.
[0019] By differentiating the error equation from the kinematic relationships and combining it with the system's geometric constraints and nonholonomic constraints, the dynamic error model can be obtained:
[0020] ;
[0021] In the formula, and These represent the equivalent disturbance coefficients in the angular velocity channel and the linear velocity channel, respectively. , , Each constitutes a disturbance term ; For the desired angular velocity, For the desired linear velocity, , These are the car's current actual angular velocity and actual linear velocity, respectively. , , Each constitutes a known term ;
[0022] In the error dynamic equation These are the computable known terms of the system, determined by the current error and control input. Since these terms are available in real time during trajectory tracking, they can be used as known inputs to the ESO to help estimate disturbance components in the system that cannot be directly measured.
[0023] Based on the above error model, the observer adopts a dual-channel design, establishing nonlinear extended observation equations for the linear velocity and angular velocity subsystems respectively. In the linear velocity channel, an error is assumed... Its observer form is:
[0024] ;
[0025] In the formula, This represents the actual position along the x-axis. These are the estimated position, estimated linear velocity, and estimated total system disturbance of the vehicle along the x-axis, respectively. These are the adaptive gain parameters of the observer; It is a nonlinear continuous function. It is a non-linear exponent. This is the bandwidth adjustment coefficient.
[0026] The angular velocity channel structure is the same as itss:
[0027] ;
[0028] In the formula, These are the estimated angular velocity of the trolley, the derivative of the estimated angular velocity, and the estimated total system disturbance, respectively. These are the adaptive gain parameters of the observer;
[0029] The dual-Q module acts as an intermediate adjustment stage, using feedback error signals to adjust the gain parameters. Lightweight online adjustments are performed. This structure enables the observer to form a three-layer system consisting of a state estimation layer, a disturbance compensation layer, and a gain self-tuning layer, achieving an organic integration of multi-channel coordination and dynamic parameter correction, thereby ensuring the stability and consistency of the overall observer structure under different operating conditions.
[0030] Furthermore, the gain adaptive adjustment unit establishes an observer parameter adjustment model through the double Q learning algorithm, and models the observer gain adjustment process of the intelligent vehicle as a Markov Decision Process (MDP), aiming to achieve adaptive optimization of the observer parameters through policy learning.
[0031] To achieve the above objectives, the present invention will use the state space. Defined as a state vector obtained by multi-sensor fusion, it includes the magnitude of observation error, the rate of error change, short-time variance, and noise estimation, reflecting the instantaneous dynamic state of the system. Action This includes three types of adjustment actions for the gain of the weighted double-Q learning adaptive nonlinear extended state observer: increasing the gain, keeping it constant, and decreasing the gain.
[0032] Furthermore, in step 3, the dual-Q algorithm includes two independent sets of Q-value functions, an action selection unit, and a parameter update module. The two sets of Q-value functions are denoted as follows: and The action decision unit selects actions based on a weighted Q-value strategy, and its comprehensive evaluation function is defined as follows:
[0033] ;
[0034] In the formula These are weighting coefficients. It belongs to the set of states. It belongs to the action set.
[0035] This unit adopts - Greedy strategy for action selection, i.e., based on probability Conduct random exploration, with probability Select the one with the largest current value. The optimal action is determined to achieve a balance between exploring new strategies and utilizing existing experience. The weighted update module employs an alternating update mechanism, enabling the two sets of Q-tables to achieve collaborative convergence while learning independently. Its update form is as follows:
[0036] ;
[0037] In the formula, , For two sets of independent Q-value functions, , These represent the maximum value estimates of all possible actions for the two sets of Q-tables in the new state. For learning rate, As a discount factor, For instant rewards, These are weighting coefficients. The state before the transfer. The new state after the state transition. The current state The selected action In the state The selected action;
[0038] Through this structure, the dual-Q network can mutually correct each other while maintaining the independence of estimation, effectively avoiding the overestimation problem in single-Q learning, and realizing the adaptive adjustment and stable update of the gain parameters of the weighted dual-Q learning adaptive nonlinear extended state observer.
[0039] After updating the double Q-value function, the algorithm selects the appropriate gain adjustment action based on the current state to achieve adaptive adjustment of the observer parameters. Specifically, the action set... Composed of discrete gain adjustment factors, used to dynamically correct the gain parameters of the extended state observer. .
[0040] After the action is executed, the observer gain is updated in the following form:
[0041] ;
[0042] In the formula, For parameter-constrained operators, The adaptive gain of the observer at the current moment. and These represent the minimum and maximum values of the gain, respectively. As a reference initial value, The current state The selected action Indicates the magnitude change in gain;
[0043] Furthermore, in step 3, the reward function structure consists of an error evaluation unit, a disturbance suppression unit, and a gain balancing unit, used to comprehensively evaluate the system's control performance and observation accuracy in each interaction. The reward function is defined as:
[0044] ;
[0045] In the formula, For instant rewards, For linear velocity error, For the rate of change of speed error, For the short-time variance of the total disturbance, The adaptive gain of the observer at the current moment. Its initial reference value; These are the respective weighting coefficients, used to balance the relative impacts of error accuracy, response speed, disturbance suppression, and gain adjustment.
[0046] This structure, through a negative feedback reward design for tracking error and disturbance estimation error, makes the double-Q learning algorithm tend to select actions that reduce system error and disturbance energy during the optimization process, while constraining drastic changes in the gain parameter. This achieves a dynamic balance between stability and robustness, ensuring the convergence performance of the extended state observer and the global tracking accuracy of the control system.
[0047] Furthermore, in step 4, the final control law consists of an equivalent control term and a super-helical sliding mode term, which is used to achieve stable trajectory tracking control under uncertain disturbances.
[0048] Let the sliding surface be Combining the error system and the observer output, the equivalent control term can be obtained:
[0049] ;
[0050] In the formula, Responsible for offsetting known terms and estimated disturbances in the system. The equivalent control components related to the system error state, the computable part of the error system. This is the control channel gain matrix, used to describe the sensitivity of the control signal to changes in system error. and This is the proportionality coefficient. This is an estimate of the interference.
[0051] To suppress sliding mode chattering and improve control continuity, a super-spiral sliding mode correction term is introduced:
[0052] ;
[0053] In the formula, To ensure rapid error convergence and smooth control, For sliding surface, For sliding mode parameters, For internal correction variables;
[0054] This leads to the final control law:
[0055] ;
[0056] The observer gain parameters are learned through the double-Q learning module. Online adjustments, The estimation accuracy is optimized, making The compensation effect is more effective, thereby enhancing the adaptive and robust performance of the control law under different operating environments.
[0057] Furthermore, in step 5, each experiment is conducted under different types of simulation and actual working conditions, including smooth road surfaces, sloping road surfaces, and unstructured environments with random disturbances. Diverse test scenarios are generated by randomly setting initial positions, target trajectories, disturbance intensity, and noise parameters. Trajectory deviations, attitude changes, and control input signals are recorded, and the system is verified under multiple independent test samples to ensure experimental consistency and control characteristic analysis under different environmental conditions.
[0058] To achieve the above-mentioned objectives, the present invention also provides a high-order robust control system for a mobile robot combined with a weighted double-Q learning algorithm, the system comprising:
[0059] The MATLAB / Simulink simulation platform is configured to perform the following process: for intelligent vehicle motion scenarios with complex uncertainties such as time-varying friction, variable slope, random external disturbances and sensor noise.
[0060] The weighted double-Q learning adaptive nonlinear extended state observer module is configured to perform the following process: employing a nonlinear error feedback function. In addition, a variable bandwidth structure is introduced, along with a weighted double-Q learning algorithm to adaptively adjust the observer gain parameters based on changes in system state error and disturbance. To achieve rapid and accurate estimation and dynamic compensation of disturbances under different working conditions, and to realize an adaptive and stable trajectory tracking strategy for intelligent vehicles;
[0061] The reward function construction module is configured to perform the following process: In response to trajectory tracking errors, disturbance estimation errors and stationarity constraints caused by complex environmental changes in the mobile robot system, the intelligent vehicle autonomously learns the optimal trajectory tracking under different disturbance conditions through continuous policy exploration and parameter updates.
[0062] The Weighted Double Q-Learning Adaptive Nonlinear Extended State Observer (WDQ-A-NLESO) module is configured to perform the following process: Generating low-level control inputs based on an improved superspiral sliding mode control law, with the control law taking the form of… Control law based on sliding surface Based on this, the sliding mode parameters are updated online adaptively through the double Q learning algorithm, and the control law is corrected in real time based on the latest disturbance estimation results of the observer, so as to achieve smooth and continuous high-precision trajectory tracking and effectively suppress jitter.
[0063] Multiple operating condition simulation modules are configured to execute the following process: by establishing multiple operating condition simulation comparison schemes, simulation experiments are conducted to verify the results, and comparative analysis is performed with the traditional sliding mode control method based on extended state observers in the simulation environment.
[0064] Meanwhile, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed, it implements the steps of the method of the present invention.
[0065] Furthermore, the present invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program being configured to implement the steps of the method of the present invention when invoked by a processor.
[0066] Finally, the present invention provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method of the present invention.
[0067] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0068] 1. An adaptive nonlinear extended state observer combining a weighted double-Q learning algorithm is used. By setting up a cooperative structure between the double-Q value network and the gain adjustment stage, the observer gain parameter is adjusted. The method employs dynamic self-adjustment. Using system state error and disturbance estimation error as feedback signals, it adaptively adjusts the observation bandwidth based on the operating state, thereby improving response speed when disturbances change rapidly and reducing gain to suppress oscillations when noise is strong, effectively balancing estimation accuracy and system stability.
[0069] 2. The weighted double-Q learning algorithm introduces weighting coefficients. By fusing the two sets of Q-value functions, the overestimation problem in single-Q learning is avoided, and the underestimation defect in traditional double-Q learning is overcome. This structure can achieve a dynamic balance between exploration and exploitation, making the gain adjustment process both sufficiently exploratory and ensuring the stability of policy convergence, thereby obtaining more accurate and robust parameter update results in complex environments.
[0070] 3. Combine WDQ-A-NLESO with the superspiral sliding mode control law, and estimate the disturbance. The adaptive compensation enables the control law to counteract the effects of external disturbances and unmodeled dynamics online. This control structure utilizes a continuous sliding mode switching law to generate smooth control inputs, avoiding the chattering phenomenon caused by high-frequency switching in traditional sliding mode control, while maintaining fast convergence and strong robustness, enabling the system to achieve smooth and high-precision trajectory tracking control under multiple operating conditions. Attached Figure Description
[0071] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0072] Figure 1 This is a flowchart of the high-order robust control method for mobile robots combined with the weighted double-Q learning algorithm of the present invention.
[0073] Figure 2 This is a simulation scene diagram built on the simulation platform according to the present invention.
[0074] Figure 3 This is a system structure diagram of the present invention.
[0075] Figure 4 This is a comparison chart of trajectory tracking errors in this invention.
[0076] Figure 5 This is a comparison chart of trajectory tracking errors in this invention. Detailed Implementation
[0077] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0078] Example 1
[0079] See Figures 1 to 5 This embodiment provides the following technical solution: The present invention proposes a high-order robust control method for mobile robots combined with a weighted double-Q learning algorithm to achieve high-precision and robust trajectory tracking control in complex dynamic environments. First, a trajectory tracking simulation scenario for an intelligent vehicle is built in the Matlab / Simulink platform, comprehensively considering various dynamic factors such as road friction changes, slope disturbances, and sensor noise to simulate the complexity and uncertainty of the real operating environment. Second, a nonlinear extended state observer with a gain adaptive adjustment mechanism is designed to estimate the total system disturbance in real time and dynamically correct model errors, thereby achieving synchronous observation of state and disturbance. Then, a double-Q learning algorithm is introduced to establish a gain optimization decision model. Through online updates of state feedback and reward functions, the observer parameters are self-learned and adjusted, enabling the system to maintain control stability under different operating conditions. Furthermore, a super-spiral sliding mode control law is used to generate continuous control input signals to reduce the high-frequency chattering problem in traditional sliding mode control and improve the smoothness and accuracy of trajectory tracking. Finally, the performance of this method is compared with that of the traditional ESO-SMC control strategy in a simulation environment to verify the advantages of the algorithm in terms of convergence speed and disturbance rejection performance.
[0080] The method flow of this embodiment is as follows: Figure 1 As shown, the specific method is as follows:
[0081] 1. Construct a motion scenario for an intelligent vehicle that includes complex uncertainties such as time-varying friction, variable slope, random external disturbances, and sensor noise in the MATLAB / Simulink simulation platform (including the Simscape Multibody module).
[0082] This example uses the Matlab / Simulink and Simscape Multibody co-simulation platform to build a simulation system for verifying the trajectory tracking control algorithm of an intelligent vehicle. The vehicle model is based on a four-wheel independent drive structure design, with the front wheels used for steering control and the rear wheels providing the main driving torque to achieve differential assisted steering and attitude stability. The control algorithm is built in the Simulink environment, including NLESO, a gain adaptive adjustment module based on weighted double Q learning, and a superspiral sliding mode controller. The system acquires the vehicle's attitude angle, velocity, and displacement information through a multi-sensor fusion module. The sensing unit includes an inertial measurement unit (IMU) to provide real-time feedback signals. The simulation scene built in the Matlab simulation platform is shown in the figure below. Figure 2 As shown.
[0083] The simulation scenario includes several training condition sets, which include at least five typical conditions: flat road surface, low-friction road surface, high-friction road surface, uphill, and downhill. New combined perturbation scenarios are added to the test set to test the algorithm's generalization ability. The perturbation sources in the simulation include sinusoidal time-varying perturbations, random impulse perturbations, and Gaussian noise superposition.
[0084] 2. Design a Weighted Double Q-Learning Adaptive Nonlinear Extended State Observer (WDQ-A-NLESO). This observer employs a nonlinear error feedback function. A variable bandwidth structure is also incorporated to enhance the system's ability to estimate external disturbances. Simultaneously, a weighted double-Q learning algorithm is introduced to adaptively adjust the observer gain parameters based on changes in system state error and disturbance. This enables rapid and accurate estimation and dynamic compensation of disturbances under different operating conditions, thereby realizing an adaptive and stable trajectory tracking strategy for intelligent vehicles.
[0085] The weighted double-Q learning adaptive nonlinear extended state observer consists of three parts: a master observer module, an extended state compensation channel, and an adaptive gain adjustment unit. The master observer module, centered on the system input signal and output feedback, employs a nonlinear feedback structure to estimate the system's main state variables. The extended compensation channel identifies external disturbances and unmodeled dynamics, achieving dynamic compensation for system uncertainties. The adaptive gain adjustment unit, combined with a weighted double-Q learning strategy, adjusts the observer gain parameters online based on system state errors and disturbance estimation errors, thereby achieving a balance between observation sensitivity and response speed under different operating conditions.
[0086] For the trajectory tracking control system of a wheeled mobile robot, to achieve nonlinear dynamic compensation based on an extended state observer, an error system matching the observer structure is first established. Due to the influence of factors such as friction variations, load fluctuations, and model uncertainties, there is a deviation between the actual trajectory and the desired trajectory of the robot. To accurately describe this deviation, the desired trajectory is mapped to the robot's body coordinate system, and the attitude error vector is defined as follows:
[0087] ;
[0088] In the formula, This represents the linear position error along the direction of travel. This indicates the lateral offset error. This represents the heading angle error.
[0089] By differentiating the error equation from the kinematic relationships and combining it with the system's geometric constraints and nonholonomic constraints, the dynamic error model can be obtained:
[0090] ;
[0091] In the formula, and These represent the equivalent disturbance coefficients in the angular velocity channel and the linear velocity channel, respectively. , , Each constitutes a disturbance term ; For the desired angular velocity, For the desired linear velocity, , These are the car's current actual angular velocity and actual linear velocity, respectively. , , Each constitutes a known term ;
[0092] This error model provides the observer with a structure-matching input basis for joint compensation of state estimation and disturbance identification. In the aforementioned error dynamic equation, The known, computable quantities of the system are determined by the current attitude error and control input. Since these terms can be acquired in real time during control, they are used as input signals to the extended state observer (ESO) to assist in estimating disturbance components that cannot be directly measured in the system. Based on this error model, a weighted double-Q learning adaptive nonlinear extended state observer (ESO) is constructed. This ESO introduces nonlinear feedback and adaptive gain adjustment mechanisms to achieve parallel estimation of state and disturbance. This structure can balance fast response and high-precision estimation under different operating conditions.
[0093] Based on the above error model, the observer adopts a dual-channel design, establishing nonlinear extended observation equations for the linear velocity and angular velocity subsystems respectively. In the linear velocity channel, an error is assumed... Its observer form is:
[0094] ;
[0095] In the formula, This represents the actual position of the trolley along the x-axis. These are the estimated position, estimated linear velocity, and estimated total system disturbance of the vehicle along the x-axis, respectively. These are the adaptive gain parameters of the observer; It is a nonlinear continuous function. It is a non-linear exponent. This is the bandwidth adjustment coefficient.
[0096] The angular velocity channel structure is the same as itss:
[0097] ;
[0098] In the formula, These are the estimated angular velocity of the trolley, the derivative of the estimated angular velocity, and the estimated total system disturbance, respectively. These are the adaptive gain parameters of the observer;
[0099] The adaptive gain adjustment unit plays a central role in the entire observation structure. This unit incorporates a weighted double-Q learning algorithm to adjust the gain parameters. The algorithm performs incremental online updates. Using the system's observation error and disturbance estimation error as input, it dynamically adjusts the observer gain through a weighted action selection and update mechanism, achieving parameter self-tuning and bandwidth adaptation. This forms a three-layer collaborative structure of "state estimation—disturbance compensation—gain adjustment," enabling the observer to maintain dynamic stability and high-precision response even in complex environments.
[0100] In the observer structure, an adaptive gain adjustment unit embeds a weighted double-Q learning module to achieve adjustment of the gain parameter. The algorithm dynamically updates the system's observation errors, disturbance estimation errors, and their rates of change as state inputs, and corrects the observer parameters in real time through action selection and weighted update mechanisms. To achieve the above functions, a gain adjustment model is established, and the state set... Composed of observation error amplitude, error rate of change, short-time variance, and noise estimation, it reflects the instantaneous dynamic state of the system. Action set It includes three types of discrete adjustment behaviors: "increase gain", "keep it unchanged" and "decrease gain".
[0101] During operation, the observer utilizes a dual-Q module to optimize the gain of different channels online, ensuring stability and accuracy under varying operating conditions. The module's optimized parameters and the structure of the compensation system are shown in the diagram below. Figure 3 As shown, this structure forms an adaptive adjustment closed loop between the master observer and the expansion channel, which can automatically adjust the estimation bandwidth when external disturbances change, thereby enhancing the system's robustness and dynamic response performance to uncertain environments.
[0102] Based on step 2, a reward function is constructed to address the trajectory tracking error, disturbance estimation error, and stationarity constraints caused by complex environmental changes in the mobile robot system. Through continuous strategy exploration and parameter updates, the intelligent vehicle can autonomously learn the optimal trajectory tracking under different disturbance conditions.
[0103] The dual-Q algorithm comprises two independent Q-value functions, an action selection unit, and a parameter update module, used to adaptively optimize the gain parameters of the extended state observer. The two Q-value functions are used to independently evaluate the gain adjustment effect under different strategies. To prevent bias from arising from a single strategy during long-term learning, a weighted fusion mechanism is introduced, with the comprehensive evaluation function being:
[0104] ;
[0105] In the formula These are weighting coefficients. It belongs to the set of states. It belongs to the action set.
[0106] This unit adopts - Greedy strategy for action selection, i.e., based on probability Conduct random exploration, with probability Select the one with the largest current value. The optimal action is determined to achieve a balance between exploring new strategies and utilizing existing experience. The weighted update module employs an alternating update mechanism, enabling the two sets of Q-tables to achieve collaborative convergence while learning independently. Its update form is as follows:
[0107] ;
[0108] In the formula, , For two sets of independent Q-value functions, , These represent the maximum value estimates of all possible actions for the two sets of Q-tables in the new state. For learning rate, As a discount factor, For instant rewards, These are weighting coefficients. The state before the transfer. The new state after the state transition. The current state The selected action In the state The selected action;
[0109] Through this structure, the dual-Q learning network achieves mutual verification while maintaining estimation independence, avoiding the overestimation problem in value assessment inherent in the single-Q algorithm, and making the gain parameter adjustment more stable and reasonable. After the update is complete, the algorithm selects the appropriate gain adjustment action based on the current state, thereby achieving adaptive correction of the observer parameters. Action set It consists of three types of discrete gain adjustment factors, including "increase gain", "keep it unchanged" and "decrease gain", which are used to dynamically correct the parameters of the extended state observer.
[0110] After the action is executed, the observer gain is updated in the following form:
[0111] ;
[0112] In the formula, For parameter-constrained operators, The adaptive gain of the observer at the current moment. and These represent the minimum and maximum values of the gain, respectively. As a reference initial value, The current state The selected action Indicates the magnitude change in gain;
[0113] Meanwhile, the reward function structure consists of an error evaluation unit, a disturbance suppression unit, and a gain balancing unit, used to comprehensively measure the system's control performance and observation accuracy. The reward function is defined as follows:
[0114] ;
[0115] In the formula, For instant rewards, For linear velocity error, For the rate of change of speed error, For the short-time variance of the total disturbance, The adaptive gain of the observer at the current moment. Its initial reference value; These are the respective weighting coefficients, used to balance the relative impacts of error accuracy, response speed, disturbance suppression, and gain adjustment.
[0116] This structure, through a negative feedback reward design for tracking error and disturbance estimation error, makes the double-Q learning algorithm tend to select actions that reduce system error and disturbance energy during the optimization process, while constraining drastic changes in the gain parameter. This achieves a dynamic balance between stability and robustness, ensuring the convergence performance of the weighted double-Q learning adaptive nonlinear extended state observer and the global tracking accuracy of the control system.
[0117] 4. Design an improved superspiral sliding mode control law based on WDQ-A-NLESO to generate low-level control inputs. The control law form is as follows: The control law is based on the sliding surface. Based on this, the sliding mode parameters are updated online adaptively through the double Q learning algorithm, and the control law is corrected in real time based on the latest disturbance estimation results of the observer, thereby achieving smooth and continuous high-precision trajectory tracking and effectively suppressing chattering.
[0118] The superspiral sliding mode control structure consists of a sliding mode surface construction unit, a superspiral switching law generation unit, and a control input synthesis unit; let the sliding mode surface be... Combining the error system and the observer output, the equivalent control term can be obtained:
[0119] ;
[0120] In the formula, Responsible for offsetting known terms and estimated disturbances in the system. The equivalent control components related to the system error state, the computable part of the error system. This is the control channel gain matrix, used to describe the sensitivity of the control signal to changes in system error. and This is the proportionality coefficient. This is an estimate of the interference.
[0121] To suppress sliding mode chattering and improve control continuity, a super-spiral sliding mode correction term is introduced:
[0122] ;
[0123] In the formula, To ensure rapid error convergence and smooth control, For sliding surface, For sliding mode parameters, For internal correction variables;
[0124] This leads to the final control law:
[0125] ;
[0126] In the formula Responsible for offsetting known terms and estimated disturbances in the system. Ensure rapid convergence of errors and smooth control.
[0127] The observer gain parameters are learned through the double-Q learning module. Online adjustments, The estimation accuracy is optimized, making The compensation effect is more effective. This structure enables the control law to maintain high disturbance rejection performance and stability under complex operating conditions, ensuring both the fast convergence characteristics of sliding mode control and improving the overall performance of the control system. The adaptive optimization enhances the system's robustness and dynamic response capability.
[0128] 5. Establish multiple sets of working condition simulation comparison schemes, conduct simulation experiments to verify the results, and conduct comparative analysis with the traditional sliding mode control method based on extended state observer (ESO-SMC) in the simulation environment.
[0129] Verification simulation experiments were conducted on the MATLAB / Simulink and Simscape Multibody co-simulation platform to evaluate the robustness and accuracy of the proposed weighted double-Q learning adaptive nonlinear extended state observer combined with the superhelical sliding mode control algorithm under different trajectory and perturbation conditions. The constructed simulation model is a four-wheel differential drive mobile robot, and the dynamics part considers nonlinear effects such as ground friction coefficient fluctuations, rolling resistance changes, slope interference, and sensor noise. The system sampling period was set to 1 ms, the simulation time was 20 s, and each set of data is the Monte Carlo average result of 20 different random perturbation conditions to ensure the statistical stability and repeatability of the experimental results.
[0130] To verify the performance differences of the algorithm at different control levels, three control strategies were designed for comparison: ① NLESO-SMC: a low-order sliding mode control based on a nonlinear extended state observer, used to verify the basic performance of a fixed-gain ESO under disturbance compensation; ② NLESO-STSMC: a super-helical sliding mode control based on a nonlinear extended state observer, introducing a second-order sliding mode law under the same ESO structure to reduce chattering and improve control continuity; ③ WQ-NLESO-STSMC: an improved super-helical sliding mode control based on weighted double-Q learning adaptive NLESO, using a double-Q network to adjust the observer gain parameters. To achieve adaptive adjustment and balance rapid response and noise resistance stability under different operating conditions, the robot was tested by running along a "figure-eight curve trajectory" and a "circular trajectory" with three perturbation scenarios superimposed: periodic sinusoidal perturbation, step impact perturbation, and random friction perturbation.
[0131] Table 1 shows the Monte Carlo average results of twenty independent MATLAB simulations with different random seeds for the NLESO-SMC and WQ-NLESO-STSMC strategies on straight and circular trajectories.
[0132] Table 1
[0133]
[0134] Table 2 shows the Monte Carlo average results of twenty independent MATLAB simulations with different random seeds for the NLESO-SMC and WQ-NLESO-STSMC strategies on sinusoidal and figure-eight trajectories.
[0135] Table 2
[0136]
[0137] Tables 1 and 2 present a comparison of the comprehensive performance indicators of the NLESO-SMC and WQ-NLESO-STSMC strategies under different trajectory and disturbance conditions, including position mean square error, maximum angle error, average control energy consumption, and disturbance estimation error. The results show that both algorithms can maintain system controllability under all four types of disturbances, but there are significant differences in disturbance estimation accuracy and trajectory tracking error. Under irregular noise conditions, NLESO-SMC, due to the observer gain parameter... The fixed position, while highly sensitive to noise, causes fluctuations in observation errors and slight oscillations in the trajectory; whereas the WQ-NLESO-STSMC, through a dual-Q learning mechanism, […]. Real-time adaptive adjustment enables the observer bandwidth to change dynamically with noise intensity, thereby effectively suppressing high-frequency interference while maintaining response speed, and reducing the average steady-state error by about 42%.
[0138] Under impact disturbance conditions, the system is subjected to transient external forces, causing a sharp increase in tracking error. The fixed observation gain of the NLESO-SMC cannot quickly track changes in disturbance, with a recovery time of approximately 3.2 seconds; while the WQ-NLESO-STSMC utilizes... The channel's rapid estimation of the total disturbance enables online compensation, reducing the response time to 2.0 s and lowering the peak error by approximately 45%. This demonstrates that the dual-Q self-learning strategy can maintain the system's rapid recovery capability and control smoothness under impact disturbances.
[0139] When the load changes and the system mass suddenly increases, the traditional NLESO-SMC exhibits observation lag, leading to overcompensation in the control output and instantaneous overshoot; while the WQ-NLESO-STSMC corrects this online. The observer's dynamic gain automatically matches the new inertial characteristics, while the superhelical sliding mode control law utilizes... The estimated value compensates for the unmodeled dynamics in real time, thereby maintaining the continuity of the control signal and the stability of the output, and reducing the mean square error by about 37%.
[0140] In the most complex compound disturbance scenarios, the system is simultaneously subjected to the superposition of random noise and impulse interference, resulting in a significant decrease in the observation performance of the traditional NLESO-SMC and the appearance of high-frequency chattering in the output. In contrast, the WQ-NLESO-STSMC balances the learning contributions of the two Q tables through a weighted strategy of the dual-Q algorithm, making the gain adjustment more stable. Combined with the continuous control structure of the super-helical sliding mode law, it effectively suppresses the disturbance amplification effect, reducing the trajectory error variance by about 50% compared with the benchmark method.
[0141] In summary, WQ-NLESO-STSMC demonstrates stronger robustness and adaptability compared to traditional NLESO-SMC under all four types of disturbances. Its self-learning gain adjustment mechanism enhances the observer's response to rapidly changing disturbances, while its... The disturbance compensation term ensures the smoothness and robustness of the control law, thereby achieving high-precision, low-chatter trajectory tracking control.
[0142] The comparison results under different perturbation conditions in the table above can be further combined Figure 4 The trajectory error curves were analyzed. It can be seen that under dynamic disturbances such as irregular noise and load changes, the error curve of NLESO-SMC fluctuates significantly, exhibiting multiple peaks, with the maximum instantaneous error approaching ±0.18m. In contrast, the error change of WQ-NLESO-STSMC is more stable, with the peak value consistently controlled within ±0.03m, and it can achieve rapid error convergence within a short time. Especially under complex disturbance scenarios, the error curve of traditional NLESO-SMC exhibits significant hysteresis and oscillation, while WQ-NLESO-STSMC can recover stability in about 0.5s, with the trajectory error gradually converging to near zero, indicating that it has stronger disturbance suppression and recovery capabilities under complex nonlinear conditions.
[0143] As can be seen from the line trend, the error curve of WQ-NLESO-STSMC is smoother throughout the tracking process, with almost no high-frequency fluctuations, indicating continuous control input and stable system dynamic response. The key lies in the method's adaptive adjustment of the observer gain parameters through a weighted double-Q learning mechanism. This makes the disturbance estimation More precise, enabling real-time compensation in the control law and significantly reducing the impact of transient disturbances on the system output. In contrast, NLESO-SMC, with its fixed observation gain, cannot adjust for changes in disturbance intensity, leading to amplified errors and slower convergence during the noise enhancement phase.
[0144] comprehensive Figure 4The trajectory error performance, as shown in Table 1, indicates that the WQ-NLESO-STSMC exhibits stronger robustness and convergence characteristics when dealing with various types of disturbances. This method not only significantly reduces steady-state error and peak deviation but also improves the smoothness during dynamic transitions, achieving a simultaneous improvement in trajectory tracking accuracy and control stability.
[0145] Example 2
[0146] Table 3 shows the Monte Carlo average results of twenty independent MATLAB simulations with different random seeds for the NLESO-STSMC and WQ-NLESO-STSMC strategies on straight and circular trajectories.
[0147] Table 3
[0148]
[0149] Table 3
[0150] Table 4 shows the Monte Carlo average results of twenty independent MATLAB simulations with different random seeds for the NLESO-STSMC and WQ-NLESO-STSMC strategies on sinusoidal and figure-eight trajectories.
[0151] Table 4
[0152]
[0153] Combining the data in Tables 3 and 4, we can further analyze the performance differences between NLESO-STSMC and WQ-NLESO-STSMC under various disturbance conditions. Overall, WQ-NLESO-STSMC outperforms NLESO-STSMC in terms of average trajectory error, steady-state error, and overshoot under all types of disturbances, indicating its significant advantages in dynamic compensation and stability. Taking irregular noise as an example, the maximum trajectory error of NLESO-STSMC is approximately 0.124m, while that of WQ-NLESO-STSMC is reduced to 0.056m, a reduction of approximately 55%. This demonstrates that the weighted double-Q learning mechanism effectively optimizes the observer's parameter response, enabling the system to maintain higher trajectory consistency under random noise interference.
[0154] Under impact disturbance conditions, the difference between the two control strategies becomes more pronounced. The NLESO-STSMC exhibits an error fluctuation range of approximately ±0.1m, displaying significant transient jitter; while the WQ-NLESO-STSMC recovers to a stable state within about 0.4s after the impact disturbance, with peak error controlled within ±0.03m. This indicates that the gain adjustment mechanism based on dual-Q learning can sense changes in disturbance amplitude in real time and adjust the ESO gain parameters accordingly. Improving the accuracy of disturbance estimation enhances the compensation capability of the control law and suppresses short-term oscillations caused by shocks.
[0155] For load variations and complex disturbances, NLESO-STSMC exhibited significant overshoot during the response phase, with a large peak error and slow steady-state convergence. In contrast, WQ-NLESO-STSMC reduced the steady-state error by approximately 40% and shortened the response time by approximately 30% under the same conditions. This is mainly attributed to the dynamic optimization of ESO parameters by weighted double-Q learning, which enables the system to adaptively adjust the observer bandwidth under different loads and complex external disturbances, maintaining synchronization between the state estimate and the actual system changes.
[0156] The comprehensive comparison results show that WQ-NLESO-STSMC not only exhibits stronger stability and suppression capabilities under random noise and transient shock disturbances, but also demonstrates faster recovery speed and smaller steady-state error in scenarios with sudden load changes and complex disturbances. Overall, this method achieves dual optimization of disturbance estimation accuracy and control smoothness by introducing weighted double-Q learning into the gain adjustment of the nonlinear ESO, thereby maintaining superior robust control performance under complex nonlinear conditions.
[0157] Combination Figure 5 The position error piecewise linear curves further verify the performance differences of different control strategies in Table 2. As can be seen from the figure, both the NLESO-STSMC and WQ-NLESO-STSMC control methods can achieve error convergence under the same trajectory tracking task, but there are significant differences in steady-state accuracy and response speed. The NLESO-STSMC has a peak error of approximately 0.10m at the initial stage of system startup, and the error convergence time is close to 1.0s; while the maximum error of the WQ-NLESO-STSMC is only about 0.04m, and it can quickly converge to below 0.01m within about 0.6s, reducing the peak error by about 60% and shortening the convergence time by about 40%.
[0158] After the system enters the steady-state phase, the error curve of NLESO-STSMC still exhibits slight fluctuations, with its steady-state error remaining within ±0.02m, indicating that its suppression effect on small disturbances and noise is limited. In contrast, the WQ-NLESO-STSMC curve is smoother, with the error stabilizing within ±0.005m, indicating that the weighted dual-Q learning adjustment mechanism optimizes the gain parameters of the extended state observer. It can dynamically adjust based on real-time errors, enabling the system to maintain high estimation accuracy and control stability under different operating conditions. Simultaneously, the control law introduces... The disturbance estimation compensation term further weakens high-frequency switching jitter, making the position error change more continuous.
[0159] comprehensive Figure 5The trajectory error variation compared with the data in Table 2 shows that the WQ-NLESO-STSMC algorithm exhibits better robustness and convergence performance under various perturbation conditions. This method utilizes weighted double-Q learning to achieve adaptive adjustment of the observer gain, improving the estimation accuracy of time-varying perturbations; combined with... The disturbance-compensated superspiral sliding mode control makes the system more stable and accurate in dynamic response. Overall results show that this strategy outperforms traditional control methods in reducing peak error and improving trajectory tracking stability, and can maintain stable and reliable tracking performance in complex environments.
[0160] Example 3: This example proposes a high-order robust control system for a mobile robot combined with a weighted double-Q learning algorithm. The system includes:
[0161] The MATLAB / Simulink simulation platform is configured to perform the following process: for intelligent vehicle motion scenarios with complex uncertainties such as time-varying friction, variable slope, random external disturbances and sensor noise.
[0162] The weighted double-Q learning adaptive nonlinear extended state observer module is configured to perform the following process: employing a nonlinear error feedback function. In addition, a variable bandwidth structure is introduced, along with a weighted double-Q learning algorithm to adaptively adjust the observer gain parameters based on changes in system state error and disturbance. To achieve rapid and accurate estimation and dynamic compensation of disturbances under different working conditions, and to realize an adaptive and stable trajectory tracking strategy for intelligent vehicles;
[0163] The reward function construction module is configured to perform the following process: In response to trajectory tracking errors, disturbance estimation errors and stationarity constraints caused by complex environmental changes in the mobile robot system, the intelligent vehicle autonomously learns the optimal trajectory tracking under different disturbance conditions through continuous policy exploration and parameter updates.
[0164] The Weighted Double Q-Learning Adaptive Nonlinear Extended State Observer (WDQ-A-NLESO) module is configured to perform the following process: generating low-level control inputs based on an improved superspiral sliding mode control law, with the control law taking the form of... Control law based on sliding surface Based on this, the sliding mode parameters are updated online adaptively through the double Q learning algorithm, and the control law is corrected in real time based on the latest disturbance estimation results of the observer, so as to achieve smooth and continuous high-precision trajectory tracking and effectively suppress jitter.
[0165] Multiple operating condition simulation modules are configured to execute the following process: by establishing multiple operating condition simulation comparison schemes, simulation experiments are conducted to verify the results, and comparative analysis is performed with the traditional sliding mode control method based on extended state observers in the simulation environment.
[0166] Example 4: This example proposes an electronic system, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method steps of the present invention.
[0167] Example 5: This example proposes a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the method described in this invention, which will not be repeated here.
[0168] Example 6: This example proposes a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps of the method described in this invention, which will not be repeated here.
[0169] It should be noted that the processing flow of embodiments 3-6 corresponds to the specific steps of the method provided in embodiment 1 of the present invention, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the method provided in embodiment 1 of the present invention.
[0170] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0171] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A high-order robust control method for a mobile robot combining a weighted double-Q learning algorithm, characterized in that, Includes the following steps: Step 1: Construct a motion scenario for the intelligent vehicle in the MATLAB / Simulink simulation platform, which includes complex uncertainties such as time-varying friction, variable slope, random external disturbances, and sensor noise. Step 2: Design a weighted double-Q learning adaptive nonlinear extended state observer; Step 3: Based on Step 2, a reward function is constructed to address the trajectory tracking error, disturbance estimation error, and stationarity constraints caused by complex environmental changes in the mobile robot system. Through continuous strategy exploration and parameter updates, the intelligent vehicle autonomously learns the optimal trajectory tracking under different disturbance conditions. In step 3, the reward function structure consists of an error evaluation unit, a disturbance suppression unit, and a gain balancing unit, used to comprehensively evaluate the control performance and observation accuracy of the system in each interaction. The reward function is defined as follows: In the formula, For instant rewards, For linear velocity error, For the rate of change of speed error, For the short-time variance of the total disturbance, The adaptive gain of the observer at the current moment. Its initial reference value; These are the respective weighting coefficients, used to balance the relative impacts of error accuracy, response speed, disturbance suppression, and gain adjustment amplitude; Step 4: Design an improved superspiral sliding mode control law based on a weighted double-Q learning adaptive nonlinear extended state observer to generate low-level control inputs. The control law form is as follows: Control law based on sliding surface Based on this, the sliding mode parameters are updated online adaptively through the double Q learning algorithm, and the control law is corrected in real time based on the latest disturbance estimation results of the observer, so as to achieve smooth and continuous high-precision trajectory tracking and effectively suppress jitter. In step 4, the final control law consists of an equivalent control term and a super-helical sliding mode term, which is used to achieve stable trajectory tracking control under uncertain disturbances. Let the sliding surface be Combining the error system and the observer output, the equivalent control term is obtained: ; In the formula, Responsible for offsetting known terms and estimated disturbances in the system. The equivalent control components related to the system error state, the computable part of the error system. This is the control channel gain matrix, used to describe the sensitivity of the control signal to changes in system error. and This is the proportionality coefficient. This is an estimate of the interference. To suppress sliding mode chattering and improve control continuity, a super-spiral sliding mode correction term is introduced: ; In the formula, To ensure rapid error convergence and smooth control, For sliding surface, For sliding mode parameters, For internal correction variables; This leads to the final control law: ; Step 5: Establish multiple sets of working condition simulation comparison schemes, conduct simulation experiments to verify the results, and conduct comparative analysis with the traditional sliding mode control method based on extended state observer in the simulation environment.
2. The high-order robust control method for mobile robots combined with a weighted double-Q learning algorithm according to claim 1, characterized in that, In step 1, the simulation scenario includes several training condition sets, which include at least 5 typical conditions: flat road surface, low friction road surface, high friction road surface, uphill and downhill. New combined disturbance scenarios are added to the test set. The sources of disturbance in the simulation include sinusoidal time-varying disturbance, random impulse disturbance and Gaussian noise superposition.
3. The high-order robust control method for mobile robots combined with the weighted double-Q learning algorithm according to claim 1, characterized in that, In step 2, the weighted double-Q learning adaptive nonlinear extended state observer consists of three parts: the master observer module, the extended state compensation channel, and the adaptive gain adjustment unit. The master observer module takes the system input and output signals as its core and uses a nonlinear feedback structure to estimate the state variables of the controlled object; the extended compensation channel is used to identify and compensate for external disturbances and unmodeled dynamics in the system. The adaptive adjustment unit dynamically optimizes the observer gain based on the dual-Q learning strategy to achieve a balance between observation accuracy and response speed. An error system matching the observer structure is established. During actual operation, the robot's attitude state is determined by its position and heading angle. The desired trajectory is mapped onto the robot's body coordinate system, and the attitude error vector is defined as follows: ; In the formula, This represents the linear position error along the direction of travel. This indicates the lateral offset error. This refers to the heading angle error; By differentiating the error equation from the kinematic relationships and combining it with the system's geometric constraints and nonholonomic constraints, the dynamic error model is obtained. ; In the formula, and These represent the equivalent disturbance coefficients in the angular velocity channel and the linear velocity channel, respectively. , , Each constitutes a disturbance term ; For the desired angular velocity, For the desired linear velocity, , These are the car's current actual angular velocity and actual linear velocity, respectively. , , Each constitutes a known term ; In the error dynamic equation These are the computable known terms of the system, which are jointly determined by the current error and the control input. Since these terms are obtained in real time during trajectory tracking, they serve as known inputs for the extended state observer to help estimate the disturbance components in the system that cannot be directly measured. A weighted double-Q learning adaptive nonlinear extended state observer is constructed, and the system state and total disturbance are estimated in parallel by introducing nonlinear feedback and adaptive gain adjustment mechanisms. The observer employs a dual-channel design, establishing nonlinear extended observation equations for the linear velocity and angular velocity subsystems respectively. In the linear velocity channel, an error is assumed. Its observer form is: ; In the formula, This represents the actual position along the x-axis. These are the estimated position, estimated linear velocity, and estimated total system disturbance of the vehicle along the x-axis, respectively. These are the adaptive gain parameters of the observer; It is a nonlinear continuous function. It is a non-linear exponent. This is the bandwidth adjustment factor; The angular velocity channel structure is the same as itss: ; In the formula, These are the estimated angular velocity of the trolley, the derivative of the estimated angular velocity, and the estimated total disturbance of the system, respectively. These are the adaptive gain parameters of the observer; The dual-Q module acts as an intermediate adjustment stage, using feedback error signals to adjust the gain parameters. Lightweight online adjustments are made, and the structure enables the observer to form a three-layer system consisting of a state estimation layer, a disturbance compensation layer, and a gain self-tuning layer. This achieves the organic integration of multi-channel coordination and dynamic parameter correction, thereby ensuring the stability and consistency of the overall observer structure under different operating conditions.
4. The high-order robust control method for mobile robots combined with the weighted double-Q learning algorithm according to claim 3, characterized in that, In step 2, the observer gain adjustment process of the intelligent vehicle is modeled as a Markov decision process. state space Composed of observable multidimensional information during robot operation, including observation error amplitude, error change rate, short-time variance, and noise estimation, it reflects the instantaneous dynamic state of the system and its action space. Defined as three discrete adjustment operations for the observer gain parameter in the extended state: "increase gain", "keep it unchanged", and "decrease gain", the double-Q learning module achieves the adjustment of the observer gain by selecting the corresponding action in different states. The adaptive updates are designed to balance the estimation accuracy and dynamic response performance of the system under different operating conditions.
5. The high-order robust control method for mobile robots combined with the weighted double-Q learning algorithm according to claim 4, characterized in that, In step 3, the double-Q algorithm includes two independent Q-value functions, an action selection unit, and a parameter update module; The two sets of Q-value functions are denoted as follows: and , where the state set Composed of key state variables obtained by the mobile robot during operation, including position error, attitude error, disturbance estimate, observation gain parameter and its rate of change, used to reflect the current working state of the system; action set Defined as three discrete adjustment operations for the gain of a nonlinear extended state observer: "increase gain", "keep it unchanged", and "decrease gain". The action decision unit selects actions based on a weighted Q-value strategy, and its comprehensive evaluation function is defined as follows: ; In the formula, These are weighting coefficients. It belongs to the set of states. It belongs to the action set; This unit adopts - Greedy strategy for action selection, i.e., based on probability Conduct random exploration, with probability Select the one with the largest current value. The optimal action is determined by an alternating update mechanism in the weighted update module, which enables the two sets of Q-tables to achieve collaborative convergence while learning independently. The update form is as follows: ; In the formula, , For two sets of independent Q-value functions, , These represent the maximum value estimates of all possible actions for the two sets of Q-tables in the new state. For learning rate, As a discount factor, For instant rewards, These are weighting coefficients. The state before the transfer. The new state after the state transition. The current state The selected action In the state The selected action; After updating the double Q-value function, the algorithm selects the appropriate gain adjustment action based on the current state to achieve adaptive adjustment of the observer parameters. (Action set) Composed of discrete gain adjustment factors, used to dynamically correct the gain parameters of the extended state observer. Let the action space be: ; In the formula This corresponds to different amplitudes of gain variation; After the action is executed, the observer gain is updated in the following form: ; In the formula, For parameter-constrained operators, The adaptive gain of the observer at the current moment. and These represent the minimum and maximum values of the gain, respectively. As a reference initial value, The current state The selected action Indicates the magnitude change of gain; Through this mechanism, the dual-Q algorithm achieves discretization and adaptive adjustment of the ESO gain while maintaining convergence stability, enabling the observer to have good disturbance estimation accuracy and dynamic response capability under different operating conditions.
6. A high-order robust control system for a mobile robot combined with a weighted double-Q learning algorithm, characterized in that, The system comprising the steps of applying the method of claim 1, wherein the system includes: The MATLAB / Simulink simulation platform is configured to perform the following process: for intelligent vehicle motion scenarios with complex uncertainties such as time-varying friction, variable slope, random external disturbances and sensor noise. The weighted double-Q learning adaptive nonlinear extended state observer module is configured to perform the following process: adopting a nonlinear error feedback function and a variable bandwidth structure, while introducing a weighted double-Q learning algorithm, the observer gain parameter is adaptively adjusted according to the changes in system state error and disturbance, so as to achieve fast and accurate estimation and dynamic compensation of disturbance under different operating conditions, and realize the intelligent vehicle's adaptive and stable trajectory tracking strategy. The reward function construction module is configured to perform the following process: In response to trajectory tracking errors, disturbance estimation errors and stationarity constraints caused by complex environmental changes in the mobile robot system, the intelligent vehicle autonomously learns the optimal trajectory tracking under different disturbance conditions through continuous policy exploration and parameter updates. A weighted double-Q learning adaptive nonlinear extended state observer module is configured to perform the following process: designing a low-level control input based on an improved superspiral sliding mode control law, wherein the control law is of the form of... Control law based on sliding surface Based on this, the sliding mode parameters are updated online adaptively through the double Q learning algorithm, and the control law is corrected in real time based on the latest disturbance estimation results of the observer, so as to achieve smooth and continuous high-precision trajectory tracking and effectively suppress jitter. Multiple operating condition simulation modules are configured to execute the following process: by establishing multiple operating condition simulation comparison schemes, simulation experiments are conducted to verify the results, and comparative analysis is performed with the traditional sliding mode control method based on extended state observers in the simulation environment.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed, it implements the steps of the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is configured to implement the steps of the method according to any one of claims 1 to 5 when invoked by a processor.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Global fixed time type backstepping second-order sliding mode control method for industrial hybrid robot
CN118131623A
Mobile robot tracking control method based on multi-target point information fusion
CN118131628A