Trajectory Tracking Method for Variable-Load Mobile Manipulator Based on Reinforcement Learning and Model Predictive Control

By combining reinforcement learning and model prediction control, a mobile robotic arm system model is established and control input is optimized, the trajectory tracking problem under variable load conditions is solved, achieving high accuracy and stability improvement.

CN119748459BActive Publication Date: 2025-07-25NANJING TECH UNIV

Patent Information

Application Number
CN202510192115.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-25
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively control the tracking of the trajectory of the mobile robot under variable load conditions, and the traditional method lacks control accuracy and stability when dealing with nonlinear and time-varying loads.

Method used

Combining reinforcement learning and model prediction control, by establishing a mobile robot arm system model, using the Lagrangian equation to calculate kinetic energy, potential energy and generalized force, building a prediction model for model prediction control, and combining the Q-learning algorithm of reinforcement learning, dynamically adjusting the weight coefficients to achieve coordinated optimization of control inputs.

Benefits of technology

Implement high-precision trajectory tracking under variable load conditions, improve the operating efficiency and stability of the system, and significantly improve control accuracy and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119748459B_ABST
    Figure CN119748459B_ABST
Patent Text Reader

Abstract

The present invention discloses a trajectory tracking method for a variable-load mobile manipulator based on reinforcement learning and model predictive control. First, a system model of the mobile manipulator established based on Lagrange's equation is obtained, and the dynamic equation is derived through the system kinetic energy, potential energy, and generalized force formulas. The model predictive control (MPC) module is based on this model, uses the discrete state equation to predict the future state, and obtains the control input by optimizing the cost function. The reinforcement learning (RL) module defines a specific state space, action space, and reward function, and trains the agent using the Q-learning algorithm. In the collaborative control stage, the weight coefficient is dynamically adjusted according to the system state, and the control inputs of MPC and RL are fused. The present invention effectively solves the problem of trajectory tracking of variable-load mobile manipulators, significantly improves the trajectory tracking accuracy, enhances the system adaptability and robustness, and has broad application prospects in the fields of industrial manufacturing, logistics warehousing, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to intelligent robot control technology, and specifically relates to a trajectory tracking method for a variable-load mobile manipulator based on reinforcement learning and model predictive control. Background Art

[0002] As an intelligent robot system that combines the mobility of a mobile platform and the operation flexibility of a manipulator, the mobile manipulator plays an increasingly important role in modern production and life. In industrial manufacturing scenarios, it can flexibly shuttle in the production line, carry parts of different specifications and weights, and achieve precise assembly; in the field of logistics and warehousing, it can complete rapid sorting and handling of goods, greatly improving the efficiency of warehousing operations; in rescue and disaster relief situations, it can enter dangerous areas to perform key tasks such as material transportation, search and rescue.

[0003] However, in the actual operation process, the mobile manipulator often faces a variable-load working environment, which poses extremely high requirements on its control system. Traditional control methods, such as PID controllers and fuzzy controllers, have many limitations in dealing with such complex situations. The PID controller relies on accurate system model parameters. For the nonlinear and time-varying mobile manipulator system, it is difficult to adjust the parameters in real time to adapt to the influence brought by variable loads, resulting in a significant reduction in control accuracy. Although the fuzzy controller can handle uncertainties to a certain extent, it lacks in-depth analysis of the system's dynamic characteristics. When the load changes rapidly, it cannot make control decisions in a timely and accurate manner, challenging the stability of the system.

[0004] In recent years, the control strategy combining reinforcement learning and model predictive control (MPC) has become a research hotspot. Many researchers have carried out a large number of explorations around this. For example, some research has effectively improved the motion accuracy of the mobile manipulator under variable-load conditions by introducing a new calibration device and using its unique structure and sensor configuration, but this method may increase the complexity and cost of the system. There is also research that combines the BP neural network with other control methods to approximate the complex nonlinear relationship of the system through the powerful learning ability of the neural network, thereby improving the control accuracy. However, the training process of the neural network is usually time-consuming and prone to falling into local optima, and its real-time performance and reliability in practical applications need to be further improved. Another research designed a neural network control method to identify the variable-load situation at the end of a space robot, but its application scenario is relatively specific, and for a general mobile manipulator system, a large amount of adjustment and optimization may be required. Summary of the Invention

[0005] Objective of the Invention: The objective of the present invention is to solve the deficiencies existing in the prior art and provide a trajectory tracking method for a variable-load mobile manipulator based on reinforcement learning and model predictive control. By giving full play to the synergistic advantages of reinforcement learning and model predictive control, high-precision trajectory tracking of the mobile manipulator under variable-load working conditions is achieved, and the operating efficiency and stability of the system are improved.

[0006] Technical Solution: A trajectory tracking method for a variable-load mobile manipulator based on reinforcement learning and model predictive control of the present invention includes the following steps:

[0007] Step 1: Obtain the established mobile manipulator system model, which is established based on Lagrange's equation and includes the calculation of system kinetic energy, potential energy, and generalized forces, as well as the dynamic equation;

[0008] The formula for system kinetic energy is: ;

[0009] In the above formula, is the mass of the mobile manipulator's chassis, is the velocity of the centroid of the mobile manipulator's chassis, is the rotational inertia of the th joint of the mobile manipulator, is the angular velocity of the th joint of the mobile manipulator, is the mass of the th joint of the mobile manipulator, is the velocity of the centroid of the th joint of the mobile manipulator;

[0010] The formula for system potential energy is: ;

[0011] In the above formula, is the acceleration due to gravity, is the height of the centroid of the th joint relative to the reference plane;

[0012] The generalized forces include the generalized frictional force (including but not limited to the friction with the ground) and the generalized gravitational force ;

[0013] Furthermore, the dynamic equation is obtained as: ;

[0014] In the above formula, is the mass matrix, is the Coriolis force and centripetal force matrix, is the gravitational vector;

[0015] Step 2: Based on the mobile manipulator system model in Step 1, construct a prediction model for model predictive control. Let the discrete state equation of the system at the sampling time k+1 be ;

[0016] Set the prediction horizon to , and at each sampling time According to the current state Predict the system state at the next time instants. Define the cost function , and solve for the control input sequence that minimizes the cost function using quadratic programming algorithm, and apply the first control input at the current time to the system; refers to the control input, that is, the rotational speed control quantity and the driving torque of the manipulator joints; the function of the prediction model is state prediction, optimizing the control input, and providing a decision-making basis;

[0017] is the error vector between the system state and the reference trajectory at time , , , are the weight matrices;

[0018] Step 3: Determine the state space , action space , reward function of reinforcement learning, and use the Q-learning algorithm for training. The Q-value function update formula is:

[0019] ;

[0020] where, is the current pose of the mobile manipulator, is the joint angle of the manipulator, is the relative position between the end effector and the target object, is the system load information, is the system speed information; the rotational speeds of the left and right wheels , and are the rotational speeds of the left and right wheels, the control inputs of the manipulator joints (including joint angles, joint angular velocities, compensation information), respectively; , , , , are the weight coefficients adjusted according to the system performance requirements, is the maximum acceptable distance of the target object, is the speed error, is the angular error, is the tipping detection variable;

[0021] Step 4, Combine the control input of the model predictive control and the action selected by the reinforcement learning using the combination formula to generate the final control input ;

[0022] where is the weight coefficient, and its value range is , and it is dynamically adjusted according to the current state of the system and the task requirements.

[0023] Furthermore, in the system discrete state equation of Step 2, the current state includes the cart position coordinates , the heading angle , the angles of each joint of the robotic arm and their speed information, including the rotational speed control amount of the cart drive wheels and the driving torque of the robotic arm joints, is the error vector between the system state and the reference trajectory at time , , , are the corresponding weight matrices.

[0024] Furthermore, the specific training method in Step 3 is as follows:

[0025] Step 301, Initialize the Q-value table: At the beginning of training, initialize a Q-value table, which is used to store the Q-values of all state-action pairs; initially, all Q-values can be set to 0 or a random small value;

[0026] For the discrete state space and the action space , the Q-value table can be represented as a two-dimensional array , where , ;

[0027] At the same time, set the hyperparameters: the learning rate and the discount factor ; The learning rate controls the step size when updating the Q-value each time, and its value range is usually between 0 and 1. The learning rate determines the influence degree of the new experience information on the old Q-value; the discount factor is used to balance the importance of the current reward and the future reward, and its value range is also between 0 and 1. When is close to 0, the agent pays more attention to the current reward; when When approaching 1, the agent pays more attention to future rewards;

[0028] Step 302: At the beginning of each training episode, the agent is in the initial state of the environment ; In each state, the agent selects an action based on -greedy policy to execute; -greedy policy selects a random action with probability for exploring new actions, and selects the action with the maximum current Q-value with probability for exploiting existing experience; As the training progresses, the value of usually gradually decreases, enabling the agent to explore more in the early stage of training and exploit existing experience more in the later stage of training;

[0029] Step 303: After the agent selects an action in the state , it applies the action to the environment, and the environment transfers to the next state according to the execution result of the action, and returns an immediate reward ; This reward is the feedback of the environment to the agent's action, which reflects the quality of the action;

[0030] Step 304: According to the update rule of Q-learning, it is necessary to calculate the TD (Temporal Difference) error, that is, the difference between the Q-value of the current state-action pair and the target Q-value;

[0031] The target Q-value is calculated based on the Bellman optimal equation, and the formula is:

[0032] ;

[0033] Among them, is the immediate reward obtained by taking the action in the current state, is the discount factor, is the maximum value among the Q-values of all possible actions in the next state;

[0034] The TD error is used to update the Q-value, and the update is: That is ;

[0035] Among them, is the learning rate, which controls the update step size of the Q-value; By continuously updating the Q-value, the agent gradually learns the optimal value of each state-action pair;

[0036] Step 305: Update the current state to the next state, that is , then determine whether the current state is a termination state; if it is a termination state, the training round ends and the next training round begins; if it is not a termination state, return to step 302 to continue selecting and executing actions;

[0037] Continuously repeat the above steps 302 - 305 for multiple training rounds until the Q value converges or the preset number of training times is reached;

[0038] As the training progresses, the agent gradually learns the optimal action strategy, enabling it to select actions with the maximum long - term cumulative reward in each state.

[0039] In step 4, is dynamically adjusted according to the current state of the system and task requirements, and different values are taken under different working conditions; when the system is in a stable state and the load change is small, increase the value so that the final control input depends more on the precise planning of model predictive control; when the system state changes violently and the load changes greatly, decrease the value to give more play to the adaptive ability of reinforcement learning; specific judgment conditions and adjustment rules can be given, such as dynamically adjusting the value according to indicators such as the amplitude of load change and the magnitude of system error.

[0040] Beneficial effects: The present invention realizes high - precision trajectory tracking under variable - load working conditions, effectively enhancing the operation efficiency and stability of the system. Compared with traditional control methods such as PD and fuzzy controllers, the RL - MPC - based controller of the present invention performs better in terms of control accuracy, system stability, and real - time performance. Verified by a large number of simulations and actual tests, in the same variable - load task, the trajectory tracking error of the method of the present invention is significantly smaller, the system can adapt to load changes faster and maintain stable operation, significantly improving the working ability of the mobile manipulator in complex environments, and providing strong technical support for its wide application in multiple fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a schematic diagram of the specific trajectory tracking process of the present invention;

[0042] Figure 2 is the error of the mobile manipulator in the x - direction in the embodiment;

[0043] Figure 3 is the error of the mobile manipulator in the y - direction in the embodiment;

[0044] Figure 4 is the navigation angle error of the mobile manipulator in the embodiment Error;

[0045] Figure 5 The speed of the mobile manipulator for the embodiment Error;

[0046] Figure 6 The angular speed of the mobile manipulator for the embodiment Error;

[0047] Figure 7 The position error of the mobile manipulator for the embodiment;

[0048] Figure 8 The joint angle error of the mobile manipulator for the embodiment. Detailed implementation manners

[0049] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0050] As Figure 1 shown, a variable-load mobile manipulator trajectory tracking method based on reinforcement learning and model predictive control of the present invention includes the following steps:

[0051] Step 1, obtain the established mobile manipulator system model, the mobile manipulator system model is established based on Lagrange equation, including the calculation of system kinetic energy, potential energy and generalized force, and the dynamic equation;

[0052] The system kinetic energy formula is: ;

[0053] In the above formula, is the mass of the chassis of the mobile manipulator, is the centroid velocity of the chassis of the mobile manipulator, is the rotational inertia of the th joint of the mobile manipulator, is the angular velocity of the th joint of the mobile manipulator, is the mass of the th joint of the mobile manipulator, is the centroid velocity of the th joint of the mobile manipulator;

[0054] The system potential energy formula is: ;

[0055] In the above formula, is the gravitational acceleration, is the height of the centroid of the th joint relative to the reference plane;

[0056] The generalized force includes the frictional generalized force and the gravitational generalized force ;

[0057] Furthermore, the dynamic equation is obtained as follows: ;

[0058] In the above equation, is the mass matrix, is the Coriolis force and centripetal force matrix, is the gravity vector;

[0059] Step 2: Based on the mobile manipulator system model in Step 1, construct a predictive model for model predictive control. Let the discrete state equation of the system at the k+1 sampling time be ;

[0060] Set the prediction horizon as . At each sampling time According to the current state Predict the system state at the next time instants. Define the cost function , and solve for the control input sequence that minimizes the cost function using quadratic programming algorithm, and apply the first control input at the current time to the system;

[0061] is the error vector between the system state and the reference trajectory at time , , , are the weight matrices;

[0062] Step 3: Determine the state space , action space , reward function of reinforcement learning, and use the Q-learning algorithm for training. The Q-value function update formula is:

[0063] ;

[0064] where, is the current pose of the mobile manipulator, is the joint angle of the manipulator, is the relative position between the end effector and the target object, is the system load information, is the system speed information; the left and right wheel speeds , and are the left and right wheel speeds and the control inputs of the manipulator joints respectively;

[0065] , , , , is the corresponding weight coefficient, is the maximum acceptable distance of the target object, is the speed error, is the angle error, is the tipping detection variable;

[0066] Step 4. Fuse the control input of the model predictive control and the action selected by reinforcement learning, and use the fusion formula to generate the final control input ; is the weight coefficient.

[0067] In the system discrete state equation of Step 2, the current state includes the cart position coordinates , the heading angle , the angles of each joint of the robotic arm and their speed information, including the rotational speed control amount of the cart drive wheels and the driving torque of the robotic arm joints, is the error vector between the system state and the reference trajectory at time , , , are the corresponding weight matrices.

[0068] The specific training method of Step 3 is as follows:

[0069] Step 301. Initialize the Q-value table: At the beginning of training, initialize a Q-value table. For the discrete state space and the action space , the Q-value table is represented as a two-dimensional array , where , ;

[0070] At the same time, set the hyperparameters: the learning rate and the discount factor ;

[0071] Step 302. At the beginning of each training episode, the agent is in the initial state of the environment; in each state, the agent selects an action based on the -greedy policy to execute;

[0072] After the agent selects the action in the state, the action Applied to the environment, the environment transfers to the next state according to the execution result of the action and returns an immediate reward ;

[0073] Step 304: According to the update rule of Q-learning, calculate the TD error, that is, the difference between the Q-value of the current state-action pair and the target Q-value;

[0074] Calculate the target Q-value based on the Bellman optimal equation, and the calculation formula is:

[0075] ;

[0076] where is the immediate reward obtained by taking the action in the current state, is the discount factor, is the maximum value among the Q-values of all possible actions in the next state;

[0077] Use the TD error to update the Q-value, and update it to: That is ;

[0078] where is the learning rate, which controls the update step size of the Q-value; by continuously updating the Q-value, the agent gradually learns the optimal value of each state-action pair;

[0079] Step 305: Update the current state to the next state, that is , then determine whether the current state is a termination state; if it is a termination state, the training round ends and enters the next training round; if it is not a termination state, return to step 302 to continue selecting actions and executing;

[0080] Continuously repeat the above steps from step 302 to step 305 for multiple training rounds until the Q-value converges or reaches the preset number of training times.

[0081] The trajectory tracking method of the mobile manipulator in this embodiment includes the following steps:

[0082] Step 1: Construct a mobile manipulator system.

[0083] The mobile manipulator system of this embodiment includes a mobile platform and a manipulator. The mobile platform is of nonholonomic constraint type, has two degrees of freedom and is rear-wheel driven. A two-finger gripper is installed at the front end of the manipulator for grasping heavy objects. The weight of the platform is 50 kg, the wheel radius is 0.5 m, and the wheelbase is 1 m. The masses of the connecting rods are 3.7 kg, 8.4 kg, 2.3 kg, 1.2 kg, 1.2 kg, 0.2 kg respectively, and the lengths of the connecting rods are 0.089 m, 0.425 m, 0.392 m, 0.109 m, 0.094 m, 0.082 m respectively.

[0084] The mobile manipulator system of this embodiment has universality, and the proposed analysis and control methods are not limited to specific robot structures. When there is no external force sensor and the system faces variable loads, the system stability is vulnerable to challenges. The designed working scenario is that the mobile manipulator successively grasps objects of 200 N, 300 N, and 400 N in a predetermined running trajectory of a cube.

[0085] Step 2: Establish a mobile manipulator system model based on the Lagrange equation. Specifically as follows:

[0086] Step 201: Calculate the system kinetic energy

[0087]

[0088] Where is the mass of the chassis, is the velocity of the mass center of the chassis, is the th joint moment of inertia, is the th joint angular velocity, is the th joint mass, is the th joint mass center velocity;

[0089] Step 202: Calculate the system potential energy

[0090]

[0091] Where is the acceleration due to gravity, is the height of the mass center of the th joint relative to the reference plane;

[0092] Step 203: Calculate the generalized force: The generalized force includes generalized forces such as control input force, friction force, and gravity.

[0093] The generalized force of the friction force is ; The generalized stress of the gravity is: ;

[0094] Step 204: The dynamic equation of the mobile manipulator can be obtained through the Lagrange equation:

[0095]

[0096] where is the mass matrix, is the Coriolis and centripetal force matrix, is the gravity vector.

[0097] Step 3: Construct a prediction model for model predictive control based on the obtained accurate dynamic model of the mobile manipulator. The specific process is as follows:

[0098] Step 301: Assume the discrete state equation of the system is:

[0099]

[0100] where includes the cart position , the heading angle , the angles of each joint of the manipulator and their velocity information, comprehensively describing the state of the system at time ; includes the control quantity of the cart drive wheel speed and the drive torque of the manipulator joints, which is the control input of the system.

[0101] Step 302: Set the prediction horizon to , and at each sampling time , predict the system state at the next moments according to the current state . First, calculate the state

[0102] at the next moment from the current state and the control input at the previous moment, and then repeat this process to predict the subsequent states.

[0103]

[0104] Step 303: Define the cost function: is the error vector between the system state and the reference trajectory at time , , , are the weight matrices, used to adjust the importance of the tracking error and the control input in the cost function. Solve the control input sequence that minimizes through the quadratic programming algorithm, and apply the first control input to the system to achieve the optimization of the future state of the system.

[0105] ​Step 4: The reinforcement learning design enables the system to achieve efficient control of the mobile manipulator in a variable load environment by leveraging the intelligent decision-making advantages of reinforcement learning.

[0106] Step 401: Determination of the state space, action space, and reward function:

[0107] State space: where is the current pose of the mobile manipulator, is all the joint angles of the manipulator, is the relative position between the end effector and the target object, is the system load information, is the system speed information.

[0108] Action space: used to adjust the rotational speed of the trolley drive wheels and and the manipulator joint control input .

[0109] Reward function: , where , , , , are weight coefficients adjusted according to the system performance requirements, is the maximum acceptable distance of the target object, is the speed error, is the angle error, is the tipping detection variable.

[0110] Step 402: Q-learning algorithm training: Initialize the Q-value table: At the start of training, a Q-value table needs to be initialized. This table is used to store the Q-values for all state-action pairs. Usually, initially, all Q-values can be set to 0 or small random values. For a discrete state space and action space , the Q-value table can be represented as a two-dimensional array , where , . Some hyperparameters also need to be set, such as the learning rate and the discount factor . The learning rate controls the step size when updating the Q-value each time. Its value range is usually between 0 and 1, and it determines the influence degree of new experience information on the old Q-values. The discount factor is used to balance the importance of the current reward and future rewards. Its value range is also between 0 and 1. When is close to 0, the agent pays more attention to the current reward; when When approaching 1, the agent pays more attention to future rewards.

[0111] Step 403: At the start of each training episode, the agent is in the initial state of the environment . In each state, the agent needs to select an action to execute. The action selection strategy usually adopts -greedy strategy, which selects a random action with probability (for exploring new actions), and with probability selects the action with the largest current Q-value (for exploiting existing experience). As the training progresses, 's value usually gradually decreases, causing the agent to explore more in the early stage of training and exploit existing experience more in the later stage of training.

[0112] Step 404: After the agent selects an action in the state , it applies this action to the environment, and the environment transfers to the next state according to the execution result of the action , and returns an immediate reward ; this reward is the feedback of the environment to the agent's action, which reflects the quality of this action.

[0113] Step 405: According to the update rule of Q-learning, it is necessary to calculate the TD (Temporal Difference) error, that is, the difference between the Q-value of the current state-action pair and the target Q-value. The calculation of the target Q-value is based on the Bellman optimal equation, and the formula is: ;

[0114] where, is the immediate reward obtained by taking action in the current state, is the discount factor, is the maximum value among the Q-values of all possible actions in the next state. Use the TD error to update the Q-value, and update it to: That is where, is the learning rate, which controls the update step size of the Q-value. By continuously updating the Q-value, the agent gradually learns the optimal value of each state-action pair.

[0115] Step 406: Update the current state to the next state, that is . Determine whether the current state is a terminal state. If it is a terminal state, this training episode ends and enters the next training episode; if it is not a terminal state, return to Step 402 and continue to select actions and execute.

[0116] Step 407: Continuously repeat the above steps 402 - 405 for multiple training rounds until the Q-value converges or reaches the preset number of training times. As the training progresses, the agent gradually learns the optimal action strategy, enabling it to select the action with the maximum long-term cumulative reward in each state.

[0117] Step 5: Cooperative control strategy

[0118] Step 501: Fuse the control input of model predictive control and the action selected by reinforcement learning using the fusion formula to generate the final control input , where is the weight coefficient, and its value range is .

[0119] Step 502: Dynamic weight adjustment: is dynamically adjusted according to the current state of the system and task requirements, and will have different values under different operating conditions. Specifically as follows:

[0120] When the system is in a stable state and the load changes little, increase the value of so that the final control input depends more on the precise planning of model predictive control; when the system state changes drastically and the load changes greatly, decrease the value of to give more play to the adaptive ability of reinforcement learning. Specific judgment conditions and adjustment rules can be given, such as dynamically adjusting the value of according to indicators such as the amplitude of load change and the magnitude of system error.

[0121] As Figures 2 - 8 shows, in order to verify the effectiveness and feasibility of the variable-load mobile manipulator trajectory tracking method based on reinforcement learning and model predictive control, a comparison and simulation experiment of this control method and the traditional Fuzzy and PID methods is carried out on Matlab. All experiments are carried out on a computer equipped with 16G of memory and a 12th Gen Intel(R) Core(TM) i7-12700H.

[0122] From Figure 2 , Figure 3 and Figure 4 it can be seen that the present invention has significant advantages over the other two methods, with a smaller error amplitude. During the operation of the robot, this method can implement adjustments more quickly, compensate for errors, and thus promote the system to reach a stable state.

[0123] Analysis Figure 5 and Figure 6It can be seen that all three control methods induced vibration phenomena in the initial stage, but in the subsequent stage, their performances in terms of speed all tended to be stable. It is worth emphasizing that the present invention can reach the stable state at a faster speed and exhibits stronger robustness.

[0124] Figure 7 and Figure 8 Precisely depict the motion error situation of the entire system, including the mobile platform and the robotic arm joints, throughout the entire motion process. Obviously, under the action of the three control methods, the errors all fluctuate within a small and controllable range. However, comprehensively evaluating various indicators, the present invention generally presents better control performance and effectively provides effective support for the mobile robotic arm system to maintain stability under variable load conditions.

[0125] The present invention organically combines the intelligent decision-making ability of reinforcement learning with the precise prediction and optimization ability of model predictive control. Reinforcement learning explores and optimizes control strategies continuously according to reward feedback through the interaction between the agent and the environment, enhancing the adaptive ability of the system in complex variable load environments; model predictive control predicts future states based on the system model and determines the control input sequence by optimizing the cost function to achieve the forward-looking control of the system. The combination of the two enables the mobile robotic arm to better balance task execution and system stability under variable load conditions and achieve precise trajectory tracking.

Claims

1. A trajectory tracking method for a variable-load mobile manipulator based on reinforcement learning and model predictive control, characterized in that, It includes the following steps: Step 1: Obtain the established mobile manipulator system model. The mobile manipulator system model is established based on the Lagrangian equation, including the calculation of system kinetic energy, potential energy, and generalized force, as well as the dynamic equation; The system kinetic energy formula is as follows: In the above formula, m0 is the mass of the chassis of the mobile manipulator, v0 is the velocity of the centroid of the chassis of the mobile manipulator, I i is the moment of inertia of the i-th joint of the mobile manipulator, ω i is the angular velocity of the i-th joint of the mobile manipulator, m i is the mass of the i-th joint of the mobile manipulator, v i is the velocity of the centroid of the i-th joint of the mobile manipulator; The formula for the potential energy of the system is as follows: In the above formula, g is the acceleration due to gravity, and h i is the height of the centroid of the i-th joint relative to the reference plane; Generalized forces include the generalized frictional force and the generalized gravitational force Furthermore, the kinetic equation obtained is as follows: In the above formula, M(q) is the mass matrix, is the Coriolis force and centripetal force matrix, and G(q) is the gravity vector; Step 2: Based on the mobile manipulator system model in Step 1, construct a prediction model for model predictive control. Let the discrete state equation of the system at the k+1 sampling time be x k+1 = f d (x k , u k ); Set the prediction horizon to N, and at each sampling time k, based on the current state x k Predict the system state at the next N time instants, and define the cost function Solve for the control input sequence that minimizes the cost function J using quadratic programming algorithm And take the first control input at the current time Apply it to the system; u k Refers to the control input; e i is the error vector between the system state and the reference trajectory at time i, and Q, R, and P are weight matrices; Step 3. Determine the state space S of the reinforcement learning as S = (q, θ a , Δp, m l , v), and the action space A = (ω l , ω r , τ a ), and the reward function , Use the Q-learning algorithm for training. The Q-value function update formula is: where q = (x, y, θ) is the current pose of the mobile manipulator, and θ a is the joint angle of the manipulator, Δp is the relative position between the end effector and the target object, m l is the mass of the i-th joint, v = (v, ω) is the system velocity information; the rotational speeds ω l and ω r of the left and right wheels, and τ a are the rotational speeds of the left and right wheels and the control inputs of the manipulator joints, respectively; r1, r2, r3, r4, r5 are corresponding weight coefficients, d max is the maximum acceptable distance of the target object, e v is the speed error, e θ is the angle error, C tip is the tipping detection variable; Step 4: Fuse the control input u MPC of model predictive control and the action u RL selected by reinforcement learning, and use the fusion formula u final =βu MPC +(1 - β)u RL to generate the final control input u final ; β is the weight coefficient.

2. The trajectory tracking method for a variable-load mobile manipulator based on reinforcement learning and model predictive control according to claim 1, characterized in that The system discrete state equation x of step 2 k+1 = f d (x k , u k ), the current state x k includes the cart position coordinates (x, y), the heading angle θ, the angles of each joint of the robotic arm θ a and their speed information, u k includes the rotational speed control amount of the cart drive wheels and the driving torques of the robotic arm joints, e i is the error vector between the system state and the reference trajectory at time i, and Q, R, P are the corresponding weight matrices.

3. The trajectory tracking method for a variable-load mobile manipulator based on reinforcement learning and model predictive control according to claim 1, characterized in that The specific training method for Step 3 is: Step 301: Initialize the Q-value table. At the beginning of training, initialize a Q-value table. For the discrete state space S and action space A, the Q-value table is represented as a two-dimensional array Q(s,a), where s∈S and a∈A; At the same time, set hyperparameters: learning rate α and discount factor γ; Step 302: At the beginning of each training episode, the agent is in the initial state s0 of the environment; in each state, the agent selects an action a to execute based on the ε-greedy policy; Step 303: After the agent selects action a in the state, apply action a to the environment. The environment transfers to the next state s′ according to the execution result of the action and returns an immediate reward R(s,a); Step 304: According to the update rule of Q-learning, calculate the TD error, that is, the difference between the Q-value of the current state-action pair and the target Q-value; Calculate the target Q-value based on the Bellman optimal equation. The calculation formula is: where \(R(s, a)\) is the immediate reward obtained by taking action \(a\) in the current state, and \(\gamma\) is the discount factor, is the maximum value among the \(Q\)-values of all possible actions in the next state; Update the Q-value using the TD error, and the update is as follows: That is where α is the learning rate, which controls the update step of the Q-value; by continuously updating the Q-value, the agent gradually learns the optimal value of each state-action pair; Step 305: Update the current state to the next state, that is, s←s’, and then determine whether the current state is a terminal state; if it is a terminal state, the training episode ends and enters the next training episode; if it is not a terminal state, return to Step 302 to continue selecting actions and executing; Continuously repeat the above steps 302 to 305 for multiple training episodes until the Q-value converges or reaches the preset number of training times.

Citation Information

Patent Citations

  • Quadruped robot control method based on model predictive control optimization reinforcement learning

    CN113568422A

  • Redundant drive mechanical arm path planning method based on DSAW offline reinforcement learning algorithm

    CN118700133A

Cited By

  • Strategy gradient embedded enhancement model predictive control method

    CN121900249A