Mechanical arm tracking control method based on transferable depth increment reinforcement learning

By employing a model-free reinforcement learning method using deep incremental models and asynchronous deep value networks, the problem of control strategy transfer in a robotic arm system across different degrees of freedom was solved, achieving fast and robust tracking control while reducing computational complexity and repetitive training overhead.

CN121928547APending Publication Date: 2026-04-28HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2026-01-13
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as high computational load, weak generalization ability, and poor transferability in the optimal tracking control of robotic arms. This results in time-consuming and inefficient repetitive training of control strategies across robotic arm systems with different degrees of freedom.

Method used

A model-free reinforcement learning method based on deep incremental models and asynchronous deep value networks is adopted. An incremental model of a 1-DOF manipulator is constructed through offline learning. An incremental guidance strategy and a zero-sum game value function are designed. An adaptive dynamic programming is used to establish a transfer mechanism to achieve rapid tracking control of a high-DOF manipulator system.

Benefits of technology

It realizes the transfer of control strategies between robotic arm systems with different degrees of freedom, reduces the repetitive training process, improves the rapid tracking control capability and robustness of the robotic arm system, and reduces computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121928547A_ABST
    Figure CN121928547A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of robot intelligent control, and particularly relates to a mechanical arm tracking control method based on transferable depth incremental reinforcement learning, which effectively solves the nonlinear optimal tracking control problem of a mechanical arm in a model-free mode by constructing a transferable incremental reinforcement learning framework. The method comprises the following steps: firstly, constructing a 1-degree-of-freedom depth increment model by utilizing one-step forward data offline learning, and providing universal dynamic representation for a mechanical arm which is difficult to accurately model; furthermore, an asynchronous depth value network is designed for the one-degree-of-freedom mechanical arm, stable and rapid convergence value function approximation is achieved through a separation base layer and an adaptive layer, and a cross-mechanical-arm migration mechanism is established, so that a pre-training model and a network base layer on the one-degree-of-freedom mechanical arm can be directly migrated to a high-degree-of-freedom mechanical arm subsystem; and the system difference is compensated only by updating the adaptive layer online, so that the repeated training overhead is remarkably reduced, and meanwhile, rapid, robust and adaptive tracking control on the mechanical arms with different degrees of freedom is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot intelligent control technology, specifically relating to a robotic arm tracking control method based on transferable deep incremental reinforcement learning. Background Technology

[0002] In the field of robot control, nonlinear optimal tracking control plays a crucial role in the high-performance operation of complex robotic arms. However, its practical application is often limited due to the lack of accurate models in uncertain or unstructured environments. To address this limitation, reinforcement learning has become a powerful tool for solving nonlinear optimal tracking control problems without the need for explicit model knowledge. However, current foundations generally suffer from limited generalization ability and weak transferability. Specifically, control strategies trained for a specific degree of freedom (DOF), such as a 1-DOF robotic arm, are difficult to directly apply to robotic arms with different DEFs, such as a 3-DOF robotic arm. Whenever faced with a new robotic arm configuration or task, it is usually necessary to retrain the strategy from scratch, consuming significant computational resources and time, resulting in low learning efficiency and severely restricting the deployment of this technology in diverse and practical robotic applications. Therefore, researching a nonlinear optimal tracking control method with good transferability to improve the reusability of reinforcement learning control strategies across different robotic arm configurations is particularly important.

[0003] For example, Chinese Patent Publication No. CN119501947A discloses a fixed-time optimal control method for a single-link manipulator. By constructing a state-space model and introducing an active disturbance rejection control strategy, combined with an extended state observer, backstepping method, and reinforcement learning algorithm, adaptive optimal control and stable tracking of the single-link manipulator are achieved. This method shows some effectiveness in single-degree-of-freedom manipulator control, but its design is primarily geared towards single-link systems and does not involve extensions or applications to high-dimensional, multi-degree-of-freedom manipulator systems.

[0004] Furthermore, Chinese patent CN119057776B proposes an optimal control method for robotic arm trajectory tracking based on an all-drive system approach. This method achieves optimal control of the robotic arm system by introducing a high-order disturbance observer and has been verified in a multi-degree-of-freedom robotic arm. However, this method does not employ a control strategy transfer mechanism, requiring redesign for different robotic arm systems, resulting in high computational complexity.

[0005] In summary, existing technologies have made some progress in optimal tracking control of robotic arms, but they still suffer from problems such as high computational complexity, weak generalization ability, and poor transferability when dealing with new robotic arms. Therefore, this invention provides a robotic arm tracking control method based on transferable deep incremental reinforcement learning. Summary of the Invention

[0006] The purpose of this invention is to provide a robotic arm tracking control method based on transferable deep incremental reinforcement learning, which solves the nonlinear optimal tracking control problem in a model-free manner. By sharing and reusing deep incremental models and deep value networks among robotic arm subsystems with different degrees of freedom, the control strategy can be effectively transferred between different degree-of-freedom systems, thereby reducing the repetitive training process and improving the rapid tracking control capability of the robotic arm system.

[0007] The specific technical solution adopted by this invention is as follows: This invention specifically relates to a model-free reinforcement learning control method based on a deep incremental model and an asynchronous deep value network, used to achieve control strategy transfer and fast tracking control between different robotic arm systems; specifically including the following steps: S1: Establish a state-space model of the dynamically unknown robotic arm system. Using a time delay estimation method, and based on one-step backward data, train the incremental gain matrix of the 1-DOF robotic arm through offline deep learning. Construct a deep incremental model; S2: Designing an incremental guidance strategy based on a depth increment model for a 1-DOF robotic arm. This simplifies the learning process to reduce complexity and improves control performance. S3: Establish an incremental residual strategy for a 1-DOF robotic arm and incremental model error The zero-sum game problem, designing the value function of the zero-sum game. The corresponding Hamilton-Jacobi-Isax equation is derived, and the optimal incremental residual strategy is solved. And worst incremental model error Strategy; S4: Utilize adaptive dynamic programming to establish an asynchronous deep value network and optimize the incremental residual strategy. And worst incremental model error The strategy is approximated, and the network weight update rate is designed; S5: Utilizing the dynamic similarity between the decoupled subsystem of the high-degree-of-freedom robotic arm and the 1-degree-of-freedom robotic arm, a transfer mechanism for the depth increment model and depth value network is established to achieve tracking control of the high-degree-of-freedom robotic arm system.

[0008] The technical effects achieved by this invention are as follows: This invention effectively solves the nonlinear optimal tracking control problem of robotic arms in a model-free manner by constructing an incremental reinforcement learning framework; it utilizes one-step forward data to learn a deep incremental model offline, providing a general dynamic representation for robotic arms that are difficult to model precisely; further, it designs an asynchronous deep value network, which achieves stable and fast convergence of value function approximation by separating the base layer and the adaptive layer, and establishes a cross-robotic arm transfer mechanism, enabling the pre-trained model and the network base layer to be directly transferred to the high-degree-of-freedom robotic arm subsystem, compensating for system differences only by updating the adaptive layer online. Thus, it achieves fast, robust, and adaptive tracking control of robotic arms with different degrees of freedom while significantly reducing the overhead of repeated training. Attached Figure Description

[0009] Figure 1 This is a flowchart of a model-free reinforcement learning control method based on a deep incremental model and an asynchronous deep value network, according to the present invention. Figure 2 The simulation results of this invention on a 1-DOF machine show that the proposed deep incremental reinforcement learning tracking control method can enable the robotic arm to track the reference trajectory. Figure 3 This is a schematic diagram of the simulation results of transferring the depth increment model and depth value network to a 3-DOF robotic arm according to the present invention. Detailed Implementation

[0010] To make the objectives and advantages of this invention clearer, the invention will be specifically described below with reference to embodiments. It should be understood that the following text is merely used to describe one or more specific embodiments of the invention and does not strictly limit the scope of protection specifically claimed by the invention.

[0011] Example 1: As Figure 1 As shown, a model-free reinforcement learning control method based on a deep incremental model and an asynchronous deep value network is presented, comprising the following steps: S1: Establish a state-space model of the dynamically unknown robotic arm system. Using a time delay estimation method, and based on one-step backward data, train the incremental gain matrix of the 1-DOF robotic arm through offline deep learning. Construct a deep incremental model; The specific process is as follows: The dynamic behavior of the robotic arm is described using the following continuous-time nonlinear system: (1) in And each state represents a position-velocity representation as follows: Control input is represented as ; Represents the nonlinear term of the system; This represents the system's control gain. Assume... and All are unknown and locally Lipschitz continuous. The desired trajectory of the system is set as follows: (2) in It is a Lipschitz continuous function.

[0012] An incremental model of a 1-DOF robotic arm is constructed using historical data, and a state-dependent gain matrix is ​​introduced. Multiply both sides of formula (1) by the Moore-Penrose pseudo-inverse. , (3) in, It includes all the unmodeled dynamics in equation (1). At sufficiently high frequencies, the unknown... The data can be estimated using a one-step forward approach: (4) in , , This represents the sampling time. This leads to the incremental model expression: (5) in This represents the model estimation error. This represents an incremental control strategy.

[0013] Establish an offline deep learning network by collecting the system's input-output dataset. To train and obtain the state-related incremental gain weight matrix : (6) S2: Designing an incremental guidance strategy based on a depth increment model for a 1-DOF robotic arm. This simplifies the learning process to reduce complexity and improves control performance. The specific process is as follows: Secondly, for S2, following the residual reinforcement learning mechanism, the system's deep tracking strategy is designed as follows: (7) in This represents an incremental guidance strategy. This represents the incremental residual strategy. The tracking error of the system is defined as... Furthermore, the formula for the dynamic characteristics of incremental tracking error can be obtained: (8) Establish an incremental guidance strategy for a 1-DOF robotic arm To simplify the learning process, reduce complexity, and improve control performance, the formula is as follows: (9) in It is a defined constant matrix. Note that the explicit form of the depth tracking strategy and the incremental guidance strategy of this invention is independent. For the convenience of theoretical analysis, the form of formula (9) was chosen.

[0014] S3: Establish an incremental residual strategy for a 1-DOF robotic arm and incremental model error The zero-sum game problem, designing the value function of the zero-sum game. The corresponding Hamilton-Jacobi-Isax equation is derived, and the optimal incremental residual strategy is solved. And worst incremental model error Strategy; The specific process is as follows: To overcome the problem of poor adaptability of incremental guidance strategies in dynamic environments, an incremental residual strategy is introduced to improve tracking accuracy and robustness.

[0015] For ease of analysis, record Then formula (8) can be written as (10) Establish an incremental residual strategy and incremental model error The zero-sum game problem is used to consider the impact of model error on control performance. It is considered to minimize the number of participants, with the aim of reducing tracking error and control volume; while The value function of a zero-sum game is defined as maximizing the participant's representation of the worst-case model error. (11) in , It is a positive definite matrix. It uses positive scalar control weights. It is a predefined constant. It is an integral virtual time variable. , , Representing virtual time The tracking error, incremental residual strategy, and model error are all factors to consider.

[0016] The optimal value function satisfies the following formula: (12) The value function satisfies the following Hamilton-Jacobi-Isaacs equation: (13) The strategies that satisfy the above equation (13) correspond to potential zero-sum game saddle points, which satisfy the following inequality: (14) Saddle Point It can be obtained through the following equation: (15) Based on this, the optimal residual strategy is obtained. and worst model error The solution is: (16) (17) S4: Utilize adaptive dynamic programming to establish an asynchronous deep value network and optimize the incremental residual strategy. And worst incremental model error The strategy is approximated, and the network weight update rate is designed; The specific process is as follows: Since solving the Hamilton-Jacobi-Isax equation is a nonlinear partial differential equation, it is difficult to obtain an analytical solution. Therefore, in S4, an asynchronous deep value network is used to approximate the cost function, which has the following form: (18) in and Let represent the ideal weights and activation function of the adaptive layer, respectively. The base layer is represented as follows: ,in These are the base layer weights. Indicate its activation function, The number of floors.

[0017] Based on the approximation equation (18), the present invention can obtain the following formula: (19) in and They are and The partial derivatives. Substituting into the Hamilton-Jacobi-Isax equation, we get: (20) in This represents the residual approximation error introduced by the deep value network. Since the system dynamics (1) is Lipschitz continuous, the residual term is bounded, denoted as... Further, we can obtain (twenty one) in .

[0018] Based on this, the approximation form of (21) can be obtained as follows: (twenty two) This leads to the error form, specifically the formula: (twenty three) The weights of the base layer are updated at a relatively slow rate to ensure training stability. The specific formula for its update law is as follows: (twenty four) in This is the learning rate of the base layer. This slow-time-scale update aims to optimize and accurately estimate... This plays a crucial role in ensuring the stability of the overall training process.

[0019] The adaptive layer weights are updated at a relatively fast rate to ensure convergence. The specific formula for its update law is as follows: (25) Where the coefficient , It is an introduced scalar gain used to balance the respective contributions of real-time data and empirical data in the online learning of adaptive layer weights; It is a constant positive definite gain matrix; This indicates the number of samples in the empirical data.

[0020] Assuming this invention considers the empirical buffer matrix It is full rank, that is This assumption differs from the traditional continuous incentive condition by introducing a rank condition as an online verifiable metric for the data richness required for adaptive weight convergence.

[0021] Based on the above steps, the optimal residual strategy can be obtained. And worst model error strategy The approximate solution is: (26) (27) The final control strategy for the 1-DOF robotic arm system is as follows: (28) S5: Utilizing the dynamic similarity between the decoupled subsystem of the high-dimensional robotic arm and the 1-DOF robotic arm, a transfer mechanism for the depth increment model and depth value network is established to achieve tracking control of the high-DOF robotic arm system.

[0022] The specific process is as follows: In S5, instead of redesigning the controller of the high-dimensional robotic arm by repeating the above design process, the dynamic similarity between the decoupled subsystem and the 1-DOF robotic arm system is utilized to transfer the depth increment model and depth value network trained on the 1-DOF robotic arm to other high-dimensional robotic arms. The specific implementation strategy is as follows.

[0023] for n The robotic arm system is first decoupled into n subsystems: (29) in It is the first The state of each subsystem; It is the first Control inputs for each subsystem; Indicates the first The combined effect of unknown dynamics and coupling terms of individual subsystems; Indicates the first The control input gain matrix of each subsystem.

[0024] The depth incremental model corresponding to the subsystem is: (30) The first high-dimensional robotic arm system The base layers of the deep value network of each subsystem are directly shared by the base layer network trained on the 1-DOF robotic arm, as shown in the following formula: (31) Similarly, the first The adaptive layer weights of each subsystem are initialized to the adaptive layer weights trained on a 1-DOF robotic arm, using the following formula: (32) Furthermore, similar to (25), online data is used for updating. To adapt to environmental changes, the specific formula is as follows: (33) Therefore, the first The value function estimate of each subsystem is expressed as: (34) No. The deep incremental model of each subsystem is transferred from a deep network trained on 1 degree of freedom. Specifically, the incremental gain matrix is ​​directly shared. (35) Thus, the strategy transfer from a 1-DOF robotic arm to an n-DOF robotic arm was achieved, as shown below: (36) (37) (38) (39) Example

[0025] The stability, or feasibility, of the proposed scheme is proven using the Lyapunov function method.

[0026] For the optimal residual strategy in step S3 And worst model error strategy Under the condition that the following inequality holds (40), the tracking error can be adjusted to a small neighborhood of zero: (41) Substituting the variables into the equations, we obtain the following inequality: (42) Assumption There exists an upper bound. ,Right now Then the following inequality holds: (43) When condition (40) is met, the present invention can be obtained. (44) in Therefore, when Sometimes, . This represents the smallest eigenvalue of a symmetric real matrix. The system state of the incremental model will eventually converge to the residual set: (45) In S4, the present invention needs to derive a convergence proof regarding the tracking error and the weighting error.

[0027] Assume there exists a constant , , , , , , making , , , , , .

[0028] Assuming the user selects the base layer features and satisfy ,in It is the function reconstruction error of the base layer, and ,in It is a bounded constant.

[0029] Under the above assumptions, the basic layer error The convergence of the network has been guaranteed, therefore the convergence analysis of the weights in the asynchronous deep value network is simplified to the adaptive layer weight error. The analysis will be resolved in the following derivation.

[0030] This invention selects the following Lyapunov functions: (46) For ease of proof, let's remember... , First of all Taking the derivative, we get: (47) According to the Hamilton-Jacobi-Isax equation, we can further obtain: (48) For one of them Based on formula (27), the following inequality can be further derived from this invention: (49) in .

[0031] Similar to (49), the present invention can yield the following three inequalities: (50) (51) (52) Substituting inequalities (49)-(52) into (48) yields: (53) in,

[0032] for The derivative of this invention yields... (54) For ease of derivation, let's say... Then the first term of formula (54) can be derived as: (55) Approximately, for the second term, the following inequality holds: (56) Substituting (55) and (56) into (54) yields: (57) in

[0033] Finally obtained The final form is: (58) in It is positive definite. , Choosing parameters makes ,because It is positive definite. If it satisfies formula (59), then the Lyapunov derivative is negative: (59) Finally, the weight learning error converges to the following residual set: (60) Therefore, the tracking error and weight estimation error of a 1-DOF robotic arm are both consistent and eventually bounded.

[0034] The theoretical analysis of the tracking error and weight estimation error of each subsystem in the high-dimensional robotic arm system can be referred to the above derivation process, which will not be repeated here.

[0035] This invention was simulated and verified on the following 1-DOF robotic arm, and its system can be described as follows: (61) in , and All are unknown positive numbers. For ease of subsequent simulation, let... and Thus, the robotic arm system (61) can be converted into the same form as (1), wherein (62) The reference trajectory is selected as The sampling rate is set to 1000Hz.

[0036] according to Figure 2Simulation results on a 1-DOF machine show that the proposed deep incremental reinforcement learning tracking control method enables the robotic arm to track the reference trajectory with excellent tracking performance, and its weights eventually converge. Figure 3 Simulation results show that the depth increment model and depth value network have been transferred to a 3-DOF robotic arm. It can be seen that the tracking error converges within a small interval near 0, indicating excellent tracking performance. The corresponding weights also eventually converge. Therefore, the robotic arm tracking control method based on transferable depth increment reinforcement learning proposed in this invention can enable a robotic arm system with unknown dynamics to track a given reference trajectory.

[0037] This invention provides an incremental reinforcement learning framework that solves the tracking control problem of nonlinear systems in a model-free manner, and achieves fast tracking control of robotic arms by sharing a deep incremental model and a depth value network among different robotic arms.

[0038] This invention uses one-step forward data offline learning of a deep incremental model to provide model-agnostic representations for robotic arms with almost no modeling. The deep incremental model pre-trained on 1 degree of freedom can be directly transferred to the subsystems of a high-dimensional robotic arm.

[0039] This invention designs an asynchronous depth value network to achieve general and accurate value function approximation, and further establishes a transfer mechanism that enables the depth value network to be shared between different robotic arms. Specifically, the base layer is directly transmitted to other robotic arm subsystems, while the adaptive layer is updated online to take into account the specific differences of the robotic arms.

[0040] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.

Claims

1. A robotic arm tracking control method based on transferable deep incremental reinforcement learning, characterized in that: Includes the following steps: S1: Establish a state-space model of the dynamically unknown robotic arm system. Using a time delay estimation method, and based on one-step backward data, train the incremental gain matrix of the 1-DOF robotic arm through offline deep learning. Construct a deep incremental model; S2: Designing an incremental guidance strategy based on a depth increment model for a 1-DOF robotic arm. This simplifies the learning process to reduce complexity and improves control performance. S3: Establish an incremental residual strategy for a 1-DOF robotic arm and incremental model error The zero-sum game problem, designing the value function of the zero-sum game. The corresponding Hamilton-Jacobi-Isax equation is derived, and the optimal incremental residual strategy is solved. And worst incremental model error Strategy; S4: Utilize adaptive dynamic programming to establish an asynchronous deep value network and optimize the incremental residual strategy. And worst incremental model error The strategy is approximated, and the network weight update rate is designed; S5: Utilizing the dynamic similarity between the decoupled subsystem of the high-degree-of-freedom robotic arm and the 1-degree-of-freedom robotic arm, a transfer mechanism for the depth increment model and depth value network is established to achieve tracking control of the high-degree-of-freedom robotic arm system.

2. The robotic arm tracking control method based on transferable deep incremental reinforcement learning according to claim 1, characterized in that: The specific process of S1 is as follows: The dynamic behavior of the robotic arm is described using the following continuous-time nonlinear system: (1) in And each state represents a position-velocity representation as follows: Control input is represented as ; Represents the nonlinear term of the system; Represents the system's control gain; The system's expected trajectory is set as follows: (2) in It is a Lipschitz continuous function; An incremental model of a 1-DOF robotic arm is constructed using historical data, and a state-dependent gain matrix is ​​introduced. Multiply both sides of formula (1) by the Moore-Penrose pseudo-inverse. , (3) in, It includes all the unmodeled dynamics in formula (1); at sufficiently high frequencies, the unknown The data can be estimated using a one-step forward approach: (4) in , , This represents the sampling time. This leads to the incremental model expression: (5) in This represents the model estimation error. This represents an incremental control strategy; First, an incremental model of a 1-DOF robotic arm is constructed using historical data, and a state-dependent gain matrix is ​​introduced. An offline deep learning network was established using the collected system input-output dataset. To train and obtain the state-related incremental gain weight matrix : (6)。 3. The robotic arm tracking control method based on transferable deep incremental reinforcement learning according to claim 1, characterized in that: The specific process of S2 is as follows: For S2, following the residual reinforcement learning mechanism, the system's deep tracking strategy is designed as follows: (7) in This represents an incremental guidance strategy. Represents the incremental residual strategy; the tracking error of the system is defined as... The formula for the dynamic characteristics of incremental tracking error is obtained as follows: (8) Establish an incremental guidance strategy for a 1-DOF robotic arm The formula is as follows: (9) in It is a defined constant matrix.

4. The robotic arm tracking control method based on transferable deep incremental reinforcement learning according to claim 1, characterized in that: Establish an incremental residual strategy and incremental model error The zero-sum game problem is used to consider the impact of model error on control performance; It is considered to minimize the number of participants, with the aim of reducing tracking error and control volume; while The value function of a zero-sum game is defined as follows: It is considered to maximize the participant's representation of the worst model error. (11) in , It is a positive definite matrix. It uses positive scalar control weights. It is a predefined constant. It is an integral virtual time variable. , , Representing virtual time Tracking error, incremental residual strategy, and model error; The optimal value function satisfies the following formula: (12) The value function satisfies the following Hamilton-Jacobi-Isaacs equation: (13) The strategies that satisfy the above equation (13) correspond to potential zero-sum game saddle points, which satisfy the following inequality: (14) Saddle Point It can be obtained through the following equation: (15) Based on this, the optimal residual strategy is obtained. And worst model error The solution is: (16) (17)。 5. The robotic arm tracking control method based on transferable deep incremental reinforcement learning according to claim 1, characterized in that: In S4, an asynchronous deep value network is used to approximate the cost function, which takes the form of: (18) in and Let these represent the ideal weights and activation function of the adaptive layer, respectively; the base layer is represented as... ,in These are the base layer weights. Indicate its activation function, Let be the number of layers; the weights of the base layers are updated at a slower rate to ensure training stability, and the specific formula for its update law is: (24) in It is the learning rate of the base layer; this slow timescale update is designed to optimize and accurately estimate... This plays a crucial role in ensuring the stability of the overall training process; The adaptive layer weights are updated at a relatively fast rate to ensure convergence. The specific formula for its update law is as follows: (25) Where the coefficient , It is an introduced scalar gain used to balance the respective contributions of real-time data and empirical data in the online learning of adaptive layer weights; It is a constant positive definite gain matrix; Indicates the sample size of the empirical data; Obtain the optimal residual strategy And worst model error strategy The approximate solution is: (26) (27) The final control strategy for the 1-DOF robotic arm system is as follows: (28)。 6. The robotic arm tracking control method based on transferable deep incremental reinforcement learning according to claim 1, characterized in that: In S5, instead of redesigning the controller of the high-dimensional robotic arm by repeating the above design process, the dynamic similarity between the decoupled subsystem and the 1-DOF robotic arm system is utilized to transfer the depth increment model and depth value network trained on the 1-DOF robotic arm to other high-dimensional robotic arms; the specific implementation strategy is as follows: for n The robotic arm system is first decoupled into n subsystems: (29) in It is the first The state of each subsystem; It is the first Control inputs for each subsystem; Indicates the first The combined effect of unknown dynamics and coupling terms of individual subsystems; Indicates the first Based on the control input gain matrix of each subsystem, the depth increment model corresponding to the subsystem can be obtained as follows: (30) The first high-dimensional robotic arm system The base layers of the deep value network of each subsystem are directly shared by the base layer network trained on the 1-DOF robotic arm, as shown in the following formula: (31) Similarly, the first The adaptive layer weights of each subsystem are initialized to the adaptive layer weights trained on a 1-DOF robotic arm: (32) Furthermore, similar to (25), online data is used for updating. To adapt to environmental changes, the specific formula is as follows: (33) Therefore, the first The value function estimate of the individual subsystems is expressed as follows: (34) No. The deep incremental model of each subsystem is transferred from a deep network trained on 1 degree of freedom. Specifically: (35) Thus, the strategy transfer from a 1-DOF robotic arm to an n-DOF robotic arm was achieved, as shown below: (36) (37) (38) (39)。

Citation Information

Patent Citations

  • An optimal control method for robot trajectory tracking based on all-wheel drive system

    CN119057776B

  • Fixed time optimal control method and system for single-connecting-rod mechanical arm

    CN119501947A