Robot control method and system based on model-guided offline reinforcement learning

By constructing a linear incremental model of nonlinear robots and generating synthetic data sets, the robot data set is expanded, and the problems of transfer difference and data deviation of reinforcement learning strategies in the existing technology are solved, and efficient and accurate robot control is achieved.

CN119511740BActive Publication Date: 2025-05-13NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510091187.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

In the prior art, when implementing robot control based on reinforcement learning algorithms, the reinforcement learning strategy that is first trained by the emulator and then deployed with hardware has poor migration performance, and the traditional offline reinforcement learning algorithm has data deviation problems.

Method used

Using the offline reinforcement learning method based on model guidance, by constructing a linear incremental model of a nonlinear robot, a linear incremental model is learned and a synthetic data set is used to generate a synthetic data set, and the pre-collected robot data set is expanded to form an enhanced data set for training reinforcement learning strategies.

Benefits of technology

It improves the efficiency and accuracy of robot control, enhances adaptability and flexibility, alleviates the problem of poor transfer of reinforcement learning strategies, and improves data bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119511740B_ABST
    Figure CN119511740B_ABST
Patent Text Reader

Abstract

The present invention discloses a robot control method and system based on model-guided offline reinforcement learning, the method steps include: step S01. constructing a linear increment model of a nonlinear robot and constructing a Q function; step S02. using pre-collected training data to iteratively solve the optimal increment strategy corresponding to the control input increment, and learning the linear increment model at the same time; step S03. using the learned linear increment model to perform forward prediction to generate a synthetic data set, and adding it to the robot data set to form an enhanced data set; step S04. using the enhanced data set to train the robot's reinforcement learning strategy to control the robot in real time. The present invention has the advantages of simple implementation method, high control efficiency and precision, strong adaptability and flexibility, etc., can alleviate the problem of poor migration of the traditional reinforcement learning strategy of first simulator training and then hardware deployment, and improve the data deviation problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot control technology, and in particular to a robot control method and system based on model-guided offline reinforcement learning. Background Art

[0002] Reinforcement learning is a mechanism for learning from experience. Robot control based on reinforcement learning algorithms is to learn the optimal control strategy through the interaction between the intelligent agent and the environment, so that the robot can achieve autonomous learning and control in complex and unknown environments. This learning mechanism can not only improve the adaptability and flexibility of the robot, but also reduce the dependence on precise hardware calibration, making the robot control more flexible and efficient.

[0003] In the prior art, when implementing robot control based on reinforcement learning algorithms, it is usually adopted to first train in the simulator and then deploy the reinforcement learning strategy in the hardware. That is, in the simulation environment, the robot learns the control strategy by interacting with the virtual environment constructed by the simulator. Through the reinforcement learning algorithm, the robot's strategy will be continuously optimized; then, in the hardware deployment stage, the strategy trained in the simulation environment is migrated to the control system of the real robot. However, this method of first training in the simulator and then deploying the reinforcement learning strategy in the hardware has poor migration performance. Summary of the invention

[0004] The technical problem to be solved by the present invention is: in response to the technical problems existing in the prior art, the present invention provides a robot control method and system based on model-guided offline reinforcement learning, which has a simple implementation method, high control efficiency and accuracy, and strong adaptability and flexibility. It can not only alleviate the problem of poor migration of traditional reinforcement learning strategies that are first trained on a simulator and then deployed on hardware, but also improve the data bias problem of traditional offline reinforcement learning algorithms.

[0005] In order to solve the above technical problems, the technical solution proposed by the present invention is:

[0006] A robot control method based on model-guided offline reinforcement learning, comprising the following steps:

[0007] Step S01: Construct a linear incremental model of the nonlinear robot, wherein the linear incremental model includes the state input of the robot x and the increment of the control input , based on the linear incremental model of the nonlinear robot Function, Functions include state input x and the increment of the control input ;

[0008] Step S02: Iteratively solve using the pre-collected robot dataset , k represents the time step, and the increment of the control input is obtained The corresponding optimal incremental strategy and the linear incremental model are learned at the same time;

[0009] Step S03. Use the linear incremental model learned in step S02 to perform forward prediction to generate a synthetic data set of state input and control input, and add it to the pre-collected robot data set to expand the robot data set to form an enhanced data set;

[0010] Step S04: Using the enhanced data set to train the robot's reinforcement learning strategy to control the robot in real time.

[0011] Furthermore, the linear incremental model is constructed as:

[0012]

[0013] in, x , u Represent the robot's state and control input respectively, represents the increment of the control input, B represents the input matrix, A represents the state transfer matrix, k represents the time step;

[0014] The linear incremental model is expanded to form an augmented incremental system, and the augmented linear incremental model is constructed as follows:

[0015]

[0016]

[0017] in, I represents the unit matrix, , Respectively express A. B The matrix obtained after linear augmentation is: For the unit array.

[0018] Further, the constructed The expression of the function is:

[0019]

[0020]

[0021]

[0022] in, P is a symmetric positive definite matrix, Q and R are the cost function weight matrices for state and input respectively, is the incremental strategy matrix;

[0023] By solving Get the optimal incremental strategy matrix for:

[0024]

[0025] The corresponding optimal incremental strategy is .

[0026] Furthermore, step S02 also includes solving the matrix Z , by following the formula Iterative Learning , and obtain the solution in least squares form:

[0027]

[0028] in, Indicates the matrix in the jth iteration Z The vectorized representation of is the penalty function, , represents the penalty function value of the kth iteration, express Y ( k ), express L The vectorized representation of l Indicates the data sequence number in the data set.

[0029] Further, The build steps include:

[0030] according to , ,get The Bellman equation for the function is as follows:

[0031]

[0032] according to get:

[0033]

[0034] Further converted to:

[0035]

[0036]

[0037] The final build is: .

[0038] Further, step S03 includes:

[0039] Step S301. Matrix obtained by iterative solution Z Calculate , according to the calculated Learn to get a linear incremental model;

[0040] Step S302. Perform forward prediction based on the learned linear incremental model to obtain a synthetic data set;

[0041] Step S303: Add the synthesized dataset to the pre-collected robot dataset to form an enhanced dataset.

[0042] Further, in step S301, according to the formula Calculate , and then determine the linear incremental model .

[0043] Furthermore, in step S302, forward prediction is performed according to the following formula:

[0044]

[0045]

[0046] Where i represents the number of steps to perform forward prediction;

[0047] Obtained from prediction The synthetic dataset is formed.

[0048] A robot control system based on model-guided offline reinforcement learning comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.

[0049] A computer-readable storage medium storing a computer program, wherein the computer program implements the above method when executed by a processor.

[0050] Compared with the prior art, the advantages of the present invention are: the present invention constructs a linear incremental model of the nonlinear robot, constructs a Q function based on the linear incremental model, directly trains the task strategy on a pre-collected offline data set, obtains the optimal incremental strategy through iterative solution and simultaneously learns the linear incremental model, uses the learned linear incremental model to generate a synthetic data set to expand the pre-collected offline data set, thereby increasing the diversity of data, and realizes the model guidance mechanism in the linear space by guiding the Q learning method, which can effectively improve the dynamic adaptability of the strategy obtained by offline training when it is deployed online, which can not only alleviate the problem of poor migration of the traditional reinforcement learning strategy of first training with a simulator and then deploying it on hardware, but also improve the data deviation problem of traditional offline reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic diagram of the implementation flow of the robot control method based on model-guided offline reinforcement learning in this embodiment. DETAILED DESCRIPTION

[0052] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.

[0053] In this embodiment, model-based offline reinforcement learning includes the following three aspects: first, the control strategy in linear space is determined by a linear incremental model to achieve the control of nonlinear robots; second, the pre-collected offline robot dataset is used to solve function, realizing simultaneous online learning of the optimal incremental strategy and the linear incremental model. Third, the learned linear incremental model is used to generate additional synthetic data sets (including state input and control input data). The synthetic data sets are added to the pre-collected offline robot data sets for reinforcement learning strategy training to solve the data bias problem.

[0054] like Figure 1 As shown, the steps of the robot control method based on model-guided offline reinforcement learning in this embodiment include:

[0055] Step S01: Construct a linear incremental model of the nonlinear robot, wherein the linear incremental model includes the state input of the robot x and the increment of the control input , based on the linear incremental model of the nonlinear robot Function, Functions include state input x and the increment of the control input ;

[0056] Step S02: Iteratively solve using the pre-collected robot dataset , k represents the time step, and the increment of the control input is obtained The corresponding optimal incremental strategy and the linear incremental model are learned at the same time;

[0057] Step S03. Use the linear incremental model learned in step S02 to perform forward prediction to generate a synthetic data set of state input and control input, and add it to the pre-collected robot data set to expand the robot data set to form an enhanced data set;

[0058] Step S04: Use the enhanced data set to train the robot's reinforcement learning strategy to control the robot in real time.

[0059] Through the above steps, this embodiment first constructs a linear incremental model of the nonlinear robot, constructs a Q function based on the linear incremental model, directly trains the task strategy on a pre-collected offline data set, obtains the optimal incremental strategy through iterative solution and simultaneously learns the linear incremental model, uses the learned linear incremental model to generate a synthetic data set to expand the pre-collected offline data set, thereby increasing the diversity of the data, and realizes the model guidance mechanism in the linear space by guiding the Q learning method, which can effectively improve the dynamic adaptability of the strategy obtained by offline training when it is deployed online, which can not only alleviate the problem of poor migration of the traditional reinforcement learning strategy of first training on the simulator and then deploying on the hardware, but also improve the data deviation problem of traditional offline reinforcement learning.

[0060] In this embodiment, the motion characteristics of the robot system can be described by the following nonlinear system:

[0061] (1)

[0062] in, x , u Represent the robot's state and control input respectively; , represent the state matrix and input matrix respectively.

[0063] The traditional robot nonlinear system model is difficult to accurately model the robot motion characteristics in the real world. Therefore, this embodiment uses historical data to construct an equivalent linear incremental system of the nonlinear robot, so as to facilitate the subsequent implementation of the model guidance mechanism to solve the data deviation problem. The detailed construction steps are as follows:

[0064] First, in order to separate the unknown model information of the robot system from the known items, multiply the constant matrix on both sides B ,get:

[0065] (2)

[0066] in, Contains all unknown model information. Unknown The historical data available are estimated as follows:

[0067] (3)

[0068] Thus, the linear incremental model can be obtained:

[0069] (4)

[0070] in, is the estimation error, which comes from the difference between the data at adjacent sampling times.

[0071] Considering that the gap between two adjacent steps is small in a data set constructed at a high sampling frequency, the estimation error is small and the impact on the strategy performance is negligible. Therefore, this embodiment will develop an offline reinforcement algorithm for the robot based on the following linear incremental model:

[0072] (5)

[0073] In order to further build a stable incremental task strategy, the above preliminary linear incremental model is expanded to form an augmented incremental system. The augmented linear incremental model is constructed as follows:

[0074] (6)

[0075] (7)

[0076] in, I represents the identity matrix of suitable dimension, , Respectively express A. B The matrix obtained after linear augmentation is: For the unit array.

[0077] In this embodiment, the following steps are followed to construct function:

[0078] First, by defining the following cost function Reflects robot task performance requirements:

[0079] (8)

[0080] in, is a penalty function that reflects the requirements of a given task. , Q and R are the cost function weight matrices for state and input respectively.

[0081] The following value function is further used to reflect the robot's incremental task strategy: Performance after the next k steps:

[0082] (9)

[0083] in, is the incremental strategy matrix,

[0084] Further build up The function is:

[0085] (10)

[0086] Considering that the value function can be further transformed into:

[0087] (11)

[0088] in is a symmetric positive definite matrix.

[0089] Can The function can be further transformed into:

[0090] (12)

[0091] The above formula can be further transformed into:

[0092] (13)

[0093] (14)

[0094] (15)

[0095] By solving The optimal incremental strategy matrix is:

[0096] (16)

[0097] The corresponding optimal incremental strategy is .

[0098] Through the above derivation, the optimal incremental strategy solution is transformed into a matrix Estimates.

[0099] To realize the matrix The following steps can be used to estimate:

[0100] First, according to , we can get The Bellman equation for the function is as follows:

[0101] (17)

[0102] According to the above derivation, , we can get:

[0103] (18)

[0104] In practice, for convenience, the above formula is converted to:

[0105]

[0106]

[0107] in, The matrices are The vectorized form can facilitate the solution of the algorithm. The next step is to iterate the learning Therefore, formula (19) can be further written as:

[0108] (19)

[0109] For estimation , the least squares solution of the above equation is:

[0110]

[0111] (20)

[0112] in, represents the penalty function value of the kth iteration, express Y ( k ), express L The vectorized representation of l Indicates the data sequence number in the data set.

[0113] In this embodiment, step S03 includes:

[0114] Step S301. Matrix obtained by iterative solution Z Calculate , according to the calculated Learn to get a linear incremental model;

[0115] Step S302. Perform forward prediction based on the learned linear incremental model to obtain a synthetic data set;

[0116] Step S303: Add the synthesized dataset to the pre-collected robot dataset to form an enhanced dataset.

[0117] In the above step S301, due to the linear incremental system is the unit matrix, which can be obtained from the iterative solution The matrix can be calculated as follows :

[0118] (twenty one)

[0119] Extracting the build increment system Required model information , and then determine the linear incremental model .

[0120] In the above step S302, forward prediction is performed specifically according to the following formula:

[0121]

[0122]

[0123] Among them, i represents the number of times the prediction is performed;

[0124] Obtained from prediction The synthetic dataset is formed.

[0125] In a specific application embodiment, the following algorithm 1 can be used to implement offline reinforcement learning iterative learning: It is understandable that in actual applications, in order to prevent problems such as singular values, the matrix Z can also be iteratively learned using gradient descent algorithms such as the Adam algorithm.

[0126] Algorithm 1: Model-guided offline reinforcement learning algorithm

[0127] Input: Threshold: ; Offline dataset ;

[0128] Output: Incremental strategy matrix

[0129] 1. if , then

[0130] 2. Based on the dataset Conduct a strategy assessment:

[0131] 3.

[0132] 4. Strategy improvement:

[0133] 5.

[0134] 6. Model extraction:

[0135] 7.

[0136] 8. Model forward prediction:

[0137] 9. Generate a data set through forward prediction of Algorithm 2

[0138] 10.

[0139] 11. end if

[0140] Then, based on the incremental model obtained above, forward prediction is performed through Algorithm 2 to obtain the synthetic dataset , as shown in lines 3-5 of Algorithm 2. Synthetic Dataset Add to pre-collected offline datasets To expand the existing data set and obtain a data set with further enhanced data richness .

[0141] Algorithm 2: Model forward prediction algorithm

[0142] enter:

[0143] Output:

[0144] 1. = {}

[0145] 2. for , do the following

[0146] 3.

[0147] 4.

[0148] 5. Add to middle

[0149] 6. end

[0150] This embodiment further provides a robot control system based on model-guided offline reinforcement learning, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.

[0151] It is understandable that the above method of this embodiment can be executed by a single device, such as a computer or server, etc., and can also be applied to a distributed scenario and completed by multiple devices in cooperation with each other. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps in the above method of this embodiment, and multiple devices interact to complete the above method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing related programs to implement the above method of this embodiment. The memory can be implemented in the form of a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device. The memory can store an operating system and other applications. When the above method of this embodiment is implemented by software or firmware, the relevant program code is stored in the memory and called and executed by the processor.

[0152] This embodiment further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0153] Those skilled in the art should understand that the above-mentioned embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0154] The above is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.

Claims

1. A robot control method based on model-guided offline reinforcement learning, characterized in that the steps include: Step S01: Construct a linear incremental model of the nonlinear robot, wherein the linear incremental model includes the state input of the robot x and the increment of the control input , based on the linear incremental model of the nonlinear robot Function, Functions include state input x and the increment of the control input , constructed as described The expression of the function is: in, P is a symmetric positive definite matrix, Q and R are the cost function weight matrices for state and input respectively, is the incremental strategy matrix, B represents the input matrix, A represents the state transfer matrix, which is a known unit matrix, k represents the time step, x Indicates the status of the robot. I represents the unit matrix, , Respectively express A.B The matrix obtained after augmentation; Step S02: Iteratively solve using the pre-collected robot dataset , k represents the time step, and the increment of the control input is obtained The corresponding optimal incremental strategy and the linear incremental model are learned at the same time; Step S03. Use the linear incremental model learned in step S02 to perform forward prediction to generate a synthetic data set of state input and control input, and add it to the pre-collected robot data set to expand the robot data set to form an enhanced data set; Step S04: Use the enhanced data set to train the robot's reinforcement learning strategy, and use the trained reinforcement learning strategy to control the robot in real time.

2. The robot control method based on model-guided offline reinforcement learning according to claim 1, characterized in that: The linear incremental model is constructed as: in, x Indicates the status of the robot. represents the increment of the control input, B represents the input matrix, A represents the state transfer matrix, which is a known unit matrix, and k represents the time step; The linear incremental model is expanded to form an augmented incremental system, and the augmented linear incremental model is constructed as follows: in, I represents the unit matrix, , Respectively express A.B The matrix obtained after augmentation.

3. The robot control method based on model-guided offline reinforcement learning according to claim 2, characterized in that: By solving Get the optimal incremental strategy matrix for: The corresponding optimal incremental strategy is .

4. The robot control method based on model-guided offline reinforcement learning according to claim 3, characterized in that: Step S02 also includes solving the matrix Z , by following the formula Iterative Learning , and obtain the solution in least squares form: Indicates the matrix in the jth iteration Z The vectorized representation of is the penalty function, , represents the penalty function value of the kth iteration, express Y ( k ), express L The vectorized representation of l Indicates the data sequence number in the data set.

5. The robot control method based on model-guided offline reinforcement learning according to claim 4, characterized in that: The build steps include: according to , ,get The Bellman equation for the function is as follows: according to get: Further converted to: The final build is: .

6. The robot control method based on model-guided offline reinforcement learning according to claim 4 or 5, characterized in that: Step S03 includes: Step S301. Matrix obtained by iterative solution Z Calculate , according to the calculated Learn to get a linear incremental model; Step S302. Perform forward prediction based on the learned linear incremental model to obtain a synthetic data set; Step S303: Add the synthesized dataset to the pre-collected robot dataset to form an enhanced dataset.

7. The robot control method based on model-guided offline reinforcement learning according to claim 6, characterized in that: In step S301, according to the formula Calculate , and then determine the linear incremental model .

8. The robot control method based on model-guided offline reinforcement learning according to claim 6, characterized in that: In step S302, forward prediction is performed according to the following formula: Where i represents the number of steps to perform forward prediction; Obtained from prediction The synthetic dataset is formed.

9. A robot control system based on model-guided offline reinforcement learning, comprising a processor and a memory, wherein the memory is used to store a computer program, characterized in that: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Unmanned ship fault-tolerant control method based on model reference reinforcement learning

    CN114296350A

  • A longitudinal control method for fixed-wing UAVs based on output feedback Q-learning

    CN114935944A