Space robot target capturing method based on model-assisted reinforcement learning
By introducing deep Lagrangian neural network into the traditional model-free reinforcement learning method, training a state transfer model and generating virtual experience, the problem of difficulty in analyzing dynamic model and high cost of machine learning training in traditional methods is solved, and a more efficient and accurate controller design for free-flying robots to capture mobile targets is achieved.
Patent Information
- Application Number
- CN202510220146.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Traditional modeling and control methods are difficult to analyze dynamic models, and the method combined with machine learning also has the problems of high training costs, low success rate for complex models and inaccurate training, especially in the controller design of free-flying robots to capture mobile targets.
Using a method based on deep Lagrangian neural network assisted model reinforcement learning, we use the dynamic constraints of the spatial free-flying robot system, build an environmental model, and design a deep Lagrangian neural network framework to train the system dynamic model to obtain a state transfer model, and then generate virtual experience assisted reinforcement learning training.
The sample efficiency of reinforcement learning, convergence speed and the speed and accuracy of the controller model training to capture mobile targets are improved. It is suitable for a variety of systems and operating conditions and has strong generalization.
Smart Images

Figure CN120068646A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of spatial free-flying robot control, and particularly relates to a dynamic capture method for a spatial flying robot based on a deep Lagrangian neural network-assisted model reinforcement learning. Background Art
[0002] Free-flying space robots are a type of robot system that can operate autonomously or semi-autonomously in a microgravity environment. As the core execution unit in space missions, they have extensive applications in fields such as satellite maintenance, space station auxiliary operations, and deep space exploration.
[0003] In the design of the controller for a free-flying robot to grasp a moving target, traditional methods mainly rely on accurate mathematical models, usually using Newton's classical mechanics and material mechanics to deduce the dynamic equations, such as using Lagrangian equations, Hamiltonian equations, etc. However, these methods face many challenges in practical applications, mainly due to the uncertainty of model parameters and the difficulty of modeling in complex environments.
[0004] With the rapid development of artificial intelligence technology in the field of robotics, the design of the controller for free-flying robots has gradually introduced machine learning methods, mainly including the following two categories: The first category is the data-driven modeling method, that is, by giving the dynamic constraints and equation structures of the robot, using deep learning techniques (including deep Lagrangian neural networks and Manhattan neural networks) to obtain the parameters or equations of the robot system. Subsequently, combined with the trained system model, traditional control algorithms such as PID and MPC are used for controller design. However, the difficulty of this method lies in how to obtain an accurate and reliable system model. The second category is to use reinforcement learning to train the controller for the robot to complete specific tasks end-to-end. However, pure model-free reinforcement learning methods have problems such as long training cycles, difficult parameter tuning, and insufficient controller accuracy.
[0005] In recent years, to improve the training sample efficiency and model accuracy of reinforcement learning, scholars have proposed a model-based reinforcement learning method, which mainly generates virtual experiences by having the robot learn the environment model (mainly including the state transition model and the environment reward model), and combines real experiences and virtual experiences for sampling to update the policy network. The key to this model reinforcement learning algorithm lies in the accuracy of the robot's learning environment model.
[0006] In summary, traditional modeling control methods are difficult to analyze the dynamic model, and the methods combined with machine learning still have some problems such as high training costs, low success rates and inaccuracies in training complex models. Summary of the Invention
[0007] To solve the above technical problems, the present invention proposes a mobile target capture method based on deep Lagrangian neural network-assisted model reinforcement learning, and the method specifically includes the following steps:
[0008] Step 1: Derive the dynamic constraints of the space free-flying robot system and build an environment model;
[0009] Step 2: Set appropriate parameters and objective functions for model-free reinforcement learning training;
[0010] Step 3: Design a deep Lagrangian neural network framework and train a system dynamics model, and obtain a state transition model;
[0011] Step 4: With the parameters in Step 2 as a reference, add the state transition model obtained in Step 3 to generate virtual experiences for the model to assist in completing the reinforcement learning training process.
[0012] The specific steps of Step 1 are as follows:
[0013] Step 1.1: The general expression of the dynamic model of the space free-flying robot system is as follows:
[0014]
[0015] where q is the generalized state; H(q), G(q) represent the inertia matrix, Coriolis matrix of the system, and the torque generated due to gravity respectively; τ is the generalized force or torque.
[0016] Step 1.2: Build an environment model, and the model includes the position and motion equation of the moving target, the position where the robot system is located, time, time step interval, and the number of training rounds;
[0017] Step 1.3: Determine the state space, action space, and reward function, and the experience pool set is defined as follows:
[0018]
[0019] where, max t represents the maximum number of steps in each round; p represents the maximum number of training rounds.
[0020] s represents the state variable of the state space, and is set to a represents the action variable of the action space, and is set to a = τ.
[0021] In Step 1.3, t, δt, and ep respectively represent the moment after t time steps in each round, the time step interval, and the number of training rounds; the pair (t, ep) represents that the system is currently at the t-th time step of the ep-th round; P eef , Pbase , respectively represent the spatial coordinates of the positions of the end effector, the base, the base target, and the moving target in the inertial coordinate system; the reward function includes sparse rewards, distance rewards, etc.
[0022] The specific steps of step 2 are as follows:
[0023] Step 2.1: Determine the control policy function of the robot system and the reinforcement learning neural network, set the training parameters of the reinforcement learning for model-free offline reinforcement learning training, and use the experience data recorded in to update the policy function according to the objective function J(π);
[0024] Step 2.2: Record the rewards obtained in each episode during training and select a reward threshold r m , after m episodes of training, the reward reaches r m , then use the controller model trained in the m-th episode to drive the robot system to grasp the moving target, collect the experience information with ep = m and form a data set:
[0025]
[0026] Among them, the variable with a subscript t represents the assignment of this variable at the t-th time step, q t , and τ t can be directly obtained from the state variable s t or the action variable a t in the experience pool, is defined as:
[0027]
[0028] The reinforcement learning neural network described in step 2.1 is an MLP neural network, and the policy function is π θ , where θ is the network parameter of the control policy function; the reinforcement learning training parameters include the number of network layers L, the number of neurons n in each layer L , the maximum number of experiences batch that the experience pool set can accommodate, the learning rate α learn , the maximum number of training episodes p 1 , the discount rate γ, and extracting i 1 experiences every i 2 steps for updating, etc.
[0029] The specific steps of step 3 are as follows:
[0030] Step 3.1: Set the loss function according to the Lagrangian equation Set the training parameters, including the number of network layers \(l\) and the maximum number of training epochs \(epoch\), etc., and use a subset of \(S\) to train the deep Lagrangian neural network to obtain the system dynamics model
[0031] Step 3.2: According to the dynamics model Obtain the state transition model Furthermore, obtain a virtual experience.
[0032] The state transition model described in Step 3.2 Is expressed as \((s t ,a t )→(r t f ,s t+1 f ), that is, according to the current state and action \((s t ,a t ), predict the state at the next moment and the obtained reward value \((r t f ,s t+1 f ). The virtual experience includes \(s t ,a t ,r t f ,s t+1 f ;
[0033] The process of generating a virtual experience is as follows:
[0034] According to \((s t ,a t ) to obtain \(q t , P eef,t ,P base,t , τ t ; According to q t , τ t Predict According to the predicted To obtain \(q t+1 , Use the following formula:
[0035]
[0036] According to \(q t+1 Obtain \(P eef,t+1 And \(P base,t+1 ; Both can be obtained by sensors; According to \(s t+1 Calculate \(r t .
[0037] Among them, according to q t+1 obtain P eef,t+1 and P base,t+1 There are two methods: one is to derive it according to the kinematic model of the robot system, that is, use the Jacobian matrix of the space robot system to obtain the poses of the end effector and the base according to the joint space position, so as to obtain P eef,t+1 , P base,t+1 ; the other is to use the method of supervised learning. First, record a large number of q t in s t and [P eef,t , P base,t , and then use the BP neural network to train the Jacobian matrix. Since the relationship between the two is linear and relatively simple, it is very easy to implement prediction with the BP neural network.
[0038] The specific steps of step 4 are as follows:
[0039] Step 4.1: Generate virtual experience: Assume that the robot system starts to generate virtual experience when it is in state s t . Here, can be obtained by the policy function π θ . From it is concluded that generate a virtual experience obtain and then use to generate the next virtual experience and so on to generate n f virtual experiences, which are represented by the set and are defined as follows:
[0040]
[0041] Step 4.2: Carry out model reinforcement learning training with reference to the parameters set in step 2. After the training is completed, a controller that can accurately grasp the moving target is obtained.
[0042] The specific parameter settings of step 4.2 are as follows:
[0043] The setting of n f can balance the ratio of virtual experience and real experience in the experience pool. An overly large n f will lead to a longer training time per round. n f can be set to a value close to i 1 ; the number of training rounds p 2 can be appropriately reduced on the basis of p 1 . Specifically, its selection can be based on the following process:
[0044] Use the formula time MF ×p 1 =time MF ×m + time MB ×P 2 + time DeLaN Obtain the number of training rounds P required for model-based reinforcement learning with a time consumption similar to that of traditional model-free reinforcement learning methods 2 , where time MF and time MB respectively represent the time required for each round in the model-free reinforcement learning and the model-based reinforcement learning after adding ; time DeLaN represents the time consumed for training the deep Lagrangian neural network. Then, select a suitable p according to the following constraints 2 :
[0045] P 2 <p 2 + m < p 1 , p 2 ∈Z;
[0046] All other parameters can be the same as those set in Step 2
[0047] Step 4.2 The update method of the parameter θ of the policy function during the training process is as follows
[0048] Every round, the robot system generates n 1 virtual experiences every i f steps Then integrate them into , and then the robot system extracts i experience data to update the parameter θ of the policy function according to the optimization objective function J(π) 2
[0049] Beneficial effects
[0050]
[0052] By adding a state transition model trained by a deep Lagrangian neural network on the basis of traditional model-free reinforcement learning, the present invention realizes the improvement of the sample efficiency, convergence speed of reinforcement learning, and the rapidity and accuracy of training a controller model for grasping a moving target
[0051] The overall logic and principle of the present invention are clear and simple, applicable to a large number of different systems and various working conditions, and have strong generalization
[0052] Aiming at the problem of using offline reinforcement learning to train a controller for a space free-flying robot to grasp a moving target, the present invention makes the state transition model reliable by using a deep Lagrangian neural network Description of the drawings
[0053] Figure 1 is the flow chart of the method of the present invention;
[0054] Figure 2 is the model diagram of the simulation test;
[0055] Figure 3 is the acceleration error curve diagram of each joint of the deep Lagrangian training dynamics model in the simulation test;
[0056] Figure 4 is the reward change diagram of comparing this method and traditional model-free reinforcement learning in the training process during the simulation test;
[0057] Figure 5 is the error change diagram of comparing this method and traditional model-free reinforcement learning in the grasping process during the simulation test. Specific embodiments
[0058] The present invention will be specifically described below in conjunction with specific embodiments.
[0059] This embodiment is a case of a 12-degree-of-freedom free-space flying robot (hereinafter referred to as a space robot) based on a deep Lagrangian neural network-assisted SAC reinforcement learning algorithm for capturing a moving target at a constant speed. The space robot system is set to ignore gravity, and the 12 degrees of freedom include 3 rotational degrees of freedom of the base, 3 translational degrees of freedom, and 6 degrees of freedom of the robotic arm (as Figure 2 shown). Figure 1 is the flow chart of the method of the present invention, and the specific steps of the method are as follows:
[0060] Step 1: For a space free-flying robot system that ignores gravity, has a fully controllable base, and is equipped with a 6-degree-of-freedom robotic arm, the Lagrangian dynamics equation is
[0061]
[0062] where q is defined as:
[0063] q = [q m1 , q m2 ,..., q m6 , q b1 , q b2 ,..., q b6
[0064] where [q m1 , q m2 ,..., q m6 represents the joint angles of the 6 joints of the 6-degree-of-freedom robotic arm; [q b1 , q b2 , q b3 ,1] represents the Euler angles of the spherical joint of the controllable base; q b4 ,q b5 ,q b6 respectively represent the distances that the controllable base moves along the X-axis, Y-axis, and Z-axis. H(q) and respectively represent the inertia matrix and the Coriolis matrix of the system; τ is defined as the vector composed of the forces or torques of the 12 joints of the system.
[0065] Step 2: Set δt = 0.002s, the maximum number of training episodes p 1 = 100 (it is necessary to ensure the convergence of the training process and the relative accuracy of the trained controller model), and the maximum step length max of each episode t = 1500.
[0066] Furthermore, set the coordinates of the base target position as:
[0067]
[0068] Set the moving target trajectory equation as follows:
[0069]
[0070] The initial position of the base is (0, 0, 0), and the initial position of the end effector is approximately (0.17 (m), 0.1 (m), 0.1 (m))
[0071] Step 3: Define the state variable s and the action variable a as follows:
[0072]
[0073] a = τ
[0074] where [A, B,...] represents the new matrix obtained by concatenating the columns of matrices A, B, etc. with the same number of rows. P eef ,P base ∈R 1 ×3 is defined as:
[0075]
[0076] where the matrix with the superscript T in the upper right corner represents the transpose of the matrix, (x eef ,y eef ,z eef ) and (x base ,y base ,z base ) respectively represent the position coordinates of the end effector and the base in the three-dimensional inertial coordinate system;
[0077] They are the spatial coordinates of the base target and the moving target in the inertial coordinate system respectively; i is defined as:
[0078]
[0079] Where d(m) represents the tracking error, and its definition is as follows:
[0080]
[0081] Where ‖.‖ 2 represents the two-norm of the. matrix. All variables involved in s t can be obtained by the sensors of the system or calculated through simple calculations.
[0082] The reward function is defined as:
[0083] r = reward sparse +reward distance +penalty action
[0084] Where:
[0085]
[0086] penalty action = -5‖a‖ 2
[0087] Then set the clamping force of the actuator to increase as d decreases to capture the moving target.
[0088] Define the experience pool set:
[0089]
[0090] Step 4: Determine the control policy function π θ according to the action space dimension of the space robot, where θ is the network parameter of the control policy function. All networks involved in reinforcement learning are represented by MLP neural networks. Set the training parameters of SAC reinforcement learning for training.
[0091] Furthermore, set i 1 = 50, i 2 = 32, L = 4, n L = 128, learning rate α learn = 0.001, entropy regularization coefficient α = 0.1, discount rate γ = 0.99. The optimization objective is:
[0092]
[0093] Where: is the estimated expected return, rj is the immediate reward obtained at the moment of j index.
[0094] Step 5: In this case, after 18 rounds of training, the reward reaches the set r m = 10000 reward value. Then collect the experience data during the capture process with ep = 18 and make a dataset:
[0095]
[0096] The following steps 6, 7, and 8 are all processes of training the system dynamics model using the deep Lagrangian neural network.
[0097] Step 6: Design the architecture of the deep Lagrangian neural network. According to the Lagrangian equation (1) of the space free - flying robot, transform H(q) into:
[0098] H(q) = L(q, φ)L(q, φ) T
[0099] where L(q, φ) T is the transpose of the matrix L(q, φ); φ is the parameter of the neural network with the MLP architecture, which can be obtained by minimizing the violation of physical laws, and its optimization goal is:
[0100]
[0101] where the loss function is defined as the mean - square error, that is:
[0102]
[0103] where f and f -1 represent the inverse kinematics function of the system, expressed as:
[0104]
[0105] represents the estimated value of the inverse kinematics output containing the parameter φ, and its definition is as follows:
[0106]
[0107]
[0108] Then define the objective function as:
[0109]
[0110] Among them, Ω(φ) is the second norm of the network weights. λ is the regularization parameter, which adjusts the model complexity and the strength of regularization. The deep Lagrangian neural network can be trained by optimizing this objective function.
[0111] Step 7: Randomly select 750 groups of data from the dataset S in Step 4 and name it the training set S 1 . Use the data in S 1 to train the dynamics model according to Step 5.
[0112] Step 8: Set the network parameters, including l = 3, epoch = 30000, λ = 0.5. We can train an estimated dynamics model
[0113] The following Steps 9 - 12 are the process of the state transition model .
[0114] Step 9: According to (s t , a t ), q t , P eef,t , P base,t , i t , τ t can be obtained. According to q t , τ t predict
[0115] Step 10: According to the predicted obtain q t+1 , . The formula is as follows:
[0116]
[0117] Step 11: Obtain P t+1 and P eef,t+1 from q base,t+1 . There are two methods: One is to derive it according to the kinematic model of the robot system, that is, use the Jacobian matrix of the space robot system to obtain the poses of the end effector and the base according to the joint space position, so as to obtain P eef,t+1 , P base,t+1 ; The other is to use the method of supervised learning. First, record a large number of q t in s t and [P eef,t , P base,t], and then use BP neural network to train the Jacobian matrix. Since the relationship between the two is linear and relatively simple, it is also very easy to use BP neural network for prediction. According to the training results and q t+1 It can be deduced that P eef,t+1 and P base,t+1 Due to the determinism of the kinetic model, this case uses the first method to complete this step.
[0118] Step 12: can be obtained by sensors, and i t+1 and r t It can be obtained from step 2. So far, a state transition model can be obtained
[0119] Steps 13 and 14 are based on and the state s at time t t The process of generating 20 virtual experiences.
[0120] Step 13: Assume that the space robot is in state s t When the virtual experience is generated, It can be represented by the policy function π θ Get, by Conclusion Generate a virtual experience
[0121] Step 14: Get Then return to step 13 to generate the next virtual experience And so on to generate 20 virtual experiences. To express this f A collection of virtual experiences, defined as follows:
[0122]
[0123] Step 15: According to time MF ×p 1 =time MF ×m+time MB ×P 2 +time DeLaN Find P 2 Furthermore, time MF =3min,time MB =4min,time DeLaN =12min to get P 2 =59. Take p 2 =72.
[0124] Step 16: Based on Steps 4 and 14, perform model-based reinforcement learning training. Set that every 50 steps during the training process, 20 virtual experiences will be generated. Then integrate into and then, in extract 32 experience data for update.
[0125] Step 17: After the training is completed, make a graph of the return change per episode of the two reinforcement learning training processes. To facilitate comparing the superiority of adding auxiliary reinforcement learning, we take the number of episodes v 1 = 32 when the return of the model-based reinforcement learning increases significantly in a short time after m episodes, and the ratio of the time consumed reaches v 1 and the number of episodes v 2 = 55 of the model-free reinforcement learning with a slightly longer time and a higher return value during the training process. Then, respectively take the number of episodes v 3 = 65 and v 4 = 88 when the returns of the model-based and model-free reinforcement learning are the largest during the training process.
[0126] Step 18: Test the ability of the controller model obtained by the model-free reinforcement learning in m, v 2 and v 4 episodes to grasp a moving target. Then test the ability of the controller model obtained by adding model-based reinforcement learning in v 1 and v 3 episodes to grasp a moving target. The performance of the controller is analyzed by the change of d with t, and the rapidity of the controller model is represented by the time step when the error is first lower than 0.05 m, and the accuracy of the controller model is represented by the time step when the minimum value of the error is reached.
[0127] Simulation experiment:
[0128] To verify the effect of the present invention, the method in the above specific implementation manner is used to test the algorithm of the present invention in the MuJoCo simulation environment. The simulation content is to test the training of a controller for a space free-flying robot with a 6-degree-of-freedom robotic arm and a fully controllable 6-degree-of-freedom base to grasp a moving target. Figure 2 is the model diagram built; Figure 3 is the graph of the change of the test acceleration error of the dynamic model trained by the deep Lagrangian neural network with the time step for 12 joints (1200 groups are extracted from 1500 time step data for testing); Figure 4 is the graph of the return change during the training process of the model-free and model-based methods in Step 17; Figure 5It is a graph showing the variation of the error during the grasping process of the five models trained in step 18 with the time steps. It can be seen that the method of adding models can train a controller that can quickly control the error below 0.05m, and the lowest error can reach about 0.02m. In contrast, the traditional model-free method is slightly inferior, and there is a tendency for the error to rise after reaching the lowest error, indicating the instability of the controller. Experiments show that adding models can optimize the training process of reinforcement learning, making the trained controller model more accurate.
[0129] Thus, it can be seen that the deep Lagrangian neural network of the present invention can well assist in training the controller for a space free-flying robot to grasp a moving target.
[0130] The above examples of the present invention are only for illustrating in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manner of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solution of the present invention still fall within the protection scope of the present invention.
[0131] The present invention can also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A space robot target capture method based on model-assisted reinforcement learning, the method specifically comprising the following steps: Step 1: Derive the dynamic constraints of the space free-flying robot system and build an environmental model; Step 2: Set appropriate parameters and objective functions for model-free reinforcement learning training; Step 3: Design a deep Lagrangian neural network framework and train the system dynamics model, and derive the state transition model; Step 4: Using the parameters of step 2 as a reference, add the state transition model obtained in step 3 to generate virtual experience to assist in completing the reinforcement learning training process.
2. The method for capturing a mobile target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 1, characterized in that: The specific steps of step 1 are: Step 1.1: The general expression of the dynamic model of the space free-flying robot system is as follows: where q is a generalized state; H(q), G(q) represents the system's inertia matrix, Coriolis matrix and the moment due to gravity; τ is the generalized force or moment; Step 1.2: Build an environment model, which includes the position and motion equation of the moving target, the position of the robot system, the time and time step interval, and the number of training rounds; Step 1.3: Determine the state space, action space and reward function, and the experience pool set is defined as follows: Among them, max t represents the maximum step size of each round; p represents the maximum number of training rounds; s represents the state variable of the state space, set a represents the action variable in the action space, and is set to a=τ.
3. The method for capturing a mobile target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 1, characterized in that: In step 1.3, t, δt and ep respectively represent the time after t time steps in each round, the time step interval and the round number of training; the number pair (t, ep) means that the system is currently at the tth time step of the epth round; P eef , P base , The spatial coordinates of the positions of the end effector, base, base target, and mobile target in the inertial coordinate system are respectively represented; the reward function includes sparse rewards and distance rewards, etc.
4. The method for capturing a moving target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 1, characterized in that: The specific steps of step 2 are: Step 2.1: Determine the control strategy function and reinforcement learning neural network of the robot system, set the training parameters of reinforcement learning, and perform model-free offline reinforcement learning training. The empirical data recorded in is used to update the strategy function according to the objective function J(π); Step 2.2: Record the rewards obtained in each round of training and select a reward threshold r m , after m rounds of training, the return reaches r m Then use the controller model trained in the mth round to drive the robot system to grasp the moving target, collect ep=m experience information and make a data set: The variable with the subscript t indicates the value of the variable at the time step t, q t , and τ t The state variable s in the experience pool can be t or action variable a t Directly obtain from Defined as:
5. The method for capturing a mobile target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 4, characterized in that: The reinforcement learning neural network described in step 2.1 is an MLP neural network, and the strategy function is π θ , where θ is the network parameter of the control strategy function; reinforcement learning training parameters include the number of network layers L, the number of neurons in each layer n L , the maximum number of experience batches that the experience pool can accommodate, and the learning rate α learn , the maximum number of training rounds p1, the discount rate γ, and the extraction of i2 pieces of experience data for update every i1 steps, etc.
6. The method for capturing a moving target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 4, characterized in that: The specific steps of step 3 are: Step 3.1: Set up the loss function based on the Lagrange equation And set the training parameters, including the number of network layers l and the maximum number of training rounds epoch, etc., and use the subset of S to train the deep Lagrangian neural network to obtain the system dynamics model Step 3.2: Based on the kinetic model Get the state transition model And then get a virtual experience.
7. The method for capturing a moving target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 6, characterized in that: State transition model described in step 3.2 Expressed as (s t ,a t )→(r t f ,s t+1 f ), that is, according to the current state and action (s t ,a t ) predicts the state and reward value at the next moment (r t f ,s t+1 f ), the virtual experience includes s t ,a t ,r t f ,s t+1 f .
8. The method for capturing a moving target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 1, characterized in that: The specific steps of step 4 are: Step 4.1: Generate virtual experience: Assume that the robot system is in state s t When the virtual experience is generated, It can be represented by the policy function π θ Get, by Conclusion Generate a virtual experience get Reuse Generate the next virtual experience And so on to generate n f Virtual experience, with collection To represent, the definition is as follows: Step 4.2: Perform reinforcement learning training on the model with reference to the parameters set in step 2. After the training is completed, a controller that can accurately grasp the moving target is obtained.
9. The method for capturing a moving target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 8, characterized in that: Step 4.2 The specific parameters are set as follows: n f Set it to a value close to i1; the number of training rounds p2 is appropriately reduced based on p1, and the selection process is as follows: Use the formula time MF ×p1=time MF ×m+time MB ×P2+time DeLaN Calculate the number of training rounds P2 required for model reinforcement learning that is similar to the traditional model-free reinforcement learning method in terms of time consumption, where time MF and time MB Represent model-free reinforcement learning and join The time required for each round of the model reinforcement learning process; time DeLaN Represents the time spent on training the deep Lagrangian neural network, and then takes a suitable p2 according to the following constraints: P2<p2+m<p1, p2∈Z; Other parameters can be consistent with those set in step 2.
10. The method for capturing a mobile target based on deep Lagrangian neural network assisted model reinforcement learning according to claim 8, characterized in that: Step 4.2 The parameter θ of the training process policy function is updated as follows: In each round, the robot system generates n f Virtual Experience Then Integration Then the robot system Extract i2 pieces of experience data and update the parameters θ of the strategy function according to the optimization objective function J(π).
Citation Information
Patent Citations
Dynamics control method and system of space non-cooperative target navigation acquisition
CN108469737A
Target tracking method and device, unmanned aerial vehicle and storage medium
CN113554680A
Track planning method and device for data collection of unmanned aerial vehicle, equipment and medium
CN114840021A
Unmanned aerial vehicle modeling and control method, system, equipment and medium
CN117666359A
Multi-agent federated reinforcement learning-based vehicle-road collaborative control system and method under complex intersection
WO2024016386A1