Single-foot jumping robot motion control method based on coupling flow model driving

By constructing intelligent agent strategies, environmental dynamics models and coupled flow models, and utilizing the distribution difference measurement and dynamic reward mechanism of the coupled flow model, efficient environmental model learning of a single-legged jumping robot is achieved under limited real samples, solving the problem of strategy failure caused by model bias in existing technologies and improving learning efficiency and safety.

CN120802715APending Publication Date: 2025-10-17CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510832530.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies find it difficult to learn accurate environmental models from limited real samples, resulting in delayed convergence or divergence of bipedal robot motion control strategies. Existing methods also have the risk of strategy failure due to model bias.

Method used

Construct intelligent agent strategies, environmental dynamics models and coupled flow models, and through the interaction and optimization of real samples and simulated samples, utilize the distribution difference measurement and dynamic reward mechanism of the coupled flow model to achieve adaptive calibration and efficient learning of the environmental model.

Benefits of technology

It significantly improves the efficiency and reliability of strategy optimization of unipedal jumping robots under a small amount of real interaction data, reduces training costs and hardware loss risks, improves the prediction accuracy and generalization performance of the model in complex environments, and avoids safety issues in strategy training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120802715A_ABST
    Figure CN120802715A_ABST
Patent Text Reader

Abstract

The invention discloses a single-foot jumping robot motion control method based on coupling flow model driving, and the method comprises the steps: outputting a decision action through an intelligent agent strategy model, and carrying out the online interaction with a single-foot jumping robot, and collecting a real environment sample; an environment dynamic model is constructed through a probabilistic neural network, and the model is trained in a supervised learning mode to simulate the motion state of the robot; building a coupling flow model by utilizing a multi-layer coupling structure, converting the output distribution difference into a reward signal, and realizing dynamic optimization of an environment model by reconstructing a Markov decision process and combining with a reinforcement learning algorithm; and finally, interacting the intelligent agent strategy with the calibrated environment model to generate a high-precision simulation sample, and combining with a real sample to complete strategy iteration updating. According to the method, accumulated rewards equivalent to model-free reinforcement learning can be achieved only through a small amount of environment interaction, and the learning efficiency and the model generalization ability of a single-foot jumping robot strategy are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of robot motion control, and relates to a single-leg hopping robot motion control method based on a coupled flow model driving. BACKGROUND

[0002] How to learn the optimal strategy to control the motion of a biped robot in a shorter time is a challenging control problem. The problem can model the motion of a biped robot as a Markov decision process, and then solve it by using a reinforcement learning strategy optimization algorithm. Model-based reinforcement learning is a method of reinforcement learning, which can generate simulated samples based on a model, thereby reducing the need for interaction with the actual environment and improving learning efficiency.

[0003] The key of the model-based method lies in the accuracy of the model. If there is a deviation between the model and the actual environment, the generated simulated samples will contain errors, which not only delays the convergence of the strategy, but also may cause the strategy to diverge in a serious case. Although model integration, uncertainty perception exploration and online model updating can improve the adaptability of the algorithm to model errors, and realize a more reliable and efficient quadruped robot motion control strategy. However, the accuracy of the model is still a key challenge to realize efficient quadruped robot control. SUMMARY

[0004] The purpose of the application is to overcome the deficiencies in the prior art, and provide a single-leg hopping robot motion control method based on a coupled flow model driving, which solves the problem of how to learn an accurate environment model only by using limited real samples, and promotes efficient learning of the strategy based on the model.

[0005] To achieve the above-mentioned purpose, the technical scheme of the application is as follows: a single-leg hopping robot motion control method based on a coupled flow model driving, comprising the following steps:

[0006] Step 1, constructing and initializing an agent strategy, an environment dynamics model and a coupled flow model, the agent strategy taking the state dimension of the single-leg hopping robot as input and outputting the action dimension, the environment dynamics model taking the state dimension and the action dimension of the single-leg hopping robot as input and outputting the state dimension of the next state of the single-leg hopping robot and the reward dimension, and the coupled flow model taking the state-action pair of the real sample and the state-action pair of the simulated sample as input and outputting the state-action density estimation value, the state-action pair including the state dimension and the action dimension;

[0007] Step 2, setting the current time step t = 0, and sampling the initial state s0 from the initial state distribution P(s0) of the single-leg hopping robot;

[0008] Step 3, generate real samples from the interaction between the agent policy and the single leg hopping robot task and store them in the real sample experience replay pool;

[0009] Step 4, generate simulated samples from the interaction between the agent policy and the environment dynamics model and store them in the simulated sample replay pool;

[0010] Update the parameters of the agent policy, the environment dynamics model, and the coupled flow model in steps 5-15 to the maximum episode number:

[0011] Step 5, update the parameters of the environment dynamics model using real samples in the real sample experience replay pool;

[0012] Step 6, update the parameters of the coupled flow model using real samples in the real sample experience replay pool and simulated samples in the simulated sample replay pool;

[0013] Step 7, reconstruct the Markov decision process (S, A, R, T) to update the parameters of the environment dynamics model using the reinforcement learning algorithm SAC to maximize the cumulative reward R, where S is the temporary state space, the temporary state in the state space is the concatenation of the sample state and the sample action in the real sample experience replay pool, A is the temporary action space, the action in the temporary action space is the concatenation of the predicted next state and reward value based on the sample state and sample action, is the temporary reward function for distribution difference measurement, is the first regularization flow model of the coupled flow model, is the second regularization flow model of the coupled flow model, k , a l is the state-action pair in the real sample experience replay pool, and T is the sample trajectory length;

[0014] Update the agent policy parameters in steps 8-14 to the maximum time step:

[0015] Step 8, get the action a t from the agent policy with the state s t , the single leg hopping robot performs the action a t to get the next state s t+1 and the reward r t+1 , and construct the real sample (s t , a t , s t+1 , r t+1 );

[0016] Step 9, put the real sample (s t , a t , s t+1 , r t+1 ) into the real sample experience replay pool and update the current state to st =s t+1 ;

[0017] Step 10: Randomly select a real sample s from the real sample experience replay pool k As state s′ t , the agent strategy is based on the state s′ t Get action a′ t , the environment dynamics model obtains the next state s′ according to the action state pair t+1 and reward r′ t+1 , repeat H time steps to get the simulation trajectory

[0018] Step 11: Simulate the trajectory The included H simulated samples are stored in the simulated sample replay pool, and steps 10 and 11 are repeated until the number of simulated samples in the simulated sample replay pool exceeds the set value;

[0019] Step 12: Calculate the coupling flow loss of the simulation trajectory, and eliminate some simulation trajectories based on the coupling flow loss;

[0020] Step 13: Update the agent strategy parameters based on the real samples in the real sample experience replay pool and the simulated trajectories in the simulated sample replay pool;

[0021] Step 14: After the time steps are accumulated to the maximum time step, the loop from step 8 to step 14 is ended;

[0022] Step 15: After the number of episodes is accumulated to the maximum number of episodes, the loop from step 5 to step 15 is ended;

[0023] Step 16: Based on state s t The updated agent strategy outputs action a t .

[0024] Furthermore, the calculation method of the coupling flow loss of the simulation trajectory in step 12 is:

[0025]

[0026] Furthermore, in step 12, part of the simulation trajectories are eliminated according to the coupling flow loss, which is to sort the simulation trajectories according to the magnitude of the coupling flow loss and eliminate the simulation trajectories whose coupling flow loss is greater than a set threshold.

[0027] Furthermore, the set threshold is determined by the coupling flow losses corresponding to the first 80% of the coupling flow losses in ascending order.

[0028] Furthermore, the parameters of the agent strategy, environment dynamics model, and coupling flow model are updated using the gradient descent algorithm, where the parameter of the agent strategy is θ and the update formula is: is the gradient of the agent strategy loss function L(θ) with respect to the parameter θ; the parameter of the environment dynamics model is ω, and the update formula is is the gradient of the loss function L(ω) of the environmental dynamics model with respect to the parameter ω; the parameter of the coupled flow model is ∈, and the update formula is is the gradient of the loss function L(∈) of the coupled flow model with respect to the parameter ∈, and α1, α2, and α3 are the set learning rates.

[0029] Furthermore, the agent strategy loss function L(θ) is

[0030]

[0031] Among them, D KL represents KL divergence calculation, Q represents action value function, and V represents state index function.

[0032] Furthermore, the loss function L(ω) of the environmental dynamics model is

[0033]

[0034] Among them, (μ ω (s t, a t )+s t ) is the s for the next state t+1 Estimation, μ ω (s t, a t ) represents the state s at the next moment t+1 The mean of s represents the next state t+1 The variance of .

[0035] Furthermore, the loss function L(∈) of the coupled flow model is

[0036]

[0037] Furthermore, in step 9, the real sample (s t ,a t ,s t+1 ,r t+1 ) is put into the real sample experience replay pool, when the number of samples in the real sample experience replay pool exceeds the set value, the real sample (s t ,a t ,s t+1 ,rt+1 ) replacing the sample that has been in the real sample experience replay pool the longest.

[0038] Further, the step 1 constructs the agent policy, the environment dynamics model and the coupled flow model, and then uniformly samples and initializes the parameters of the agent policy, the environment dynamics model and the coupled flow model between [-0.1, 0.1].

[0039] Compared with the prior art, the application has the following advantages:

[0040] The application first outputs a decision action through the agent policy model, and collects real environment samples through online interaction with a single-leg jumping robot; secondly, an environment dynamics model is constructed by a probabilistic neural network, and the model is trained in a supervised learning manner to simulate the motion state of the robot; then, a coupled flow model is built by using a multi-layer coupled structure, the distribution difference output by the coupled flow model is converted into a reward signal, and through reconstruction of a Markov decision process and combination of a reinforcement learning algorithm, dynamic optimization of the environment model is realized; finally, the agent policy interacts with the calibrated environment model to generate high-precision simulation samples, and the policy iteration is updated jointly with the real samples.

[0041] The application only needs a small amount of real interaction data as a benchmark, and combines simulation data generated by the model to realize high-precision iterative update of the environment dynamics model, thereby effectively reducing the experimental cost and hardware loss risk in the modeling process of the single-leg jumping robot.

[0042] Based on the precisely calibrated environment dynamics model, the simulation data used in the policy learning stage can more accurately reflect the dynamic characteristics of the real system. This characteristic avoids the problem of misleading policy training caused by model deviation, significantly improves the performance of the reinforcement learning algorithm in sample utilization efficiency, and enables the single-leg jumping robot to converge to the optimal control policy in fewer training steps, thereby greatly improving the efficiency and reliability of policy optimization.

[0043] The application integrates distribution difference measurement into the model optimization framework, and realizes adaptive calibration of the environment model through construction of a dynamic reward mechanism. This data-driven optimization mode significantly enhances the representation ability of the model for the complex nonlinear dynamics characteristics of the single-leg jumping robot, so that the model can maintain high prediction accuracy and generalization performance under different dynamic working conditions such as terrain and load.

[0044] Based on the high-quality simulation data generated by the precise environment model, the exploration risk in the policy training process is effectively reduced. In real physical experiments, the single-leg jumping robot may fall, collide and cause other safety problems due to improper policy, but the application can avoid potential risks in advance through reliable simulation rehearsal, thereby ensuring the safety of the equipment and significantly improving the feasibility and stability of algorithm deployment. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 A model schematic diagram of a single-leg hopping robot.

[0046] Figure 2 A flowchart schematic diagram of a single-leg hopping robot motion control method based on a coupled flow model.

[0047] Figure 3 A network structure schematic diagram of an environment dynamic model in an embodiment of the present application.

[0048] Figure 4 A structure schematic diagram of a coupled flow model in an embodiment of the present application.

[0049] Figure 5 A schematic diagram of verifying the effectiveness of the coupled flow in a single-leg hopping robot simulation experiment (coupled flow model) in an embodiment of the present application.

[0050] Figure 6 A schematic diagram of verifying the effectiveness of the coupled flow in a single-leg hopping robot simulation experiment (single-leg hopping robot task) in an embodiment of the present application.

[0051] Figure 7 A cumulative average reward comparison graph of using a coupled flow model and other methods in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The present application will be further described below in conjunction with embodiments, but not as a limitation to the present application.

[0053] The embodiment provides a single-leg hopping robot motion control method based on a coupled flow model, which controls the 4 connecting rods and 3 hinge connectors of the single-leg hopping robot (such as shown in Figure 1 ) through the optimal strategy learned, applies torque to the hinge connectors through the strategy, coordinates the bending of the single leg to jump forward, so that the final intelligent agent strategy takes control of the single-leg hopping robot to continuously move forward without falling down, and obtains more rewards. Please refer to Figure 2 , the specific motion control method steps are as follows:

[0054] Step 1, initialize hyperparameters: construct the intelligent agent strategy π θ and the environment dynamic model T ω and the coupled flow model

[0055] The input of the intelligent agent strategy is the 12 state dimensions of the single-leg hopping robot, and the output is the 3 action dimensions of the single-leg hopping robot. The state space composed of 12 state dimensions of the single-leg hopping robot and the action space composed of 3 action dimensions are shown in Table 1 and Table 2, respectively.

[0056] Table 1 State space of the single leg hopping robot

[0057]

[0058] Table 2 Action space of the single leg hopping robot

[0059] Number Action Control minimum value Control maximum value Unit 0 Torque acting on the thigh joint -1 1 Torque (N m) 1 Torque acting on the shank joint -1 1 Torque (N m) 2 Torque acting on the foot joint -1 1 Torque (N m)

[0060] The input of the environment dynamics model is 15 dimensions, which contains 12 state dimensions and 3 action dimensions of the single leg hopping robot, and the output is 13 dimensions of the next state, which contains 12 state dimensions and 1 reward dimension. The input of the coupled flow model is 15 dimensions, which contains 12 state dimensions and 3 action dimensions of the single leg hopping robot, and the output is 1 dimension of the state-action density estimate.

[0061] Agent policy π θ A fully connected neural network is used, the input of which is the state s of the single leg hopping robot with 12 dimensions, and the output is the mean and variance of the corresponding action a under the state s, which has 3 dimensions.

[0062] As shown in Figure 3 , the environment dynamics model T ω is a deep probabilistic neural network composed of three fully connected layers and input and output layers ω (s t+1 a t, ) t ) = N(μ ω (s t, a t ), Each fully connected layer of the network has 256 network nodes. The input of the environment dynamics model is the state-action pair (s t, a t ) at time t, which has 15 dimensions, and the output is the mean μ t+1 (s t, a t ) and variance of the state s ω at the next time with 12 dimensions, and one dimension representing the reward value, a total of 13 dimensions. The samples learned by the environment dynamics model come from the real samples (s t ,a t ,s t+1 ) generated by the actual interaction of the agent policy and the single leg hopping robot, 1≤t≤T, T is the length of the sample trajectory.

[0063] As shown in Figure 4 , the coupled flow model It contains two regularized flow models, each of which is implemented by stacking multiple coupling layers, and each coupling layer is composed of a 3-layer fully connected neural network with 128 neurons. The input is the state-action pair of the real sample (s t, a t ) and the second regularized flow model of the coupled flow model The input is a simulated sample state action pair (s′ t ,a′ t ), the input of the model with two regularized streams has 15 dimensions, and the output is a state-action density estimate.

[0064] Agent strategy π θ , Environmental Dynamics Model T ω and coupled flow models The corresponding network parameters are θ, ω and ∈ respectively.

[0065] The parameters θ, ω, and ∈ of the agent strategy, environment dynamics model, and coupling flow model are uniformly sampled between [-0.1, 0.1] to initialize these parameters. The learning rate of the agent strategy is set to α1 = 0.001, and the environment dynamics model T ω The learning rate of the coupled flow model is α2 = 0.001 and the learning rate of the coupled flow model is α3 =

[0066] 0.0001, the length of the sampling trajectory is H = 3, the maximum iteration plot is N = 300, the maximum training iteration time step is E = 1000, and the initial state distribution of the single-legged jumping robot P(s0) = [0, 1.25, 0, 0, 0, 0, 0, 0, 0, 0] +

[0067] u~[-5e-3,5e-3]. The maximum number of real samples in the experience pool is M1 = 100,000, the minimum number of real samples participating in model training is m = 5,000, the minimum number of simulated samples in the simulated sample pool is M2 = 100,000, and the maximum number of simulated samples saved is M2 + H = 100,003. Set the current iteration plot n = 0 and the current iteration time step e = 0.

[0068] Step 2: Initialize the environment state: set the current time step t = 0, and sample the initial state s0 according to the initial state distribution P(s0) of the single-legged jumping robot.

[0069] Step 3: Generate initial real samples: According to the agent strategy π θ It continuously interacts with the single-legged jumping robot task to generate m real samples and stores them in the real sample experience replay pool.

[0070] Step 4, generate initial simulation samples: according to the agent policy π θ and the environment dynamics model T ω Interactively generate m simulation samples and store them in the simulation sample replay pool.

[0071] Step 5, update the environment dynamics model parameters ω: start the loop episode n, use the samples in the real sample experience replay pool to calculate the loss function L(ω) of the environment dynamics model T ω ,

[0072]

[0073] where (μ ω (s t, a t )+s t ) is the estimate of the next state s t+1 . Then calculate the gradient of the loss function L(ω) with respect to the parameters ω Update the environment dynamics model parameters ω by the gradient descent algorithm SGD

[0074] Step 6, train the coupled flow model: use the real samples in the real sample experience replay pool and the simulation samples in the simulation sample replay pool to calculate the loss function L(∈) of the coupled flow model,

[0075]

[0076] Then calculate the gradient of the loss function L(∈) with respect to the parameters ∈ Update the parameters of the coupled flow model by the gradient descent algorithm SGD

[0077] Step 7, reconstruct the Markov decision process to adaptively fine-tune the environment dynamics model: use the environment dynamics model as a temporary policy, construct a temporary state space S = (s k ,a k ), (0 < = k < T), the temporary state in the temporary state space S is the concatenation of the sample state and the sample action, which has 15 dimensions, where the state-action pair (s k ,a k ) is taken from the real sample experience replay pool; accordingly, construct a temporary action space A = (s k+1 ,r k+1 ), the action in the temporary action space A is the concatenation of the sample state and the sample reward, s k+1 and r k+1 are the predicted next state and reward value. Calculate the temporary reward function of the coupled flow difference measure: The reconstructed Markov decision is (S, A, R, T). The reinforcement learning algorithm SAC is used to maximize the cumulative reward R, so that the distribution of simulated samples generated by the environmental dynamics model is close to the distribution of real samples, thereby fine-tuning the parameters of the environmental dynamics model and achieving accurate modeling of the environmental dynamics model.

[0078] Step 8: Sampling real samples online: Start the cycle time step e, in state s t According to the agent strategy π θ The output π(a t |s t ) probability density sampling action a t , the single-legged jumping robot performs action a t Get the next state s t+1 and reward r t+1 , thus obtaining the real sample (s t ,a t ,s t+1 ,r t+1 ).

[0079] Step 9: Update the real sample experience replay pool: update the current real sample (s t ,a t ,s t+1 ,r t+1 ) is stored in the real sample experience replay pool. When the number of samples in the real sample experience replay pool is less than M1, it is directly added; otherwise, (s t ,a t ,s t+1 ,r t+1 ) replaces the first sample added to the real sample experience replay pool and updates the current state to s t =s t+1 .

[0080] Step 10: Online sampling simulation sample: Randomly select any real sample s from the real sample experience playback pool k ,(0<=k <T),令它为模拟样本的初始状态s ′ 0=s k At t=0, according to the current agent strategy input s ′ 0Get action a ′ 0=π(s ′ 0), and (s ′ 0,a ′ 0) Input to the environmental dynamics model T ω Get the next state and reward s1 ′ ,r1 ′ =T ω (s ′ 0,a ′0), forming simulated sample data (s ′ 0,a ′ 0,s1 ′ ,r1 ′ ) at t = 0; at t = 1, input s1 ′ to the current agent policy to obtain action a ′ 1 = π(s1 ′ ), and input (s1 ′ ,a ′ 1) to the environment dynamics model T ω to obtain the next state and reward s2 ′ ,r2 ′ = T ω (s1 ′ ,a ′ 1); forming simulated sample data (s1 ′ ,a ′ 1,s2 ′ ,r2 ′ ); according to the time step t, repeat the above iteration until t = H-1 to form H-step continuous simulated sample data Forming a simulated trajectory The simulated trajectory comprises a sample trajectory composed of consecutive time steps of simulated samples, and the H-step continuous simulated trajectory comprises H simulated samples.

[0081] Step 11, increasing the simulated samples in the simulated sample pool: repeating step 10, each time increasing H simulated samples into the simulated sample pool until the number of simulated samples in the simulated sample pool is greater than M2.

[0082] Step 12, selecting the optimal simulated sample: according to the loss function of the coupled flow

[0083]

[0084] All simulated trajectories in the simulated sample pool are taken out, and the total loss of each simulated trajectory is calculated, and the simulated trajectories with the smallest top 80% loss values are selected and added to the simulated sample pool according to the calculated loss values, and the simulated trajectories with the 20% larger loss values are removed.

[0085] Step 13, updating the agent policy parameters θ: taking out the real sample in the real sample experience replay pool and the simulated sample in the simulated sample pool, using the reinforcement learning algorithm SAC to calculate the gradient of L(θ) to the parameters θ Using the gradient descent algorithm SGD to update the parameters D KL represents the KL divergence calculation, Q represents the action value function, and V represents the state function.

[0086] Step 14, judge whether the maximum time step E of the episode is reached, if yes, turn to step 15; otherwise, update the current time step e = e + 1, and turn to step 8 to continue execution.

[0087] Step 15, judge whether the maximum number N of the episode is reached, if yes, turn to step 16; otherwise, update the current number n = n + 1, and turn to step 5 to continue execution.

[0088] Step 16, based on the state s t output the action a by the updated agent policy t , that is, according to the optimal parameter θ of the learned agent policy * , generate the optimal policy

[0089] Figure 5 and Figure 6 is a schematic diagram for verifying the effectiveness of the coupling flow in the simulation experiment of the single-leg hopping robot in the embodiment of the application, Figure 7 is a comparison diagram of the cumulative reward in the simulation experiment of the single-leg hopping robot in the embodiment of the application and other methods. From Figure 5 it can be seen that the temporary reward constructed has approached the maximum temporary reward 2 at 5000 real samples, and the greater the temporary reward, the better the fine-tuning effect on the environmental dynamic model; from Figure 6 it can be seen that, under the condition of using the same number of real samples, the cumulative reward obtained by the algorithm with the coupling flow module is the highest; to Figure 7 it can be seen that the MBFlow-Discriminator (the method of the application) and the MBFlow-RL (the simulation sample elimination without step 12) both integrate the coupling flow module, and compared with other mainstream algorithms, under the condition of the same real data samples, the agent policy has higher learning performance and obtains higher cumulative reward. Therefore from Figures 5-7 it can be obtained that the optimal policy learning method for the single-leg hopping robot motion control based on the coupling flow driven model-based reinforcement learning of the application only needs a small amount of real interaction data as a benchmark, and combines the simulation data generated by the model to perform joint optimization, thereby reducing the learning cost of the optimal policy of the agent; the data driven mode of the coupling flow model enhances the representation ability of the model to the complex nonlinear dynamics characteristics of the single-leg hopping robot, and improves the sample efficiency of the agent policy learning.

[0090] The above is only a preferred embodiment of the application, and it should be noted that, for those skilled in the art, without departing from the principles of the application, a number of improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the application.

Claims

1. A motion control method for a single-legged jumping robot driven by a coupled flow model, characterized in that: The steps include: Step 1: Construct and initialize an agent strategy, an environment dynamics model, and a coupled flow model. The agent strategy takes the state dimension of the unipedal jumping robot as input and outputs the action dimension. The environment dynamics model takes the state dimension and action dimension of the unipedal jumping robot as input and outputs the state dimension and reward dimension of the next state of the unipedal jumping robot. The coupled flow model takes the state-action pairs of real samples and the state-action pairs of simulated samples as input and outputs a state-action density estimate. The state-action pairs include the state dimension and the action dimension. Step 2: Set the current time step t = 0 and sample the initial state s0 from the initial state distribution P(s0) of the single-legged jumping robot; Step 3: The agent strategy interacts with the unipedal jumping robot task to generate real samples and store them in the real sample experience replay pool; Step 4: The agent strategy interacts with the environment dynamics model to generate simulation samples and store them in the simulation sample playback pool; From step 5 to step 15, update the parameters of the agent strategy, environment dynamics model, and coupled flow model to the maximum number of episodes: Step 5: Use the real samples in the real sample experience replay pool to update the parameters of the environment dynamics model; Step 6: Update the parameters of the coupled flow model using the real samples in the real sample experience replay pool and the simulated samples in the simulated sample replay pool; Step 7. Reconstruct the Markov decision as (S, A, R, T), and use the reinforcement learning algorithm SAC to maximize the cumulative reward R to update the parameters of the environment dynamics model, where S is the temporary state space, the temporary state in the state space is the splicing of the sample state and sample action in the real sample experience replay pool, A is the temporary action space, and the action in the temporary action space is the splicing of the next state and reward value predicted based on the sample state and sample action. is the temporary reward function of the distribution difference measure, is the first regularized flow model of the coupled flow model, is the second regularized flow model of the coupled flow model, (s k ,a k ) is the state-action pair in the real sample experience replay pool, and T is the sample trajectory length; Update the agent strategy parameters from step 8 to step 14 to the maximum time step: Step 8: With state s t Get action a from the agent strategy t , the single-legged jumping robot performs action a t Get the next state s t+1 and reward r t+1 , construct real samples (s t ,a t ,s t+1 ,r t+1 ); Step 9: The real sample (s t ,a t ,s t+1 ,r t+1 ) into the real sample experience replay pool and update the current state to s t =s t+1 ; Step 10: Randomly select a real sample s from the real sample experience replay pool k As state s t ′ , the agent strategy is based on the state s′ t Get action a′ t , the environment dynamics model obtains the next state s′ according to the action state pair t+1 and reward r′ t+1 , repeat H time steps to get the simulation trajectory Step 11: Simulate the trajectory The included H simulated samples are stored in the simulated sample replay pool, and steps 10 and 11 are repeated until the number of simulated samples in the simulated sample replay pool exceeds the set value; Step 12: Calculate the coupling flow loss of the simulation trajectory, and eliminate some simulation trajectories based on the coupling flow loss; Step 13: Update the agent strategy parameters based on the real samples in the real sample experience replay pool and the simulated trajectories in the simulated sample replay pool; Step 14: After the time steps are accumulated to the maximum time step, the loop from step 8 to step 14 is ended; Step 15: After the number of episodes is accumulated to the maximum number of episodes, the loop from step 5 to step 15 is ended; Step 16: Based on state s t The updated agent strategy outputs action a t .

2. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 1, characterized in that: The calculation method of the coupling flow loss of the simulation trajectory in step 12 is:

3. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 2, characterized in that: In step 12, part of the simulation trajectories are eliminated according to the coupling flow loss, which is to sort the simulation trajectories according to the magnitude of the coupling flow loss and eliminate the simulation trajectories whose coupling flow loss is greater than a set threshold.

4. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 3 is characterized in that: The set threshold is determined based on the coupling flow losses corresponding to the first 80% of the coupling flow losses in ascending order.

5. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 1, characterized in that: The parameters of the agent strategy, environment dynamics model, and coupled flow model are updated using the gradient descent algorithm, where the parameter of the agent strategy is θ and the update formula is: is the gradient of the agent strategy loss function L(θ) with respect to the parameter θ; the parameter of the environment dynamics model is ω, and the update formula is is the gradient of the loss function L(ω) of the environmental dynamics model with respect to the parameter ω; the parameter of the coupled flow model is ∈, and the update formula is is the gradient of the loss function L(∈) of the coupled flow model with respect to the parameter ∈, and α1, α2, and α3 are the set learning rates.

6. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 5, characterized in that: The agent strategy loss function L(θ) is Among them, D KL represents KL divergence calculation, Q represents action value function, and V represents state index function.

7. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 5, characterized in that: The loss function L(ω) of the environmental dynamics model is Among them, (μ ω (s t, a t )+s t ) is the s for the next state t+1 Estimation, μ ω (s t, a t ) represents the state s at the next moment t+1 The mean of s represents the next state t+1 The variance of .

8. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 5, characterized in that: The loss function L(∈) of the coupled flow model is 9. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 1, characterized in that: In step 9, the real sample (s t ,a t ,s t+1 ,r t+1 ) is put into the real sample experience replay pool, when the number of samples in the real sample experience replay pool exceeds the set value, the real sample (s t ,a t ,s t+1 ,r t+1 ) replaces the sample that was first added to the real sample experience replay pool.

10. The method for controlling the motion of a single-legged jumping robot based on coupled flow model driving according to claim 1, characterized in that: After constructing the agent strategy, the environment dynamics model and the coupling flow model in step 1, the parameters of the agent strategy, the environment dynamics model and the coupling flow model are initialized by uniformly distributed sampling between [-0.1, 0.1].