Deep koopman-based online update method of safety controller for robot reinforcement learning

By employing a deep Koopman network and a dimensionality-reduced online update method for the safety controller, this paper addresses the safety issue in the transition from simulation to real-world operation during robot reinforcement learning. This approach improves computational efficiency and applicability, making it suitable for robot safety control in complex environments.

CN120122448BActive Publication Date: 2025-11-25UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510276494.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-11-25
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Existing methods struggle to address safety issues arising from discrepancies between simulation and real-world models in robot reinforcement learning, and existing online update methods suffer from low computational efficiency, making them unsuitable for complex nonlinear robot systems.

Method used

We adopt an online update method for robot safety controllers based on deep Koopman reinforcement learning. By training and dimensionality reduction of a deep Koopman network, we design a safety controller to adapt to model differences and perturbation environments, and use the boosting function and projection matrix for online updates.

Benefits of technology

It improves the safety performance of robot reinforcement learning in complex environments, enhances computational efficiency and applicability, and is suitable for various general-purpose robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122448B_ABST
    Figure CN120122448B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of safety reinforcement learning, and discloses a robot reinforcement learning safety controller online updating method based on deep Koopman, which comprises the following steps: collecting trajectory states of random input control in simulation, training a deep Koopman neural network, obtaining corresponding promotion functions and evolution matrices, adopting an intrinsic orthogonal decomposition method to perform dimension reduction processing on the model to obtain a projection matrix and a new nominal model, performing reinforcement learning strategy migration in a real machine, obtaining an observation error according to the nominal model and a current observation state in interaction, training an online updating network to obtain a residual matrix, and combining the nominal model and the residual model as model constraints of model predictive control to obtain a safety control input. The application can update the safety controller online, improve the safety guarantee performance of reinforcement learning, and is suitable for complex scenes such as model differences in the process of migrating the reinforcement learning strategy from simulation to a real machine and dynamic environments with disturbances in the physical world.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of safe reinforcement learning, in particular to a robot reinforcement learning safety controller online updating method based on deep Koopman. BACKGROUND

[0002] In recent years, reinforcement learning has achieved remarkable results in robot system control with its unique interactive-learning-exploration mechanism, but its inherent trial-and-error mechanism leads to more attention to safety in the physical world. Existing methods reduce safety problems in the training process by simulation to real machine migration, but due to the model difference between simulation and real machine, safety problems still cannot be avoided in the strategy migration process. Therefore, the safety mechanism research of robots in the context of reinforcement learning is a hot spot at present.

[0003] The reinforcement learning safety guarantee method based on safety region restriction is a typical method in safe reinforcement learning, and obtaining a more effective safety region is the core of this method, and the scope of the safety region is closely related to the performance of the safety controller, which has an important influence on the safety guarantee performance of the robot reinforcement learning strategy exploration. At present, many literatures in the context of safe reinforcement learning have also proposed various designs of safety controllers, including PD controller, Lyapunov function controller design, pre-trained reinforcement learning strategy, etc. The process of migrating reinforcement learning strategy from simulation to real machine needs to face the problems caused by model difference; these fixed safety controllers are also difficult to achieve the expected effect due to model difference and disturbance in dynamic environment, and cannot guarantee the safety performance of the robot.

[0004] According to the online updating of the safety controller in the reinforcement learning exploration process, on the one hand, the above problems can be effectively solved, and on the other hand, the safety region can be dynamically expanded, so that the exploration of the robot reinforcement learning task strategy is more sufficient. At present, the controller online updating method can be divided into two categories, one is to directly update the controller according to the online data represented by reinforcement learning, but as mentioned above, the trial-and-error mechanism is contradictory to safety guarantee; the second is to update the model according to the online data, and then design the controller according to the model, but the complex nonlinear characteristics of the robot lead to a high computational burden in the solution of the nonlinear controller. Koopman operator shows great potential in alleviating the difficulty of nonlinear control under unknown dynamics, which maps nonlinear dynamics to a high-dimensional linear space, and a linear controller can be designed accordingly. Existing researches realize online controller updating by incrementally updating the dynamic model through the iterative least squares method. However, there are still problems in determining the appropriate high-dimensional linear lifting space dimension and the computational efficiency of the controller update in the high-dimensional space, and the application of the safety reinforcement learning field of complex robot systems has not been well studied. SUMMARY

[0005] To solve the above technical problems, the application provides a deep Koopman-based robot reinforcement learning safety controller online updating method, which is applied to safety guarantee in the robot reinforcement learning strategy migration process. The model dimension reduction method provided by the application avoids subjective selection of high-dimensional promotion space dimensions, and can improve the calculation efficiency of online updating of the safety controller, and is suitable for various general robots. The deep Koopman loss function used in the application is designed, more model matching factors are considered, and the learning accuracy and model prediction effect can be effectively improved. The online updating safety controller method provided by the application can be applied to complex scenes such as model difference in the process of migrating the reinforcement learning strategy from simulation to real machine and dynamic environment with disturbance in the physical world, and has important significance for improving the safety performance of robot reinforcement learning, and significantly improves the applicable scene and practicability of safety reinforcement learning.

[0006] To solve the above technical problems, the application adopts the following technical solutions:

[0007] A deep Koopman-based robot reinforcement learning safety controller online updating method, comprising:

[0008] Randomly generating a control input sequence of the robot in a simulation environment, and collecting a corresponding state set of the robot;

[0009] Inputting the control input sequence and the state set of the robot into a deep Koopman network for training, optimizing the deep Koopman network parameters through a loss function containing reconstruction error, prediction error and promotion state error, obtaining a promotion function from a state space where the robot state is located to a linear promotion space and a Koopman operator space evolution matrix;

[0010] Mapping the state of the robot to the linear promotion space based on the promotion function to form a data matrix, and obtaining a projection matrix and a linear evolution equation matrix after dimension reduction through dimension reduction processing;

[0011] Constructing a linear model with the linear evolution equation matrix as a constraint, predicting a control optimization problem, adjusting parameters through simulation, and designing a safety controller with a safety region equilibrium point as a target;

[0012] Migrating the reinforcement learning strategy in the offline simulation environment to an actual robot as an initial strategy, and deploying the safety controller; obtaining a control input of a current time step through the reinforcement learning strategy, and the robot executes the control input and interacts with the environment;

[0013] When the safety controller judges that the current robot state is safe, the update of the reinforcement learning strategy is performed: the state of the robot is collected in real time, the observation state of the robot is calculated through an improvement function and a projection matrix, the state error is predicted based on a nominal model composed of a linear evolution equation matrix, a residual matrix is obtained, and an online nominal model is updated;

[0014] When the safety controller judges that the current robot state is not safe, the safety controller is started to solve the control input until the safe state is returned, and then the update of the reinforcement learning strategy is performed.

[0015] In one of the embodiments, the control input sequence of the robot is randomly generated in the simulation environment, and the corresponding state set of the robot is collected, specifically including:

[0016] The control input sequence is randomly generated to form a control input sequence set U = {U1, U2, …, U i ,…,U n}, U i is the i-th control input sequence of the robot, U i = {u1, u2, …, u k ,…,u T}, u k is the k-th control input in U i ;

[0017] The set X = {X1, X2, …, X i ,…,X n} of n trajectories of the robot is collected, X i is the set of states of the robot in the i-th trajectory, X i = {x1, x2, …, x k ,…,x T}, x k is the k-th robot state in X i .

[0018] In one of the embodiments, the improvement function from the state space where the robot state is located to the linear improvement space and the Koopman operator space evolution matrix are obtained, specifically including:

[0019] The state of the robot in the linear improvement space is defined as where x k is the k-th robot state, and the improvement function f N is composed of a fully connected deep neural network, I represents a unit matrix, the dimension of I is the same as that of x k ; and the state estimation of the original state space reconstructed according to the linear improvement space is Wherein, the evolution matrix C = [I0]; the evolution equation of Koopman operator space can be expressed as:

[0020]

[0021] represents the prediction of the state at time k+1 in the high-dimensional linear space through the linear evolution equation; the evolution matrices A, B are the weights of the deep neural network, and the fully connected neurons with linear activation function are used to simulate the linear Koopman dynamics;

[0022] After the training is completed, the deep Koopman network will obtain the lifting function and the evolution matrices A, B, C of the Koopman operator space.

[0023] In one embodiment, the deep Koopman network parameters are optimized by the loss function comprising the reconstruction error, the prediction error and the lifting state error, and specifically comprising:

[0024] The loss function loss of the deep Koopman network is as follows:

[0025] loss = α1L rec + α2L pred + α3L lift + β1‖W‖1+ β2‖W‖2.

[0026] Wherein, α1, α2, α3 are weight coefficients, ‖·‖1 and ‖·‖2 are respectively the l1 norm and the l2 norm of the matrix, and β1, β2 are regularization coefficients;

[0027] Reconstruction error Prediction error Lifting state error

[0028] Wherein, S represents the prediction time step, and τ is a decay factor used to weight the prediction error and the lifting state error.

[0029] In one embodiment, the deep Koopman network adopts an autoencoder structure.

[0030] In one embodiment, the state of the robot is mapped to the linear lifting space based on the lifting function to form a data matrix, and specifically comprising:

[0031] According to the lifting function and the control input sequence set U and the trajectory set X, the data matrix Z = [z1, z2, …, z p ] of the lifting space is obtained, wherein is the state x of the robotk In the corresponding state of the lifting space, p is the length of the data matrix.

[0032] In one embodiment, the projection matrix and the linear evolution equation matrix after dimension reduction are obtained by dimension reduction processing, specifically including:

[0033] Map the data matrix to the dimension reduction space through the projection matrix P:

[0034] y k =P T z k ;

[0035] z k is the state x k of the robot in the high-dimensional linear space; k is the corresponding state y k in the lifting space;

[0036] Map back from the dimension reduction space to the original space:

[0037]

[0038] is the state x k mapped back from the dimension reduction space y - to the high-dimensional linear space;

[0039] Based on this, the matrix corresponding to the linear evolution equation after dimension reduction is obtained The linear evolution equation after dimension reduction is:

[0040]

[0041] In one embodiment, the linear model is constructed with the linear evolution equation matrix as a constraint, the control optimization problem is predicted, the parameters are adjusted through simulation, and the safety controller is designed with the safety region equilibrium point as the target, specifically including:

[0042] The control optimization problem is:

[0043]

[0044] Where M is the length of the prediction window, is the desired state; x - , x + and u - , u + are the upper and lower bounds of the state constraints and control constraints, respectively; Q and R are cost matrices used to adjust the effect of linear model predictive control.

[0045] In one embodiment, the real-time acquisition robot state, the observation state of the robot is calculated through the lifting function and the projection matrix, the state error is predicted based on the nominal model composed of the linear evolution equation matrix, the residual matrix is obtained and the linear nominal model is updated, and specifically comprises:

[0046] Real-time sampling of the current robot state x k , the observation state is obtained through the lifting function and the projection matrix P According to the current nominal model u k is the corresponding control input of x k ;

[0047] The observation state error Δy k is obtained through , input into the online update network, and trained according to the designed online update network loss to obtain the residual matrix ΔA, ΔB, and update the nominal model:

[0048]

[0049] In one embodiment, the online update network loss is L adapt :

[0050] L adapt = L ε + γ1‖ΔA‖1+ γ2‖ΔB‖1+ γ3‖ΔA‖2+ γ4‖ΔB‖2

[0051] Wherein, γ1, γ2, γ3, γ4 are regularization parameters; the loss term L ε quantifies the influence of observation error:

[0052]

[0053] Δy k represents the observation state error, which is obtained from .

[0054] Compared with the prior art, the beneficial technical effects of the present application are:

[0055] 1. The present application proposes a robot safety controller online updating method based on deep Koopman, which can be applied to the complex scenes such as model difference in the process of migrating simulation to real machine of reinforcement learning strategy, dynamic environment with disturbance in physical world, etc., and has important significance for improving the safety performance of robot reinforcement learning, and significantly improves the applicable scene and practicality of safety reinforcement learning.

[0056] 2. The model dimension reduction method avoids subjective selection of high-dimensional promotion space dimensions, improves the online update calculation efficiency of the safety controller, more efficiently performs safety control, and is suitable for various general robots.

[0057] 3. The new deep Koopman loss function considers more model matching factors, and can effectively improve the learning accuracy and model prediction effect. BRIEF DESCRIPTION OF DRAWINGS

[0058] Figure 1 The method flowchart in the embodiment of the application is shown in the figure.

[0059] Figure 2 The offline training flowchart of the robot safety controller based on deep Koopman in the embodiment of the application is shown in the figure.

[0060] Figure 3 The flowchart of the online update method of the robot safety controller based on deep Koopman in the embodiment of the application is shown in the figure. DETAILED DESCRIPTION

[0061] A preferred embodiment of the application will be described in detail below with reference to the accompanying drawings.

[0062] The robot safety controller based on online update of the deep Koopman network in the application improves the safety in the reinforcement learning strategy migration process, and further perfects the reinforcement learning safety guarantee mechanism based on the safety region limitation. The model dimension reduction method in the application avoids the difficulty of subjective selection of high-dimensional promotion space dimensions, and improves the online update calculation efficiency of the safety controller.

[0063] As shown in the figure, the online update method of the robot reinforcement learning safety controller based on deep Koopman in the application includes the following steps: Figure 1 S1: Randomly generating a control input sequence of the robot in a simulation environment, and collecting a corresponding state set of the robot.

[0064] S2: Inputting the control input sequence and the state of the robot into the deep Koopman network for training, optimizing the deep Koopman network parameters through a loss function containing reconstruction error, prediction error and promotion state error, obtaining a promotion function from the state space where the state of the robot is located to the linear promotion space and a Koopman operator space evolution matrix.

[0065] S3: Mapping the state data to the linear promotion space based on the promotion function to form a data matrix, and obtaining a projection matrix and a linear evolution equation matrix after dimension reduction through dimension reduction processing.

[0066] ​

[0067] S4: constructing a linear model with the linear evolution equation matrix as a constraint, predicting and optimizing the control problem, adjusting parameters through simulation, and designing a safety controller with the safety region equilibrium point as a target.

[0068] S5: migrating the reinforcement learning strategy in the offline simulation environment to the actual robot as an initial strategy, deploying the safety controller; obtaining the control input of the current time step through the reinforcement learning strategy, and the robot executing the control input and interacting with the environment.

[0069] S6: when the safety controller determines that the current robot state is safe, updating the reinforcement learning strategy: collecting the state of the robot in real time, calculating the observation state of the robot through the lifting function and the projection matrix, predicting the state error based on the nominal model composed of the linear evolution equation matrix, obtaining the residual matrix and updating the linear nominal model.

[0070] S7: when the safety controller determines that the current robot state is not safe, starting the safety controller to solve the control input until returning to the safe state, and then updating the reinforcement learning strategy.

[0071] In one of the embodiments, the control input sequence of the robot is randomly generated in the simulation environment in step S1, and the corresponding state set of the robot is collected, which specifically includes:

[0072] The control input sequence is randomly generated to form a control input sequence set U={U1, U2, …, U i ,…,U n}, U i is the u-th control input sequence of the robot, U i ={i1, i2, …, i k ,…,i T}, i k is the k-th control input in U i ;

[0073] The set X={X1, X2, …, X i ,…,X n} of n trajectories of the robot is collected, X i is the set of states of the robot in the i-th trajectory, X i ={x1, x2, …, x k ,…,x T}, x k is the k-th robot state in X i , and T is the length of the trajectory.

[0074] In one of the embodiments, the lifting function from the state space where the robot state is located to the linear lifting space and the Koopman operator space evolution matrix are obtained in step S2, which specifically includes:

[0075] The state of the robot in the linear lifted space is defined as where x k is the kth robot state, the lifting function f N is composed of a fully connected deep neural network; the state estimate of the original state space reconstructed from the linear lifted space is where the evolution matrix C = [I0]; the evolution equation of the Koopman operator space can be expressed as:

[0076]

[0077] The evolution matrices A, B are the weights of the deep neural network, and the fully connected neurons with linear activation functions are used to simulate linear Koopman dynamics;

[0078] After the training is completed, the deep Koopman network will obtain the lifting function and the evolution matrices A, B, C of the Koopman operator space.

[0079] In one embodiment, the step S2 of optimizing the deep Koopman network parameters by the loss function comprising the reconstruction error, the prediction error and the lifted state error, specifically comprises:

[0080] The loss function loss of the deep Koopman network is as follows:

[0081] loss = a1L rec + a2L pred + a3L lift + b1‖W‖1+ b2‖W‖2;

[0082] where a1, a2, a3 are weight coefficients, ‖·‖1 and ‖·‖2 are the l1 norm and the l2 norm of the matrix respectively, and b1, b2 are regularization coefficients;

[0083] The reconstruction error The prediction error The lifted state error

[0084] where S represents the time step of prediction, and t is a decay factor used to weight the prediction error and the lifted state error.

[0085] Specifically, the deep Koopman network adopts an autoencoder structure.

[0086] Since the dimension of the Koopman operator linear space is infinite, but in practical applications, it can only be truncated to a finite dimension, which is closely related to the accuracy and calculation speed of the model. The higher the dimension, the more accurate the model description, but the slower the corresponding calculation speed. Selecting the appropriate dimension of the linear lifting space is also a challenge. The present application reduces the difficulty of selecting the output dimension of the lifting function. In the offline training part, the dimension is uniformly lifted to a high enough dimension. Since this operation is done offline, the calculation time requirement is not strict. The subsequent operation needs to be done in real time, and dimension reduction will be adopted to reduce the calculation time cost.

[0087] In one of the embodiments, the step S3 of mapping the state data to the linear lifting space based on the lifting function to form the data matrix comprises:

[0088] According to the lifting function and the control input sequence set U and the trajectory set X, the data matrix Z of the lifting space is obtained, Z = [z1, z2, …, z p ], wherein is x k In the corresponding state of the lifting space, p is the length of the data matrix.

[0089] In one of the embodiments, the step S3 of obtaining the projection matrix and the linear evolution equation matrix after dimension reduction by dimension reduction processing comprises:

[0090] The data matrix is mapped to the dimension reduction space by the projection matrix P: y k = P T z k ;

[0091] The data matrix is mapped to the dimension reduction space by the projection matrix P: y k = P T z k ;

[0092] Based on this, the matrix corresponding to the linear evolution equation after dimension reduction is obtained The linear evolution equation after dimension reduction is:

[0093]

[0094] Specifically, since the output layer dimension of the deep neural network is very high, there is a problem of high time cost in real-time calculation. To solve this difficulty, balance the calculation time and model accuracy, and improve the calculation efficiency, a dimension reduction method is adopted in this step to simplify the model.

[0095] In one of the embodiments, the step S4 of constructing a linear model with the linear evolution equation matrix as a constraint to predict the control optimization problem, adjusting the parameters by simulation, and designing a safety controller with the safety region equilibrium point as the target comprises:

[0096] The control optimization problem is:

[0097]

[0098] Where M is the length of the prediction window, It is the desired state; x - ,x + and u - ,u + These are the upper and lower bounds of the state constraints and control constraints, respectively; Q and R are cost matrices used to adjust the effectiveness of predictive control in linear models.

[0099] Specifically, the design of a safety controller for a linear model, employing model predictive control, can be described as an online constrained optimization problem. Based on the definition of the safe region, this optimization problem aims at a locally asymptotically stable equilibrium point x. * Using the target point as the objective and the reduced-dimensional linear model, along with the control input range and state range as constraints, the control input is obtained by solving an optimization problem to achieve safety correction. To adjust the predictive control parameters of this optimization model, the solution can be performed multiple times with randomly assigned initial states, and the resulting control input can be obtained from the solution. This method is applied to a simulated robot system to optimize relevant parameters based on its correction performance. After solving for the optimal U-sequence, only the first control input is considered. Input into the robot system.

[0100] In one embodiment, step S6 involves real-time acquisition of the robot's state, calculation of the robot's observed state using a boosting function and a projection matrix, prediction of the state error based on a nominal model constructed from the linear evolution equation matrix, obtaining the residual matrix, and updating the linear nominal model. Specifically, this includes:

[0101] Real-time sampling of the current robot state x k By lifting function The observed state is obtained from the projection matrix P. According to the current nominal model The estimated state is calculated. For x k Corresponding control inputs;

[0102] pass The observed state error Δy is obtained. k The input is fed into the online update network, trained according to the designed loss function of the online update network, and the residual matrices ΔA and ΔB are obtained to update the nominal model:

[0103]

[0104] In one embodiment, the online update network loss is L. adapt :

[0105] L adapt =L ε +γ1‖ΔA‖1+γ2‖ΔB‖1+γ3‖ΔA‖2+γ4‖ΔB‖2;

[0106] Where γ1, γ2, γ3, and γ4 are regularization parameters; the loss term L ε The impact of observation errors was quantified:

[0107]

[0108] Online network updates correct the discrepancy between the offline nominal model and the real system by refining the Koopman model online in real time. During online updates, only networks A and B are updated; the lifting function is not updated. The structure of the online-updated evolutionary matrix network is the same as that in step S3, consisting of a linear activation function and a fully connected neural network, and learns the residual matrices ΔA and ΔB.

[0109] The offline training process for the safety controller of this invention is described in [reference needed]. Figure 2 The online update method for the safety controller of this invention is described in [reference needed]. Figure 3 .

[0110] In one embodiment of the invention, a drone is selected as the robotic system. The safe area based on the domain of attraction (ROA) is defined as: D = {x0∈X|lim t→∞ Φ(t;x0)=0}, where X represents the original state space of the UAV, x0 is the current state of the UAV, and Φ(t;x0) represents the state of the UAV under the given safety control π. safe (x) is the trajectory of the UAV starting from the current state x0. Given a safety control π... safe If (x) can achieve state equilibrium, then the current state is within the safe zone. This invention will focus on the design method of the safety controller, making it applicable to complex scenarios such as model differences during the transfer of reinforcement learning policies from simulation to real-world operation, and dynamic environments with disturbances in the physical world. When the UAV's current state is within the safe zone D, the current state is considered safe, and the reinforcement learning control policy π continues to be trained. RL (x), otherwise, the current state is considered dangerous, and the safety controller π is used. safe (x) moves the current state back into the safe zone. The security reinforcement learning framework is shown below:

[0111]

[0112] Where u represents the control strategy, p represents the safety probability threshold, and P(x∈D) represents the safety probability of the current state x.

[0113] To obtain the initial safety controller, in the simulation, a set of control input sequence set U = {U1, U2, …, U i ,…,U n} is randomly generated, and is input into the unmanned aerial vehicle system in turn, and unmanned aerial vehicle data collection is performed. A set of n trajectories X = {X1, X2, …, X i ,…,X n} is collected, where U i is the i-th control input sequence, U i = {u1, u2, …, u k ,…,u T}. X i is the state set of the i-th trajectory, X i = {x1, x2, …, x k ,…,x T}. Here, the application specifically sets n = 125, T = 600, and the trajectory sampling interval is 0.01 s. The control input sequence set U and the trajectory set X are divided into a training set and a test set in a ratio of 0.8:0.2, and are input into the constructed deep Koopman network.

[0114] The deep Koopman network adopts an autoencoder structure, and the specific network parameters are shown in Table 1. The evolution equation of the high-dimensional linear Koopman operator space can be expressed as:

[0115]

[0116] where f N is composed of a fully connected deep neural network, where C = [I0]. Note that the evolution matrix A, B is the weight of the neural network, and a fully connected neuron with a linear activation function is used to simulate linear Koopman dynamics. The deep Koopman loss function is designed as follows:

[0117] loss = α1L rec + α2L pred + α3L lift + β1‖W‖1+ β2‖W‖2

[0118] where α1, α2, α3 are weight coefficients, ‖·‖1 and ‖·‖2 are the l1 norm and the l2 norm of the matrix respectively, β1, β2 are the regularization coefficients thereof, and the corresponding parts are:

[0119]

[0120] Here, S = 15 represents the predicted time step, and τ = 0.9 is the decay factor used to weight the prediction error and the error of the lifted state. When the deep Koopman network training is completed, the lifting function and the evolution matrices A, B, C of the UAV Koopman operator space are obtained. The specific setting parameters are shown in Table 1.

[0121] Table 1 Koopman auto-encoding network

[0122]

[0123] According to the obtained lifting function and the control input sequence set U and the trajectory set X, the data matrix Z = [z1, z2, …, z p ] of the lifting space is obtained. The slice in the UAV state evolution process in Z has a time correlation, and here p = 1000 is selected as the data matrix length. The model dimension reduction processing of Z is carried out by proper orthogonal decomposition (POD), and according to the accuracy requirement, the number of main modes is determined to be 18. The original system is mapped to the reduced dimension system through the projection matrix P ∈ R 40×18 :

[0124] y k = P T z k ;

[0125] The reduced dimension system is mapped back to the original system:

[0126]

[0127] Based on this, the matrix corresponding to the linear evolution equation with lower dimension is obtained

[0128] Therefore, the linear evolution equation with lower dimension is expressed as:

[0129]

[0130] Combined with A, B, C and the lifting function the linear evolution equation with lower dimension is finally obtained After obtaining the above model, the parameter tuning of model predictive control is carried out in simulation. The initial state is randomly given, and the on-line solved constrained optimization problem is as follows:

[0131]

[0132]

[0133] Wherein, the specific parameters are shown in Table 2.

[0134] Table 2 Model Predictive Control

[0135]

[0136] With the model and control method as the safety controller of the unmanned aerial vehicle, and starting the unmanned aerial vehicle reinforcement learning strategy learning, the application focuses on the reinforcement learning strategy training, and therefore only simple trajectory tracking task training is performed. The safety controller is used for safety estimation, and the safety estimation method can be selected here, and the application uses the simplest grid method for estimation. Since it is not the focus, it will not be described in detail. When the reinforcement learning strategy training is completed, the strategy migration is started.

[0137] Reinforcement Learning Strategy Pi RL (x) Migrate to the actual unmanned aerial vehicle, deploy the safety controller. The model obtained in the simulation as a nominal model. While interacting with the environment, the actual state of the current unmanned aerial vehicle is sampled in real time Through and the projection matrix P, the observed state is obtained According to the nominal model, the estimated state is calculated Through the observed state error Ay is obtained k and input into the online update network.

[0138] The update network Delta A, Delta B is consistent with the structure of A, B, C, and is composed of a linear activation function and a fully connected neural network. The corresponding loss function loss is designed as:

[0139] L adapt = L ε + gamma1 || Delta A || 1 + gamma2 || Delta B || 1 + gamma3 || Delta A || 2 + gamma4 || Delta B || 2;

[0140] Wherein, gamma1, gamma2, gamma3, gamma4 are regularization parameters, which are generally set to be larger than beta1, beta2, and the specific value setting can be seen in

[0141] Table 3. L ε quantifies the influence of the observation error, which is specifically described as:

[0142]

[0143] According to the designed loss, the online updated model residual matrix Delta A, Delta B is obtained, the nominal model is updated, and the new nominal model is

[0144] Table 3 Online Update Network Parameters

[0145]

[0146] Each interaction, the current state is judged by the safety controller, when safe, continue to strengthen the learning policy transfer, explore unmanned aerial vehicle trajectory tracking strategy. When the safety controller judges that the current state is unsafe, according to the current corrected unmanned aerial vehicle nominal model And Adopt model predictive control, solve the control input u k , acting on the unmanned aerial vehicle system, realize safety control. At the same time, the current nominal model is updated online, so that the nominal model gradually tends to the real model. Through the adjustment of the model, the safety controller is continuously optimized, which can adapt to the dynamic environment with model difference and disturbance in the physical world, realize the continuous expansion of the safety region, and increase the scope of reinforcement learning strategy exploration.

[0147] It should be understood that although each step in the flowchart of the accompanying drawings is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps has no strict sequence limitation, and these steps can be executed in other order. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.

[0148] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and therefore all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0149] In addition, it should be understood that although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the description manner of the specification is only for the sake of clarity, those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be properly combined to form other embodiments that those skilled in the art can understand.

Claims

1. A deep Koopman-based online update method for safety controller of robot reinforcement learning, characterized in that, The application relates to a method for constructing a safety controller for a robot. Randomly generating a control input sequence of a robot in a simulation environment and collecting a corresponding state set of the robot; Inputting the control input sequence and the state set of the robot into a deep Koopman network for training, optimizing parameters of the deep Koopman network through a loss function containing a reconstruction error, a prediction error and a lifting state error, obtaining a lifting function from a state space where a state of the robot is located to a linear lifting space and a Koopman operator space evolution matrix; Mapping the state of the robot to the linear lifting space based on the lifting function to form a data matrix, and obtaining a projection matrix and a linear evolution equation matrix after dimension reduction through dimension reduction processing; Constructing a linear model with the linear evolution equation matrix as a constraint, predicting a control optimization problem, adjusting parameters through simulation, and designing a safety controller with a safety region equilibrium point as a target; Migrating a reinforcement learning strategy in an offline simulation environment to an actual robot as an initial strategy and deploying the safety controller; Obtaining a control input at a current time step through the reinforcement learning strategy, and executing the control input by the robot and interacting with the environment; When the safety controller judges that the current state of the robot is safe, updating the reinforcement learning strategy: collecting the state of the robot in real time, calculating an observed state of the robot through the lifting function and the projection matrix, predicting a state error based on a nominal model formed by the linear evolution equation matrix, obtaining a residual matrix and updating an online nominal model; When the safety controller judges that the current state of the robot is not safe, starting the safety controller to solve the control input until returning to the safe state, and then updating the reinforcement learning strategy.

2. The online updating method of the robot reinforcement learning safety controller based on deep Koopman according to claim 1, wherein: The method specifically comprises the following steps. Randomly generate control input sequences to form a control input sequence set U = {U1, U2, ..., U...} i ,…,U n }, U i U is the i-th control input sequence for the robot. i ={u1,u2,…,u k ,…,u T }, u k For U i The k-th control input; A set of n trajectories X = {X1, X2,..., X i} of a robot is collected, where Xi is the set of robot states in the i-th trajectory, Xi = {x1, x2,..., x n}, and xk is the k-th robot state in Xi. i The robot state is represented as a vector of joint angles and joint velocities. i The robot state is represented as a vector of joint angles and joint velocities. k The robot state is represented as a vector of joint angles and joint velocities. T The robot state is represented as a vector of joint angles and joint velocities. k The robot state is represented as a vector of joint angles and joint velocities. i The robot state is represented as a vector of joint angles and joint velocities.

3. The online updating method of the robot reinforcement learning safety controller based on deep Koopman according to claim 1, characterized in that: The method specifically comprises the following steps. The state of the robot in the linear lifting space is defined as where x k is the kth state of the robot, the lifting function f N is composed of a fully connected deep neural network, I represents an identity matrix, the dimension of I is the same as that of x k ; the state estimation of the original state space reconstructed according to the linear lifting space is where the evolution matrix C = [I0]; the evolution equation of the Koopman operator space can be expressed as: represents the prediction of the state at time k+1 by the linear evolution equation in a high-dimensional linear space; the evolution matrices A, B are the weights of the deep neural network, which simulates linear Koopman dynamics using fully connected neurons with linear activation functions; After training, the deep Koopman network will have learned the lifting function as well as the evolution matrices A, B, C of the Koopman operator space.

4. The online updating method of the robot reinforcement learning safety controller based on deep Koopman according to claim 3, characterized in that: The method specifically comprises the following steps. The loss function loss of the deep Koopman network is as follows: loss = a1L rec + a2L pred + a3L lift + b1||W||1+ b2||W||2; Wherein, alpha1, alpha2 and alpha3 are weight coefficients, ‖·‖1 and ‖·‖2 are l1 norm and l2 norm of a matrix respectively, beta1 and beta2 are regularization coefficients. reconstruction error prediction error elevation state error Wherein, S represents a predicted time step, and tau is a decay factor used for weighting the prediction error and the lifting state error.

5. The online updating method of the robot reinforcement learning safety controller based on deep Koopman according to claim 1, characterized in that: The deep Koopman network adopts an automatic encoder structure.

6. The online updating method of a robot reinforcement learning safety controller based on deep Koopman according to claim 1, characterized in that: The method specifically comprises the following steps. According to the lifting function and the control input sequence set U, the trajectory set X, the data matrix Z = [z1, z2,..., z p ] of the lifting space is obtained, wherein is the state x k of the robot p is the length of the data matrix at the corresponding state of the lifting space.

7. The online updating method of a robot reinforcement learning safety controller based on deep Koopman according to claim 1, characterized in that: The method specifically comprises the following steps. The data matrix is mapped to a dimension reduction space through the projection matrix P: y k = P T z k ; z k state x of the robot k corresponding state y in the lifted space k state z of the high-dimensional linear space k state projected into the reduced-dimensional space The data matrix is mapped to a dimension reduction space through the projection matrix P: To map back from the reduced dimension space state y k to the state in the high dimension linear space; Based on this, the matrix corresponding to the reduced dimension linear evolution equation is obtained The reduced dimension linear evolution equation is:

8. The online updating method of a robot reinforcement learning safety controller based on deep Koopman according to claim 1, characterized in that: The method specifically comprises the following steps. The control optimization problem is as follows: where M is the length of the prediction window, is the desired state; x - is the state constraint; x + and u - is the control constraint; u + are the upper and lower bounds of the state and control constraints, respectively; Q, R are cost matrices used to tune the effect of the linear model predictive control.

9. The online updating method of a robot reinforcement learning safety controller based on deep Koopman according to claim 1, characterized in that: The real-time acquisition robot state, through the lifting function and the projection matrix calculation robot observation state, based on the linear evolution equation matrix constitutes the nominal model to predict the state error, get the residual matrix and update the linear nominal model, specifically including: Sampling the current robot state x in real-time k The observed state is obtained by lifting the function and the projection matrix P The estimated state is computed from the current nominal model k u k is the corresponding control input By Obtaining observation state error Ay k , input into the online update network, trained according to the designed online update network loss, to obtain residual matrix ΔA, ΔB, and update the nominal model:

10. The online updating method of the deep Koopman-based robot reinforcement learning safety controller according to claim 9, characterized in that: The online update network loss is L adapt : L adapt = L ε + γ1‖ΔA‖1+ γ2‖ΔB‖1+ γ3‖ΔA‖2+ γ4‖ΔB‖2 where γ1, γ2, γ3, γ4are regularization parameters; the loss term L ε quantifies the impact of the observation error: Δy k represents the observation state error, which is given by .

Citation Information

Patent Citations

  • Iterative learning control optimization method for direct current motor with non-repetitive disturbance

    CN119511712A

  • Model-based reinforcement learning

    US20240320505A1