A soft robot control method fusing attention mechanism and physics-inspired neural network
By integrating the attention mechanism with physics-inspired neural networks, the problem that traditional dynamic modeling methods are difficult to describe the behavior of soft robots in complex environments is solved, and efficient and accurate soft robot dynamic modeling is achieved, which is suitable for various configurations and motion conditions.
Patent Information
- Application Number
- CN202411772144.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-12-04
AI Technical Summary
Traditional dynamic modeling methods have difficulty describing the behavior of soft robots in complex environments, while existing physics-inspired neural networks have high training time costs and large data requirements when faced with complex soft robotic systems.
The attention mechanism is integrated with physics-inspired neural networks (PINNs) for soft robot dynamics modeling. By introducing an attention layer into the multi-layer neural network, the internal structural relationship of the input data is learned, the model training process is optimized, the training time is reduced, and the prediction accuracy is improved.
It achieves efficient dynamic modeling under different configurations and motion conditions, simplifies the physical model construction steps, improves the model's long-term prediction accuracy and anti-interference ability to environmental noise, and reduces training time and data requirements.
Smart Images

Figure CN119388440B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a soft robot dynamics modeling method that integrates an attention mechanism with a physics-inspired neural network. The present invention belongs to the field of soft robot dynamics modeling. Background Art
[0002] With the widespread application of robotics in industry and daily life, soft robots, due to their unique flexibility and adaptability, have shown broad application potential in numerous fields. They can be used for tasks such as operating in complex environments, assisting with precision surgery, disaster relief, and personalized care, achieving efficient control and control. Dynamic modeling is a core component of soft robotics technology and the foundation for precise operation. Due to the high nonlinearity and complex degrees of freedom introduced by the structural and material properties of soft robots, traditional dynamic modeling methods struggle to accurately capture their dynamic characteristics. Therefore, developing methods that accurately simulate their behavior can help improve the operational precision and adaptability of soft robots. Furthermore, reducing modeling errors and improving prediction accuracy can enhance the efficiency of soft robots in complex environments, facilitating their application in a wider range of scenarios. Therefore, optimizing the accuracy and efficiency of dynamic modeling for soft robots is of great significance. Traditional modeling methods require a detailed dynamic model of the soft robot system. However, the complex coupling between the kinematics and dynamics of soft robots makes dynamic models for multi-degree-of-freedom soft robots extremely complex. Furthermore, the nonlinear properties of materials and the influence of environmental factors further complicate obtaining accurate dynamic models, making traditional modeling methods ineffective for highly flexible soft robots. As an advanced deep learning framework, physics-inspired neural networks acquire information through the interaction between models and data and optimize predictive models under the guidance of physical laws. Using this approach for soft robot dynamics modeling allows for the direct introduction of physical constraints into network training without relying entirely on traditional dynamics modeling processes. This makes the optimization objective more intuitive and the model framework broadly applicable to various types of soft robotic systems. However, accurately simulating the complex dynamic behavior of soft robots requires sufficient training data and meticulous optimization of the model's backpropagation and gradient descent processes, making the overall training process extremely complex. While directly employing physics-based deep learning models can improve prediction accuracy, this approach still requires a large number of training samples and a long training time, which can be inefficient in practical applications and thus limit its widespread application.
[0003] In summary, traditional dynamic modeling methods rely on precise physical equations and are therefore unable to describe the behavior of soft robots in complex environments. Existing physics-inspired neural networks, when faced with complex soft robotic systems, have high training costs and extremely high data requirements. Summary of the Invention
[0004] The purpose of this invention is to solve the problem that traditional dynamic modeling methods rely on precise physical equations and are difficult to describe the behavior of soft robots in complex environments. The existing physics-inspired neural networks have high training time costs and extremely high data requirements when facing complex soft robot systems. A soft robot dynamic modeling method that integrates attention mechanism and physics-inspired neural networks (PINNs) is proposed.
[0005] A soft robot control method that combines attention mechanism and physics-inspired neural network. The specific process is as follows:
[0006] Step 1: Create a soft robot training dataset; the specific process is as follows:
[0007] The soft robot training dataset includes the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot;
[0008] The inherent properties of the soft robot are: number of soft robot arm segments n, soft robot length, soft robot width, soft robot height, soft robot weight, and soft robot density;
[0009] The energy of the soft robot is: the speed of the soft robot and the acceleration of the soft robot;
[0010] Step 2: Build a deep learning network model;
[0011] The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as the input of the deep learning network model, and the mass matrix function M(q; θ M ) as the output of the deep learning network model;
[0012] The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as inputs of the deep learning network model, and the dissipation matrix function As the output of the deep learning network model;
[0013] The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as the input of the deep learning network model, and the energy matrix function U(q,u;θ U ) as the output of the deep learning network model;
[0014] The mass matrix function M(q;θ) output by the Lagrangian function and deep learning network model M ), dissipative matrix function Energy matrix function U(q,u;θ U), obtain the predicted position, velocity and acceleration of the soft robot at the next moment;
[0015] The difference between the predicted acceleration of the soft robot at the next moment and the true value of the acceleration is used as the error term of the cross-entropy loss function. The cross-entropy loss function is calculated until the loss function converges to obtain a trained deep learning network model.
[0016] Among them, θ M represents the learning parameters of the mass matrix function, θ D represents the learning parameter of the dissipation matrix function, θ U represents the learning parameters of the energy matrix function; q represents the robot position vector, represents the robot velocity vector, represents the robot acceleration vector; u represents control;
[0017] Step 3: Input the inherent properties, position and energy of the soft robot to be tested, which are the same as those in step 1, into the trained deep learning network model, and the trained deep learning network model outputs the mass matrix function M(q; θ M ), dissipative matrix function and energy matrix function U(q,u;θ U );
[0018] The mass matrix function M(q;θ) output by the Lagrangian function and deep learning network model M ), dissipative matrix function Energy matrix function U(q,u;θ U ) to obtain the position, velocity and acceleration of the soft robot at the next moment.
[0019] The beneficial effects of the present invention are:
[0020] This paper proposes a new modeling method for soft robots that integrates an attention mechanism and a Lagrangian neural network. This method is applicable to soft robots of various configurations and motion conditions, and is a universal dynamic modeling solution. By extracting dynamic relationships directly from experimental data, it avoids the complex steps of building a physical model, making the prediction process more efficient and concise. It also addresses the difficulty of traditional dynamic modeling methods in describing the behavior of soft robots in complex environments due to their reliance on precise physical equations.
[0021] This paper introduces an attention layer into a multi-layer neural network, allowing deep neural networks to automatically identify and weight the most informative features by learning the internal structural relationships of the input data. Each state input is converted into a query, and the network generates a corresponding key and value. The degree of match between the query and the key determines the weighted importance of the corresponding value. This mechanism enables the network to prioritize the information that is most critical to the current prediction target among many inputs.
[0022] The present invention combines the strategies of pre-training and main-task training. During the pre-training phase, the model learns Lagrangian dynamics and related physics matrices, and further incorporates the advantages of the encoder-decoder architecture during the main-task training process, enabling the model to establish direct dependencies between different time steps. This not only allows the model to learn the impact of the current state on the future of the system, but also considers the impact of historical states on the future, thereby optimizing the model from a global perspective and improving the long-term prediction accuracy of the model to a certain extent.
[0023] In this paper, we employ a positional encoding mechanism, enabling the model to break free from the limitations of traditional mathematical models when processing time series data, enabling simultaneous and parallel processing of all training data. Positional encoding adds additional positional information to each element of the model training set, enabling the neural network to identify and process the relative or absolute position of each element, and the model can independently process multiple data points within a single batch. As a result, the entire dataset can be fed into the model for training simultaneously without waiting for the previous training iteration to complete, saving significant training time and addressing the high training time and data requirements of existing physics-inspired neural networks for complex soft robotic systems.
[0024] The present invention aims to improve the prediction accuracy and anti-interference ability of the soft robot model to environmental noise by optimizing the Lagrangian neural network and the attention weights of the model state. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flow chart of the method of the present invention;
[0026] Figure 2 are the actual values and model predictions of the first two segments of the soft arm in the three dimensions of X, Y, and L when the number of pre-training steps is 500. a is the displacement of the first segment of the robot arm, b is the displacement of the second segment of the robot arm, and Δx real is the true value of the robot arm in the X dimension, Δy real is the true value of the robot arm in the Y dimension, ΔL real is the true value of the robotic arm in the L dimension, Δx predis the model prediction value of the robot arm in the X dimension, Δy pred is the model prediction value of the robot arm in the Y dimension, ΔL pred is the model prediction value of the robotic arm in the L dimension; the three dimensions X, Y, and L are the three-dimensional space coordinate system in the simulation software;
[0027] Figure 3 is the model's predicted value of the velocity of the first segment of the soft robotic arm under Gaussian white noise with a mean of 0.05, Time is the time, and First segment velocity is the velocity of the first segment of the soft robotic arm;
[0028] Figure 4 is the predicted value of the velocity of the first segment of the soft robotic arm under noise-free conditions, Time is the time, and Firstsegment velocity is the velocity of the first segment of the soft robotic arm;
[0029] Figure 5 is the predicted value of the velocity of the second segment soft robotic arm under noise-free conditions, Time is the time, and Secondsegment velocity is the velocity of the second segment soft robotic arm;
[0030] Figure 6 This is a heatmap of the attention scores of the input transformation matrix for the force on the second segment of the soft arm in the X dimension when the pre-training step number is 10. State TransitionMatrix Columns represents the rows of the input transition matrix, and StateTransition MatrixRows represents the columns of the input transition matrix. DETAILED DESCRIPTION
[0031] Specific embodiment 1: This embodiment is a soft robot control method that integrates attention mechanism and physics-inspired neural network. The specific process is as follows:
[0032] Step 1: Create a soft robot training dataset; the specific process is as follows:
[0033] The soft robot training dataset includes the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot;
[0034] The inherent properties of the soft robot are: number of soft robot arm segments n, soft robot length, soft robot width, soft robot height, soft robot weight, and soft robot density;
[0035] The energy of the soft robot is: the speed of the soft robot and the acceleration of the soft robot;
[0036] The position of the soft robot is prepared in advance and is necessary for training the network;
[0037] Step 2: Build a deep learning network model;
[0038] The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as the input of the deep learning network model, and the mass matrix function M(q; θ M ) as the output of the deep learning network model;
[0039] The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as inputs of the deep learning network model, and the dissipation matrix function As the output of the deep learning network model;
[0040] The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as the input of the deep learning network model, and the energy matrix function U(q,u;θ U ) as the output of the deep learning network model;
[0041] The mass matrix function M(q;θ) output by the Lagrangian function and deep learning network model M ), dissipative matrix function Energy matrix function U(q,u;θ U ), obtain the predicted position, velocity and acceleration of the soft robot at the next moment; this section represents the position, velocity and acceleration of the soft robot at the next moment predicted based on the physical method;
[0042] The difference between the predicted acceleration of the soft robot at the next moment and the true value of the acceleration is used as the error term of the cross-entropy loss function. The cross-entropy loss function is calculated until the loss function converges to obtain a trained deep learning network model.
[0043] Among them, θ M represents the learning parameters of the mass matrix function, θ D represents the learning parameter of the dissipation matrix function, θ U represents the learning parameters of the energy matrix function; q represents the robot position vector, represents the robot velocity vector, represents the robot acceleration vector; u represents control;
[0044] Step 3: Input the inherent properties, position and energy of the soft robot to be tested, which are the same as those in step 1, into the trained deep learning network model, and the trained deep learning network model outputs the mass matrix function M(q; θ M ), dissipative matrix function and energy matrix function U(q,u;θ U );
[0045] The mass matrix function M(q;θ) output by the Lagrangian function and deep learning network model M ), dissipative matrix function Energy matrix function U(q,u;θ U ) to obtain the position, velocity and acceleration of the soft robot at the next moment.
[0046] Specific implementation method 2: This implementation method differs from specific implementation method 1 in that: in step 2, a deep learning network model is constructed; the specific process is:
[0047] The deep learning network model includes self-attention mechanism and multi-layer perceptron MLP;
[0048] The self-attention mechanism includes an encoder and a decoder.
[0049] Other steps and parameters are the same as those in the first embodiment.
[0050] Specific embodiment three: The difference between this embodiment and specific embodiment one or two is that in step two, the inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as inputs of the deep learning network model, and the mass matrix function M(q; θ M ) as the output of the deep learning network model; the specific process is:
[0051] The inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot are used as inputs to the encoder, which outputs the query Q.
[0052] The mass matrix function M(q;θ) output by the deep learning network model corresponding to the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot in the previous iteration M ), dissipative matrix function and energy matrix function U(q,u;θ U ) as the input of the decoder, the decoder outputs the key K;
[0053] The query Q and key K are normalized by softmax to obtain the value V;
[0054] The value V is used as the input of the multilayer perceptron MLP, and the multilayer perceptron MLP outputs the quality matrix function M(q; θ M ).
[0055] Other steps and parameters are the same as those in the first or second embodiment.
[0056] Specific embodiment 4: The difference between this embodiment and any one of the specific embodiments 1 to 3 is that in step 2, the inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as inputs of the deep learning network model, and the dissipation matrix function As the output of the deep learning network model;
[0057] The specific process is:
[0058] The inherent properties of the soft robot, the position of the soft robot, and the force of the soft robot are used as inputs to the encoder, which outputs the query Q;
[0059] The mass matrix function M(q;θ) output by the deep learning network model corresponding to the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot in the previous iteration M ), dissipative matrix function and energy matrix function U(q,u;θ U ) as the input of the decoder, the decoder outputs the key K;
[0060] The query Q and key K are normalized by softmax to obtain the value V;
[0061] The value V is used as the input of the multilayer perceptron MLP, and the multilayer perceptron MLP outputs the dissipation matrix function
[0062] The other steps and parameters are the same as those in the first to third embodiments.
[0063] Specific embodiment 5: This embodiment differs from any one of specific embodiments 1 to 4 in that: in step 2, the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot are used as inputs of the deep learning network model, and the energy matrix function U(q,u;θ U ) as the output of the deep learning network model;
[0064] The specific process is:
[0065] The inherent properties of the soft robot, the position of the soft robot, and the force of the soft robot are used as inputs to the encoder, which outputs the query Q;
[0066] The mass matrix function M(q;θ) output by the deep learning network model corresponding to the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot in the previous iteration M ), dissipative matrix function and energy matrix function U(q,u;θ U ) as the input of the decoder, the decoder outputs the key K;
[0067] The query Q and key K are normalized by softmax to obtain the value V;
[0068] The value V is used as the input of the multi-layer perceptron MLP, and the multi-layer perceptron MLP outputs the energy matrix function U(q,u;θ U ).
[0069] The other steps and parameters are the same as those in the first to fourth embodiments.
[0070] Specific embodiment 6: This embodiment differs from any one of specific embodiments 1 to 5 in that the Lagrangian function in step 2 is:
[0071]
[0072] in, represents kinetic energy, and V(q) represents potential energy.
[0073] Other steps and parameters are the same as those in Specific Implementations 1 to 5-1.
[0074] Specific embodiment seven: This embodiment differs from any one of specific embodiments one to six in that: in step two, the difference between the predicted acceleration of the soft robot at the next moment and the true value of the acceleration is used as the error term of the cross-entropy loss function, and the cross-entropy loss function is calculated until the loss function converges to obtain a trained deep learning network model; the specific process is:
[0075] The cross entropy loss function is
[0076]
[0077] Where B is the set of training data samples; Indicates the predicted acceleration of the soft robot at the next moment corresponding to the p-th sample, represents the acceleration of the real soft robot at the next moment corresponding to the p-th sample; ||·||2 represents the Euclidean norm.
[0078] The other steps and parameters are the same as those in the first to sixth embodiments.
[0079] 1. Based on Lagrangian function The dynamic model of the soft robot is obtained; the expression is:
[0080]
[0081] in, represents kinetic energy, V(q) represents potential energy;
[0082] M(q) represents the mass matrix; represents the Coriolis force and centrifugal force matrix; G(q) represents the gravity vector;
[0083] q represents the robot position vector, represents the robot velocity vector, represents the robot acceleration vector;
[0084] Robot position vector q, robot velocity vector and the robot acceleration vector Construct the state vector Q,
[0085] u represents the control input,
[0086] Where n and m represent the dimensions of the robot position vector q and the control input u, respectively;
[0087] 2. Define the mass matrix function M(q;θ M ), dissipation matrix function Input transformation matrix function U(q,u;θ U );
[0088] θ M ,θ D ,θ U The network parameters of the multi-layer perceptron (MLP) Θ = {θ M ,θ D ,θ U};
[0089] Among them, θ M represents the learning parameters of the quality matrix function; where θ D represents the learning parameters of the dissipation matrix function; where θ U represents the learning parameters of the input transformation matrix function;
[0090] 3. Define the forward model function
[0091] Control u as the input of the forward model, the predicted robot position vector q next and the robot velocity vector As the output of the forward model;
[0092] θ F represents the learning parameters of the forward model;
[0093] 4. Define the inverse model function
[0094] The predicted robot position vector q next and the predicted robot velocity vector As the input of the inverse model, the predicted control u is the output of the inverse model;
[0095] θ Grepresents the learning parameters of the inverse model;
[0096] The forward model represents the prediction of the next state of the soft robot;
[0097] The inverse model represents the derivation of the control inputs that need to be applied;
[0098] The next state predicted by the forward model function will serve as the input of the inverse model function;
[0099] 5. Establish update function U(θ F ,θ G ; η), used to update the parameters θ of the forward model and the inverse model according to the loss function gradient during training F and θ G ;
[0100] The learning rate η is a tuning parameter in the update process;
[0101] 6. Set the number of degrees of freedom n dof (number of soft robot segments);
[0102] Set the width w of the Lagrangian function l and depth d l , dissipation matrix (white noise, etc.) function width w d and depth d d , the width w of the input transformation matrix function i and depth d i ,
[0103] Set the batch size n, diagonal perturbation ∈ diagonal , activation function RELU, learning rate α, weight decay λ decay , and the maximum number of training cycles N;
[0104] 7. Initialize the Lagrangian function parameters θ L , the learning parameter θ of the dissipation matrix function D , the learning parameters θ of the input transformation matrix function U ;
[0105] Initialize the Lagrangian objective function parameters Learning parameters of the dissipative matrix objective function Input transformation matrix objective function learning parameters
[0106] 8. Define the attention mechanism encoder Encoder attn ;
[0107] The current state s t Input attention mechanism encoder Encoder attn, attention mechanism encoder Encoder attn Output Encoder attn (s t θ enc );
[0108] θ enc Represents encoder parameters;
[0109] θ enc Learn during training to capture key features and dependencies in the state space;
[0110] 9. Define the attention mechanism decoder Decoder attn ;
[0111] The output of the encoder is Encoder attn (s t θ enc ) Input attention mechanism decoder Decoder attn , attention mechanism decoder Decoder attn (Encoder attn (s t );θ dec ) Output the predicted next state s t+1 and control behaviora t ;
[0112] θ dec Represents decoder parameters;
[0113] The decoder structure utilizes the context information provided by the encoder and uses its parameters θ dec To generate corresponding actions and state transitions;
[0114] 10. Repeat steps 8 to 9 until the attention mechanism encoder Encoder attn and attention mechanism decoder Decoder attn Converge and obtain an optimized attention mechanism encoder Encoder attn and attention mechanism decoder Decoder attn ;
[0115] 11. Initialize the weight matrix θ of query Q, key K and value V Q ,θ K and θ V ;
[0116] 12. The predicted robot position vector q next , robot velocity vector Convert the predicted control input u into the embedding vector Eq u ;
[0117] Embed the vector E q Converted to K=E q θ K , Q=E q θ Q ;
[0118] Embed the vector E u Convert to V=E u θ V ;
[0119] 13. Calculate the compatibility score between the query Q and the key K, and then determine the attention weight matrix A for each value V:
[0120]
[0121] Among them, d k represents the dimension of the key vector, T represents the matrix transpose operation; softmax() represents the activation function;
[0122] 14. Based on the attention weight matrix A and the value V, the state vector is updated by multiplication to obtain a new value V'=AV;
[0123] 15. Apply the gradient descent algorithm (SGD) to update the policy network f π , the policy network loss function L loss The parameter is θ π ,θ V ,θ loss ; Perform the following substeps after each iteration;
[0124] The network is evaluated to determine whether it meets predetermined performance criteria or whether more iterations are needed.
[0125] If the performance does not meet the predetermined standard, return to step 7 and consider adjusting the learning rate or other hyperparameters.
[0126] If the performance of the model on the validation dataset no longer improves, or the predetermined number of iterations N is reached, the iteration is stopped.
[0127] 17. Apply residual connection to the output of the multi-head attention layer to obtain feature R = A multi +E' q , and then perform layer normalization on the feature R to obtain the feature E norm =LayerNorm(R);
[0128] 18. Use feedforward network FFN to feature E norm Processing is performed to obtain feature E FFN ; The process is:
[0129] For feature Enorm Processing to obtain feature F FFN , the expression is:
[0130] F FFN =RELU(FFN(E norm ))
[0131] Among them, RELU is a nonlinear activation function;
[0132] Based on feature F FFN and feature E norm Get feature E FFN , the expression is:
[0133] E FFN =LayerNorm(F FFN +E norm )
[0134] 19. Through the Softmax layer, the feedforward neural network FFN output feature E FFN Converted to state transition probability P next =Softmax(Linear(E FFN ));
[0135] The state transition probability P next Convert to a specific action or state value q next ;
[0136] 20. If the state transition probability P next If the maximum state transition probability value in exceeds the preset threshold, the corresponding action a is selected t And update the state s t+1 ;
[0137] Otherwise, based on the state transition probability P next Adopting the exploration strategy q(a t |P next ) to select an action and update the state;
[0138] Step 21: Use the Error Analysis function (q next ,P next ) Evaluate the error between model output and actual observations;
[0139] Adjust the deep learning parameters Θ of the multi-layer perceptron (MLP) network layer according to the error results to optimize the prediction accuracy; confirm the physical model output q of the soft robot next After the predetermined performance index is met, the deep learning parameter Φ of the optimized multi-layer perceptron (MLP) network layer is obtained, that is, the optimized multi-layer perceptron (MLP) network layer is obtained;
[0140] This error analysis function is used to adjust the entire training process, which includes the attention mechanism and the ordinary neural network, to guide the training towards the point where the error is minimized.
[0141] 22. Input the control vector of the robot to be tested into the optimized multi-layer perceptron (MLP) network layer, and the optimized multi-layer perceptron (MLP) network layer outputs the predicted robot position vector and velocity vector.
[0142] Prepare to enter the practical application stage.
[0143] The soft robotic arm consists of n segments, each with three degrees of freedom: x, y, and l. The objective to be optimized is:
[0144] Φ(τ,ε)=k τ ×τ+k ε ×ε
[0145] Among them, τ is the time consumption, ε is the energy loss, k τ is the weight coefficient of time consumption, k ε is the weight coefficient of energy time consumption;
[0146] The target to be optimized is the dynamic model of the soft robot
[0147] The energy loss ε is defined as the integral sum of each degree of freedom of the robot arm:
[0148]
[0149] Among them, τ ij represents the control torque of the i-th segment on the j degree of freedom, represents the angular velocity or linear velocity of the i-th segment in the j-degree of freedom; t represents time;
[0150] Each τ ij and The product of represents the power consumption in j degrees of freedom, thus capturing the overall situation of energy loss;
[0151] The state vector only involves the intrinsic dynamics of the soft robotic arm and is expressed as:
[0152]
[0153] Among them, q x ,q y ,q l Respectively represent the positions of each segment of the soft robotic arm in the x, y and l directions, and Indicates the speed of each segment of the soft robotic arm in the x, y and l directions;
[0154] For each segment of the soft robotic arm, the Lagrangian function is defined And the Lagrangian function Decomposed into kinetic energy and potential energy V(q), represents kinetic energy and V(q) represents potential energy; it is approximated by a neural network of the following form:
[0155]
[0156] N θ is a multi-layer perceptron (MLP) network layer with parameter θ, which is used to approximate the true Lagrangian function;
[0157] The encoder in step eight is specifically:
[0158] Initialize the encoder parameters, which include the weight matrix and bias vector;
[0159] Each layer of the multi-layer perceptron MLP structure gradually extracts the features of the input state through a linear transformation followed by a nonlinear activation function.
[0160] The encoder is constructed as follows:
[0161] e enc =MLP Encoder (s t θ enc )
[0162] Among them, θ enc Represents the parameters of the encoder, s t Represents the state of the current time step; MLP Encoder represents the multi-layer perceptron corresponding to the encoder; e enc Indicates the status of the encoder output;
[0163] The decoder in step nine is specifically:
[0164] The encoder output state e enc Passed to the decoder;
[0165] The purpose of the decoder is to parse the output of the encoder and θ The three physics-inspired matrices in [1] calculate the attention scores and generate corresponding control actions or predictions. The decoder also adopts the MLP structure.
[0166] The decoder construction process is as follows:
[0167] a t =MLP Decoder (e enc θ dec )
[0168] Among them, θ dec Represents the decoder parameters, a t Represents the action or torque obtained by mapping the encoded state; MLP Decoder Represents the multi-layer perceptron corresponding to the decoder; there is an MLP network in the decoder;
[0169] The process of obtaining the soft robot dynamics model is as follows:
[0170] The self-attention score Attention in the decoder is used to predict the torque τ:
[0171] τ att =Attention(Q,K,V)
[0172] Using the learned Lagrangian function and predicted torque τ att To construct the Lagrangian function Through the Lagrangian function Get the soft robot dynamics model The differential form of
[0173] Soft robot dynamics model The differential form of the expression is:
[0174]
[0175] According to minimizing the predicted state q next and the target state q target The difference between them defines a loss function L(θ) to evaluate the accuracy of the model prediction:
[0176]
[0177] Where B is the set of training data samples;
[0178] q next,i Indicates the next state predicted by the pth sample, q target,i represents the target state of the p-th sample;
[0179] ||·||2 represents the Euclidean norm, which is used to calculate the squared difference between the predicted state and the target state;
[0180] The model will enter the strategy optimization phase:
[0181] At this stage, the policy network π φ The introduction is to directly learn the attention mechanism, the policy network π φ Take state s as input and the probability distribution of action a as output;
[0182] Policy Network π φ The loss L policy (φ) reflects the superiority of the selected action relative to the average level and is calculated as follows:
[0183]
[0184] The parameter update of the policy network follows the gradient ascent method to maximize the cumulative reward. The update rule is given by the following formula:
[0185]
[0186] Where α is the learning rate, Denotes the loss function L policy Gradient with respect to the policy network parameters φ;
[0187] Finally, the performance of the policy network is verified by the test set; the performance indicators mainly include prediction accuracy and action curve fit; specifically, by calculating the average loss on the test set To measure generalization ability:
[0188]
[0189] Among them, y att (s', r; θ') represents the target value function, B test Represents the set of test data samples;
[0190] In addition, the action sequences generated by the policy network on the test set are compared with the optimal action sequences obtained in the real world or in high-precision simulations.
[0191] The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. A soft robot control method that integrates an attention mechanism with a physics-inspired neural network, characterized by: The specific process of the method is: Step 1: Create a soft robot training dataset; the specific process is as follows: The soft robot training dataset includes the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot; The inherent properties of the soft robot are: number of soft robot arm segments n, soft robot length, soft robot width, soft robot height, soft robot weight, and soft robot density; The energy of the soft robot is: the speed of the soft robot and the acceleration of the soft robot; Step 2: Build a deep learning network model; The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as the input of the deep learning network model, and the mass matrix function M(q; θ M ) as the output of the deep learning network model; The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as inputs of the deep learning network model, and the dissipation matrix function As the output of the deep learning network model; The inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as the input of the deep learning network model, and the energy matrix function U(q,u;θ U ) as the output of the deep learning network model; The mass matrix function M(q;θ) output by the Lagrangian function and deep learning network model M ), dissipative matrix function Energy matrix function U(q,u;θ U ), obtain the predicted position, velocity and acceleration of the soft robot at the next moment; The difference between the predicted acceleration of the soft robot at the next moment and the true value of the acceleration is used as the error term of the cross-entropy loss function. The cross-entropy loss function is calculated until the loss function converges to obtain a trained deep learning network model. Among them, θ M represents the learning parameters of the mass matrix function, θ D represents the learning parameter of the dissipation matrix function, θ U represents the learning parameters of the energy matrix function; q represents the robot position vector, represents the robot velocity vector, represents the robot acceleration vector; u represents control; Step 3: Input the inherent properties, position and energy of the soft robot to be tested, which are the same as those in step 1, into the trained deep learning network model, and the trained deep learning network model outputs the mass matrix function M(q; θ M ), dissipative matrix function and energy matrix function U(q,u;θ U ); The mass matrix function M(q;θ) output by the Lagrangian function and deep learning network model M ), dissipative matrix function Energy matrix function U(q,u;θ U ) to obtain the position, velocity and acceleration of the soft robot at the next moment.
2. The method for controlling a soft robot that integrates an attention mechanism with a physics-inspired neural network according to claim 1, characterized in that: In step 2, a deep learning network model is constructed; the specific process is as follows: The deep learning network model includes self-attention mechanism and multi-layer perceptron MLP; The self-attention mechanism includes an encoder and a decoder.
3. The method for controlling a soft robot that integrates an attention mechanism with a physics-inspired neural network according to claim 2, characterized in that: In the second step, the inherent properties of the soft robot, the position of the soft robot and the energy of the soft robot are used as the input of the deep learning network model, and the mass matrix function M(q; θ M ) as the output of the deep learning network model; the specific process is: The inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot are used as inputs to the encoder, which outputs the query Q. The mass matrix function M(q;θ) output by the deep learning network model corresponding to the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot in the previous iteration M ), dissipative matrix function and energy matrix function U(q,u;θ U ) as the input of the decoder, the decoder outputs the key K; The query Q and key K are normalized by softmax to obtain the value V; The value V is used as the input of the multilayer perceptron MLP, and the multilayer perceptron MLP outputs the quality matrix function M(q; θ M ).
4. The method for controlling a soft robot that integrates an attention mechanism with a physics-inspired neural network according to claim 3, characterized in that: In step 2, the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot are used as inputs of the deep learning network model, and the dissipation matrix function As the output of the deep learning network model; the specific process is: The inherent properties of the soft robot, the position of the soft robot, and the force of the soft robot are used as inputs to the encoder, which outputs the query Q; The mass matrix function M(q;θ) output by the deep learning network model corresponding to the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot in the previous iteration M ), dissipative matrix function and energy matrix function U(q,u;θ U ) as the input of the decoder, the decoder outputs the key K; The query Q and key K are normalized by softmax to obtain the value V; The value V is used as the input of the multilayer perceptron MLP, and the multilayer perceptron MLP outputs the dissipation matrix function 5. The method for controlling a soft robot that integrates an attention mechanism with a physics-inspired neural network according to claim 4, characterized in that: In step 2, the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot are used as inputs of the deep learning network model. The energy matrix function U(q,u;θ U ) as the output of the deep learning network model; the specific process is: The inherent properties of the soft robot, the position of the soft robot, and the force of the soft robot are used as inputs to the encoder, which outputs the query Q; The mass matrix function M(q;θ) output by the deep learning network model corresponding to the inherent properties of the soft robot, the position of the soft robot, and the energy of the soft robot in the previous iteration M ), dissipative matrix function and energy matrix function U(q,u;θ U ) as the input of the decoder, the decoder outputs the key K; The query Q and key K are normalized by softmax to obtain the value V; The value V is used as the input of the multi-layer perceptron MLP, and the multi-layer perceptron MLP outputs the energy matrix function U(q,u;θ U ).
6. The method for controlling a soft robot that integrates an attention mechanism with a physics-inspired neural network according to claim 5, characterized in that: The Lagrangian function in step 2 is: in, represents kinetic energy, and V(q) represents potential energy.
7. The method for controlling a soft robot that integrates an attention mechanism with a physics-inspired neural network according to claim 6, characterized in that: In step 2, the difference between the predicted acceleration of the soft robot at the next moment and the true value of the acceleration is used as the error term of the cross entropy loss function, and the cross entropy loss function is calculated until the loss function converges to obtain a trained deep learning network model; the specific process is: The cross entropy loss function is Among them, B is the set of training data samples; Indicates the predicted acceleration of the soft robot at the next moment corresponding to the p-th sample, represents the acceleration of the real soft robot at the next moment corresponding to the p-th sample; ||·||2 represents the Euclidean norm.
Citation Information
Patent Citations
Model order reduction method for rapid simulation of pneumatic soft robot
CN114347029A
Rope-driven flexible double-joint bionic crab and control method
CN114701583A