Bionic jellyfish robot propulsion and grabbing collaborative control method

CN122807938APending Publication Date: 2026-09-25HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611267675.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明所要解决的技术问题是:现有技术中抓取与推进相互干扰及运动控制智能化不足的问题

Benefits of technology

[0051]本发明通过将推进与抓取在控制中心层面进行独立或协同的灵活调度,实现了水下浮动基座上的精准作业;利用PPO离线训练生成策略,并将最终网络固化为轻量级的动作决策表,兼具了智能决策的全局最优性和工程应用的实时性;针对SMA丝驱动特有的加热-冷却滞后特性,PPO算法配合多离散动作空间能有效处理非线性的开关控制逻辑,确保了推进动作的平滑切换。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807938A_ABST
    Figure CN122807938A_ABST
Patent Text Reader

Abstract

The application discloses a kind of bionic jellyfish robot propulsion and grabbing collaborative control method, comprising: obtaining the operating parameter of bionic jellyfish robot;Kinetics and kinematics model is constructed;Define the action space, observation space and reward function of deep reinforcement learning;Initialize a policy network with random network parameters, train the policy network using PPO algorithm with multiple discrete actions, adjust the parameters of policy network, so that the action output by policy network according to observation can maximize cumulative reward, and after training, the policy network is solidified into action decision table;Action decision table is deployed to the control center of robot, combined with the posture information collected in real time to generate propulsion control command, and according to the demand of grabbing task, the driving mechanism of grabbing tentacle is independently or cooperatively controlled by control center to execute grabbing action.The application solves the mutual interference problem of propulsion action and grabbing action in dynamics, and realizes accurate operation on underwater floating base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for coordinated control of propulsion and grasping of a biomimetic jellyfish robot, belonging to the field of biomimetic underwater robot technology. Background Technology

[0002] Bionic jellyfish, with their low noise, low disturbance, and good environmental adaptability, have high application value in underwater exploration, sample collection, and ecological monitoring. Existing underwater robots mostly use propeller propulsion and rigid end effectors for grasping, making it difficult to simultaneously meet the requirements of low-disturbance motion and flexible grasping. While some bionic jellyfish possess bionic propulsion capabilities, they still have significant shortcomings in terms of lightweight structure, allocation of grasping functions, and adaptability to grasping complex targets.

[0003] Existing biomimetic jellyfish head support structures largely rely on empirical design, resulting in excessive structural redundancy and difficulty in balancing strength, installation space, and overall weight. Especially after integrating control units, drive units, and tentacle mounting structures internally, the head structure is prone to localized load concentration and low material utilization, limiting the miniaturization and improved mobility of the entire device.

[0004] In existing technologies, the grasping and propulsion functions are often coupled in the same set of execution components or in mutually interfering layouts, making it impossible to optimize propulsion and grasping independently. When performing a grasping action, the propulsion component is easily affected, thus impacting attitude stability; when performing a propulsion action, the grasping component is easily affected by hydrodynamic disturbances, leading to reduced grasping accuracy.

[0005] Furthermore, when facing marine life, reef attachments, or other targets with complex surface shapes, traditional rigid gripping or low-contact-point gripping methods struggle to achieve stable coverage, easily leading to slippage, localized damage, or weak gripping. Meanwhile, existing biomimetic jellyfish motion control methods mostly employ pre-programmed or traditional PID methods, which are ill-suited to the nonlinearity and strong coupling of the underwater environment, as well as the thermodynamic hysteresis characteristics of SMA materials. There is a lack of an intelligent control method capable of autonomously learning the optimal driving strategy. Summary of the Invention

[0006] The technical problem to be solved by the present invention is the problem of mutual interference between grasping and propulsion and the lack of intelligent motion control in the prior art.

[0007] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution.

[0008] On one hand, the present invention provides a method for coordinated control of propulsion and grasping of a biomimetic jellyfish robot, the biomimetic jellyfish robot including a control center, grasping tentacles, and propulsive tentacles, the coordinated control including:

[0009] Obtain the structural and physical parameters of the biomimetic jellyfish robot;

[0010] A dynamic and kinematic model is constructed based on the structural and physical parameters.

[0011] Based on the aforementioned dynamics and kinematics model, an action space, observation space, and reward function for deep reinforcement learning are defined to form a decision-making process framework for training policy networks. The action space defines the action types that each pushing tentacle can choose, the observation space defines the environmental and self-state information that the agent can obtain when making decisions, and the reward function is used to quantify the quality of each action.

[0012] Based on the aforementioned decision-making process framework, a policy network with random network parameters is initialized, and the policy network is trained using the PPO algorithm with multiple discrete actions. The parameters of the policy network are adjusted so that the policy network can maximize the cumulative reward based on the actions of the observed inputs and outputs. After training, the policy network is solidified into an action decision table.

[0013] The motion decision table is deployed to the robot's control center, and propulsion control commands are generated by combining the real-time collected posture information. Based on the grasping task requirements, the control center independently or collaboratively controls the drive mechanism of the grasping tentacles to execute the grasping action.

[0014] By flexibly scheduling the propulsion tentacles and grasping tentacles independently or in coordination at the control center level, the problem of mutual interference between propulsion and grasping actions in terms of dynamics is solved, enabling precise operations on underwater floating bases.

[0015] The dynamics and kinematics model includes the physical geometry parameters, mass distribution parameters, hydrodynamic parameters, and driving characteristic parameters of the biomimetic jellyfish, which are used to simulate the robot's motion response and force state in a simulation environment.

[0016] The construction of the dynamics and kinematics model includes:

[0017] Each propelling tentacle is divided into several segments, and multiple SMA filaments are embedded in each segment. The composite curvature vector of the segment is calculated by superimposing the bending angles of each SMA filament. Then, each segment is integrated segment by segment to obtain the spatial pose of the entire tentacle, thereby establishing a mapping relationship from the driving amount of the SMA filaments to the position and attitude of the tentacle end.

[0018] The driving torque acting on the robot body is calculated, and the rotational dynamics equation of the robot body is established by combining the hydrodynamic damping torque and the gyroscopic torque. When solving the rotational dynamics equation, a semi-implicit integration method is adopted. First, the increment of the angular velocity of the driving term and the gyroscopic term is explicitly calculated, and then the hydrodynamic damping term is implicitly processed to maintain numerical stability. Finally, the rotation matrix and attitude of the robot are updated.

[0019] By segmenting the tentacles and superimposing the bending curvature of SMA filaments, and combining the segment-by-segment integration method, a high-fidelity mapping model from the microscopic SMA driving force to the macroscopic tentacle end pose was established, which makes up for the inability of the traditional rigid linkage model to describe the flexible deformation characteristics of jellyfish.

[0020] The steps for explicitly calculating the driving term are as follows:

[0021] ,

[0022] ,

[0023] in, The angular acceleration vector generated by the drive, The matrix is ​​the inverse of the moment of inertia tensor. This is the sum of undamped moments. This is the intermediate angular velocity vector after explicit updating by the driving term. This represents the time step for numerical integration.

[0024] The implicit processing of the hydrodynamic damping term is described in the following formula:

[0025] ,

[0026] in, To account for hydrodynamic damping corrections, the final updated tentacle angular velocity vector at this time step is... This is the equivalent second-order damping coefficient. , For effective rotational inertia, It is the Euclidean norm.

[0027] The action space includes: each pushing tentacle independently selects from 3 discrete actions, namely closing, bending upward and bending downward;

[0028] The observation space includes at least the current decision step index, the remaining cooling steps for each pushing tentacle, the previous action, the robot's center of mass displacement, center of mass velocity, rotation matrix elements, and training round information;

[0029] The reward function includes a basic reward at each step and a final state reward. The basic reward at each step is used to suppress invalid activation and collision behavior, and the final state reward includes a roll angle Gaussian reward, a parasitic rotation penalty, a position offset penalty, an energy-saving reward, and a collision comprehensive evaluation value.

[0030] The step of training the policy network using the PPO algorithm with multiple discrete actions includes:

[0031] The policy network collects experience samples, including state, action, reward, and next state, as the data basis for parameter updates.

[0032] The Actor network takes the observation space as input and outputs the probability distribution of the action space, while the Critic network takes the observation space as input and outputs the state value estimate. Before the action is output, an action masking mechanism is used to mask actions that do not conform to physical constraints.

[0033] Based on the collected experience samples, the generalized advantage estimation method is used to obtain the advantage function for each action; the temporal difference error is calculated based on the immediate reward, discount factor, and the Critic network's value estimation of the current state and the next state from the experience samples.

[0034] The policy loss is calculated based on the advantage function, which uses the PPO-clip objective function and incorporates an entropy regularization term to encourage exploration, thereby updating the Actor network parameters; the Critic network parameters are updated based on the temporal difference error.

[0035] The training process employs a four-round progressive training strategy, opening and exploring each segment of the propelling tentacle in rounds from near to far: in each round, only one segment is opened for the policy network to learn; the optimal action learned in the previous round is locked in subsequent rounds by fixing the probability distribution in the action space corresponding to that segment or by using the action masking mechanism, and is no longer included in the training; after each round of training, the action sequence with the highest cumulative reward is selected and saved to guide the fixed segment action in the next round of training.

[0036] The action masking mechanism includes at least one of a cooling mask, a stage mask, and a direction prohibition mask;

[0037] The cooling mask includes: the number of remaining cooling steps. The push-tent forces the action to be turned off only, and the cooling counter is set to zero the instant the power is switched from on to off. Then, decrease by 1 for each subsequent decision step;

[0038] The phase mask includes: during the forced cooling phase, all pushing tentacles are prohibited from bending upwards and downwards;

[0039] The direction prohibition mask includes: the push tentacles that performed an upward bend in the previous decision step prohibit the current step from selecting a downward bend, and vice versa.

[0040] The cooling mask forces the SMA to enter a 3-step cooling period after performing high-power up-bend / down-bend movements. This physically protects the SMA filament from being burned out by overcurrent and extends the actuator's lifespan.

[0041] The propulsion control includes:

[0042] When the control center receives a rotation command, it initiates a rotation cycle that includes a power-on drive period and a forced cooling period. During the power-on drive period, the activation range of the push tentacles is cyclically mapped according to the target rotation direction, and the action decision table is assigned to the push tentacles with the corresponding numbers.

[0043] During the forced cooling period, all SMA filaments are de-energized, and new activation requests can only be executed after the cooling time constraint is met.

[0044] By binding the decision table to the physical numbered tentacles through cyclic mapping, and combining it with explicit timing logic for the power-on drive period and forced cooling period, the rotation command can generate a stable and predictable net torque by activating tentacles in a specific range, simplifying the underlying drive logic.

[0045] The coordinated control includes:

[0046] The grasping tentacles and the pushing tentacles are driven by separate drive links. The grasping tentacles are driven by a motor to retract and extend ropes to achieve bending and wrapping, while the pushing tentacles are driven by SMA filaments to achieve propulsion and attitude adjustment.

[0047] The control center reduces or pauses the propulsion action of the tentacles when performing the grasping task, and resumes the propulsion function after the grasping is completed, so as to achieve the timing coordination of propulsion and grasping.

[0048] During the grasping process, the grasping parameters are adjusted in real time based on sensor data.

[0049] The gripping tentacles employ a separate drive link between the motor rope and the pushing tentacles (SMA), achieving functional decoupling between heavy-duty gripping and flexible propulsion.

[0050] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0051] This invention achieves precise operations on underwater floating bases by flexibly scheduling propulsion and grasping independently or collaboratively at the control center level; it utilizes the PPO offline training generation strategy and solidifies the final network into a lightweight action decision table, combining the global optimality of intelligent decision-making with the real-time performance for engineering applications; and it effectively handles nonlinear switching control logic by combining the PPO algorithm with a multi-discrete action space to address the heating-cooling hysteresis characteristic unique to SMA filament drive, ensuring smooth switching of propulsion actions. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the biomimetic jellyfish robot propulsion and grasping coordinated control method shown in Embodiment 1 of the present invention;

[0053] Figure 2 This is a schematic diagram of the theoretical response curves of SMA wire under energized heating and de-energized cooling as shown in Embodiment 2 of the present invention;

[0054] Figure 3 This is a schematic diagram of the biomimetic jellyfish robot structure shown in Embodiment 2 of the present invention;

[0055] Figure 4 This is a schematic diagram of the PPO algorithm network structure shown in Embodiment 2 of the present invention;

[0056] Figure 5 This is a schematic diagram of the four-round progressive training open sequence as shown in Embodiment 2 of the present invention;

[0057] Figure 6 This is a schematic diagram of the segmented state of the four-round progressive training as shown in Embodiment 2 of the present invention.

[0058] Explanation of reference numerals in the attached diagram: 1. Head shell; 2. Control center; 3. Driving tentacle; 4. Driving tentacle. Detailed Implementation

[0059] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0060] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0061] Example 1

[0062] like Figure 1 As shown in the figure, this embodiment introduces a method for coordinated control of propulsion and grasping of a biomimetic jellyfish robot. The biomimetic jellyfish robot includes a control center, grasping tentacles, and propulsive tentacles. The coordinated control includes:

[0063] Obtain the structural and physical parameters of the biomimetic jellyfish robot;

[0064] Construct dynamic and kinematic models based on structural and physical parameters;

[0065] Based on dynamics and kinematics models, we define the action space, observation space, and reward function for deep reinforcement learning to form a decision-making process framework for training policy networks. The action space limits the types of actions that each pushing tentacle can take, the observation space defines the environmental and self-state information that the agent can obtain when making decisions, and the reward function is used to quantify the quality of each action.

[0066] Based on the decision process framework, a policy network with random network parameters is initialized, and the policy network is trained using the PPO algorithm with multiple discrete actions. The parameters of the policy network are adjusted so that the policy network can maximize the cumulative reward based on the actions of the observed input and output. After training, the policy network is solidified into an action decision table.

[0067] The motion decision table is deployed to the robot's control center, and propulsion control commands are generated by combining the real-time collected posture information. Based on the grasping task requirements, the control center independently or collaboratively controls the drive mechanism of the grasping tentacle to execute the grasping action.

[0068] The dynamics and kinematics model includes the physical geometry parameters, mass distribution parameters, hydrodynamic parameters, and driving characteristic parameters of the biomimetic jellyfish, which are used to simulate the robot's motion response and force state in a simulation environment.

[0069] The construction of dynamic and kinematic models includes:

[0070] Each propelling tentacle is divided into several segments, and multiple SMA filaments are embedded in each segment. The composite curvature vector of the segment is calculated by superimposing the bending angles of each SMA filament. Then, each segment is integrated segment by segment to obtain the spatial pose of the entire tentacle, thereby establishing a mapping relationship from the driving amount of the SMA filaments to the position and attitude of the tentacle end.

[0071] The driving torque acting on the robot body is calculated, and the rotational dynamics equation of the robot body is established by combining the hydrodynamic damping torque and the gyroscopic torque. When solving the rotational dynamics equation, a semi-implicit integration method is adopted. First, the increment of the angular velocity of the driving term and the gyroscopic term is explicitly calculated, and then the hydrodynamic damping term is implicitly processed to maintain numerical stability. Finally, the rotation matrix and attitude of the robot are updated.

[0072] The steps for explicitly calculating the driving term are as follows:

[0073] ,

[0074] ,

[0075] in, The angular acceleration vector generated by the drive, The matrix is ​​the inverse of the moment of inertia tensor. This is the sum of undamped moments. This is the intermediate angular velocity vector after explicit updating by the driving term. This represents the time step for numerical integration.

[0076] The steps for implicitly handling the hydrodynamic damping term are as follows:

[0077] ,

[0078] in, To account for hydrodynamic damping corrections, the final updated tentacle angular velocity vector at this time step is... This is the equivalent second-order damping coefficient. , For effective rotational inertia, It is the Euclidean norm.

[0079] Specifically, effective moment of inertia , , For the moment of inertia tensor, The intermediate angular velocity vector, It is a unit vector in the direction of angular velocity.

[0080] The motion space includes: each pushing tentacle independently selects from 3 discrete actions, namely closing, bending upwards, and bending downwards;

[0081] The observation space includes at least the current decision step index, the remaining number of cooling steps for each pushing tentacle, the previous action, the robot's center of mass displacement, center of mass velocity, rotation matrix elements, and training round information;

[0082] The reward function includes a basic reward at each step and a final state reward. The basic reward at each step is used to suppress invalid activation and collision behavior, while the final state reward includes a Gaussian roll angle reward, a parasitic rotation penalty, a position offset penalty, an energy-saving reward, and a comprehensive collision evaluation value.

[0083] Specifically, the basic reward formula for each step is as follows:

[0084] ,

[0085] in, This is the base reward value for a single time step. The number of driving tentacles activated in the current step. This represents the total number of robot tentacles. This is an indicator function; it takes a value of 1 when a collision occurs and a value of 0 when there is no collision.

[0086] Specifically, the Gaussian reward formula for roll angle is as follows:

[0087] ,

[0088] in, The Gaussian reward value matched for the roll angle. It is a natural exponential function. This is the current roll angle of the aircraft. To calculate the arctangent of the roll angle using the second and third elements of the third row of the rotation matrix, The standard deviation is a Gaussian distribution. ;

[0089] Specifically, the formula for the parasitic rotation penalty is as follows:

[0090] ,

[0091] in, The penalty values ​​are for parasitic rotations in the pitch and yaw directions. This is the absolute value of the pitch angle. This is the absolute value of the yaw angle;

[0092] Specifically, the position offset penalty formula is as follows:

[0093] ,

[0094] in, This is the penalty value corresponding to the position offset. This represents the current position vector of the biomimetic jellyfish robot. The desired target position vector for the biomimetic jellyfish robot;

[0095] Specifically, the energy-saving reward formula is as follows:

[0096] ,

[0097] in, Optimize energy consumption with corresponding energy-saving reward values;

[0098] Specifically, the formula for the comprehensive collision evaluation value is as follows:

[0099] ,

[0100] in, The comprehensive reward / penalty value corresponding to the collision scenario. This represents the number of tentacles that have collided at the current moment.

[0101] The steps for training the policy network using the PPO algorithm with multiple discrete actions include:

[0102] Experience samples, including state, action, reward, and next state, are collected through a policy network to serve as the data basis for parameter updates.

[0103] The Actor network takes the observation space as input and outputs the probability distribution of the action space, while the Critic network takes the observation space as input and outputs the state value estimate. Before the action is output, an action masking mechanism is used to mask actions that do not conform to physical constraints.

[0104] Based on the collected empirical samples, the generalized dominance estimation method is used to obtain the dominance function for each action; the temporal difference error is calculated based on the immediate reward, discount factor, and the Critic network's value estimation of the current state and the next state from the empirical samples.

[0105] The policy loss is calculated based on the advantage function, which adopts the PPO-clip objective function and incorporates an entropy regularization term to encourage exploration in order to update the Actor network parameters; the Critic network parameters are updated based on the temporal difference error.

[0106] Specifically, the objective function of PPO-clip includes:

[0107] Actor strategy loss:

[0108] ,

[0109] in, Let the policy loss function of the Actor network be... The number of samples in a training batch. This represents the probability ratio between the current strategy and the old strategy. , For the sample The advantage estimate, For the clipping function, This is the PPO cutting factor. These are the entropy regularization weighting coefficients. The information entropy function of the policy distribution. For the first The probability distribution of the current policy corresponding to each action dimension. For the first The state (observation) corresponding to each training sample. For parameters The current strategy, For the first The sample at the th Actions in each action dimension This is the old strategy with fixed parameters;

[0110] Specifically, the Critic network parameters are updated based on the temporal difference error, and its value loss function is independent of the aforementioned PPO-clip objective function, as follows:

[0111] ,

[0112] in, Let be the value loss function of the Critic network. For the current state of the Critic network The value estimate, For the sample Instant rewards The preset discount factor, To update the old Critic network for the next state Value estimate.

[0113] The training process employs a four-round progressive training strategy, opening up and exploring each segment of the propelling tentacle in rounds from near to far: only one segment is opened for the policy network to learn in each round, and the optimal action learned in the previous round is locked in subsequent rounds by fixing the probability distribution in the action space corresponding to that segment or by using the action masking mechanism, and is no longer involved in training; after each round of training, the action sequence with the highest cumulative reward is selected and saved to guide the fixed segment action in the next round of training.

[0114] Action masking mechanisms include at least one of cooling masks, stage masks, and orientation prohibition masks;

[0115] The cooling mask includes: the number of cooling steps remaining. The push-tent forces the action to be turned off only, and the cooling counter is set to zero the instant the power is switched from on to off. Then, decrease by 1 for each subsequent decision step;

[0116] The phase mask includes: during the forced cooling phase, all pushing tentacles are prohibited from bending upwards or downwards;

[0117] Directional prohibition masks include: the push tentacles that pushed the previous decision step to bend upwards prohibit the current step from choosing to bend downwards, and vice versa.

[0118] Propulsion control includes:

[0119] When the control center receives a rotation command, it initiates a rotation cycle that includes a power-on drive period and a forced cooling period. During the power-on drive period, the activation range of the push tentacles is cyclically mapped according to the target rotation direction, and the action decision table is assigned to the push tentacles with the corresponding numbers.

[0120] During the forced cooling period, all SMA filaments are de-energized, and new activation requests can only be executed after the cooling time constraint is met.

[0121] By binding the decision table to the physical numbered tentacles through cyclic mapping, and combining it with explicit timing logic for the power-on drive period and forced cooling period, the rotation command can generate a stable and predictable net torque by activating tentacles in a specific range, simplifying the underlying drive logic.

[0122] Coordinated control includes:

[0123] The grasping tentacles and the pushing tentacles adopt separate drive links. The grasping tentacles are driven by a motor to retract and extend the rope to achieve bending and wrapping, while the pushing tentacles are driven by SMA filaments to achieve propulsion and attitude adjustment.

[0124] The control center reduces or pauses the propulsion action of the tentacles when performing the grasping task, and resumes the propulsion function after the grasping is completed, so as to achieve the timing coordination of propulsion and grasping.

[0125] During the grasping process, the grasping parameters are adjusted in real time based on sensor data.

[0126] Example 2

[0127] This embodiment introduces a biomimetic jellyfish robot propulsion and grasping coordinated control method. Figure 3 The application of the biomimetic jellyfish robot shown.

[0128] Step 1: Obtain the four propulsion tentacles of the biomimetic jellyfish robot, with each tentacle having a total length of... cylinder diameter Divided into Segment, length of each segment Each embedded segment SMA filaments are distributed at 90° intervals on the cylindrical surface of the tentacle (0° at the back, 180° at the front, 90° at the top, and 270° at the bottom), with a filament diameter of... The material is a nickel-titanium alloy with a density of Each filament is independently controlled to switch on and off;

[0129] Step 2: Construct a dynamic and kinematic model for the biomimetic jellyfish robot:

[0130] Step 2.1: Single-segment mass It consists of a cylindrical matrix and 4 SMA filaments:

[0131] ,

[0132] in, , , Density of 3D printing materials;

[0133] Step 2.2: As Figure 2 As shown, the bending angle θ of each SMA wire evolves independently according to the energizing state, using a first-order linear model;

[0134] Maximum bending angle Heating time depends on driving voltage : hour , hour Intermediate value linear interpolation, heating process:

[0135] ,

[0136] in, The initial bending angle of the SMA wire before it is electrically heated. This is the current real-time of the system. The moment when the SMA filament begins to heat up after being energized;

[0137] The cooling process (after power failure) linearly decays to zero; cooling time constant. :

[0138] ,

[0139] in, The bending angle of the SMA at the instant the power is cut off (when heating stops). The moment when the SMA cuts off the drive voltage and begins natural cooling;

[0140] The tentacle configuration was reconstructed using a piecewise constant curvature model, for the first Segment, synthesized curvature vector The bending angles of the four SMA wires are superimposed:

[0141] ,

[0142] Each segment is further divided into 5 microsegments, microsegment length Micro-segment rotation increment Rotation angle Rotation Increment Matrix Calculated using the axis-angle formula. Integrating segment by segment from the base, the coordinates of the centerline point of the entire tentacle and the corresponding rotation matrix are obtained;

[0143] Step 2.3: The rotational dynamics equations in the body coordinate system are:

[0144] ,

[0145] in, It is the total moment of inertia tensor (including dry inertia and fluid-added moment of inertia). The angular velocity in the body coordinate system. Undamped torques (thrust torque, gravity / buoyancy torque, contact torque, etc.) To and The hydrodynamic damping torque is proportional to the torque.

[0146] Step 2.4: To avoid numerical divergence of explicit integration under quadratic damping, a semi-implicit integration method is adopted:

[0147] Step 2.4.1: Ignore secondary damping and perform explicit Euler calculations on the driving and gyroscope terms:

[0148] ,

[0149] ,

[0150] in, The rigid body angular acceleration vector is generated by the combined driving torque and gyroscopic torque. For time step;

[0151] Step 2.4.2: Let The equivalent second-order damping coefficient is around Effective inertia in the direction ,in The damping is simplified to a scalar decay equation in this direction. The analytical solution is given as follows:

[0152] ,

[0153] Where ωnew is the three-dimensional angular velocity vector that is finally updated after taking into account the secondary damping;

[0154] Step 2.4.3: Based on the corrected angular velocity vector and time step, update the robot's rotation matrix and attitude using axis angle representation;

[0155] Specifically, time step ;

[0156] Step 3: As Figure 4 As shown, a deep reinforcement learning (PPO) algorithm is used to train the SMA driving strategy of the biomimetic jellyfish to propel its tentacles;

[0157] Step 3.1: Define the agent's observation space and action space:

[0158] Movement space: There are 4 tentacles that can be pushed. Each of the push tentacles independently selects from three discrete actions: 0 - close, 1 - bend upwards, 2 - bend downwards. The policy network output... Dimension logits;

[0159] Observation space: 41-dimensional vector, including the current decision step index normalization, remaining cooldown steps (4-dimensional), previous action (4-dimensional), centroid displacement (3-dimensional), centroid velocity (3-dimensional), 9 elements of the rotation matrix, training round normalization, and fixed segment activity (4-dimensional).

[0160] Step 3.2: Design a multi-objective reward function, which consists of a basic reward for each step and a final state reward;

[0161] Basic reward per step: ,in, This is the base reward value for a single time step. The number of driving tentacles activated in the current step. This represents the total number of robot tentacles. For indicator functions;

[0162] Final state reward: Roll angle Gaussian reward: ,in, The Gaussian reward value matched for the roll angle. It is a natural exponential function. This is the current roll angle of the aircraft. To calculate the arctangent of the roll angle using the second and third elements of the third row of the rotation matrix, The standard deviation is a Gaussian distribution. Parasitic Spin Penalty: ,in, The penalty values ​​are for parasitic rotations in the pitch and yaw directions. This is the absolute value of the pitch angle. The absolute value of the yaw angle; position offset penalty: ,in, This is the penalty value corresponding to the position offset. This represents the current position vector of the biomimetic jellyfish robot. The desired target position vector for the biomimetic jellyfish robot; energy-saving reward: ,in, Energy-saving reward value corresponding to energy consumption optimization; Collision comprehensive evaluation value: ,in, The comprehensive reward / penalty value corresponding to the collision scenario. The number of tentacles that have collided at the current moment;

[0163] Step 3.3: Establish an action masking mechanism:

[0164] The cooling mask includes: the number of cooling steps remaining. The push-tent forces the action to be turned off only, and the cooling counter is set to zero the instant the power is switched from on to off. Then, decrease by 1 for each subsequent decision step;

[0165] Phase mask: During the forced cooling phase, all pushing tentacles are prohibited from bending upwards or downwards;

[0166] Directional prohibition mask: The push tentacles that pushed upwards in the previous decision step prohibit the current step from choosing downwards, and vice versa;

[0167] Step 3.4: Based on structural symmetry and kinematic transmission characteristics, prioritize training the segments with the greatest impact, employing a four-round progressive training strategy, such as... Figure 5 , Figure 6 As shown:

[0168] Round 1: Open the near-end segment Train for 200 rounds;

[0169] Round 2: Opening the 3rd section The optimal move obtained in the first round is fixed;

[0170] Round 3: Opening the 4th section The optimal moves obtained in rounds 1 and 2 are fixed;

[0171] Round 4: Opening remote segments The optimal moves obtained in rounds 1, 2, and 3 are fixed;

[0172] Specifically, within the current round, only the policy network is allowed to update the SMA-driven policy of the open segment. The other fixed segments directly look up the table to call the previous optimal action and do not participate in gradient updates. After each training round, the action matrix (integer matrix, elements ∈ {0,1,2}) with the highest cumulative reward is selected from 200 episodes and saved as the permanent parameters for that round. In the next training round, these permanent parameters are frozen and only used as part of the environment state for the policy network of the newly opened segment to refer to.

[0173] Step 3.5: In each round of progressive training, the standard PPO algorithm is used to iteratively optimize the current open policy network;

[0174] Step 3.5.1: Network structure of the policy network:

[0175] Actor Network: Input layer (41-dimensional) → Fully connected layer (256-dimensional, ReLU) → Fully connected layer (256-dimensional, ReLU) → Output (12-dimensional);

[0176] Critic network: Input layer (41-dimensional) → Fully connected layer (256-dimensional, ReLU) → Fully connected layer (256-dimensional, ReLU) → Output layer (1-dimensional);

[0177] Step 3.5.2: Calculate the dominance function. Within each batch, calculate recursively from the end to the beginning of each time step:

[0178] ,

[0179] ,

[0180] in, For the first The timing difference error at time 10:00. For the first Instant rewards earned at any time As a reward discount factor, The state value function estimated for the Critic network. For the first The state at any given moment, For the first The advantage estimate at time, The attenuation coefficient is the estimate of the generalized advantage. For the first The advantage estimate at that moment;

[0181] Specifically, setting It is 0.99. It is 0.95;

[0182] Step 3.5.3: Loss Function Design:

[0183] Actor strategy loss: ,in, Let the policy loss function of the Actor network be... The number of samples in a training batch. This represents the probability ratio between the current strategy and the old strategy. , For the sample The advantage estimate, For the clipping function, This is the PPO cutting factor. These are the entropy regularization weighting coefficients. The information entropy function of the policy distribution. For the first The probability distribution of the current policy corresponding to each action dimension. For the first The state (observation) corresponding to each training sample. For parameters The current strategy, For the first The sample at the th Actions in each action dimension This is the old strategy with fixed parameters;

[0184] Specifically, It is 0.2. It is 0.1;

[0185] The Critic value loss is calculated using the mean squared error between the state value estimate and the discounted cumulative reward, where the discounted cumulative reward is calculated starting from the corresponding training state.

[0186] Step 3.5.4: Use the Adam optimizer, with the learning rate set to... After collecting the complete trajectory of each episode, the parameters of the Actor and Critic networks are updated in batches.

[0187] Step 3.6: After all four rounds of progressive training are completed, a complete action decision table is finally obtained, in which the row index corresponds to the number of the pushing tentacle (1~4); the column index corresponds to the decision step (1~5); the matrix element is 0, 1 or 2, which indicates the action that the tentacle should perform in that step.

[0188] Step 4: After training, save the optimal action sequence as a strategy file. During keyboard control, directional control is achieved by cyclically mapping the touch index: When the user presses Q + direction keys (W / A / S / D), the system uses the mapping formula:

[0189] ,

[0190] in, This refers to the tentacle index number corresponding to the original action sequence within the strategy file. For modulo operation function, Enter the desired tentacle number after pressing Q + arrow keys on the keyboard. This is the index cycle offset. This represents the total number of tentacles.

[0191] The activation interval in the standard strategy is mapped to the target pushing tentacle, thereby causing the robot to rotate in the specified direction with a rotation period of [missing information]. (forward Power-on drive, then Forced cooling), automatically rotates continuously when pressed and held.

[0192] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0193] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0194] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0195] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0196] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for coordinated control of propulsion and grasping of a biomimetic jellyfish robot, characterized in that, The biomimetic jellyfish robot includes a control center, grasping tentacles, and propulsive tentacles; the coordinated control includes: Obtain the structural and physical parameters of the biomimetic jellyfish robot; A dynamic and kinematic model is constructed based on the structural and physical parameters. Based on the aforementioned dynamics and kinematics model, an action space, observation space, and reward function for deep reinforcement learning are defined to form a decision-making process framework for training policy networks. The action space defines the action types that each pushing tentacle can choose, the observation space defines the environmental and self-state information that the agent can obtain when making decisions, and the reward function is used to quantify the quality of each action. Based on the aforementioned decision-making process framework, a policy network with random network parameters is initialized, and the policy network is trained using the PPO algorithm with multiple discrete actions. The parameters of the policy network are adjusted so that the policy network can maximize the cumulative reward based on the actions of the observed inputs and outputs. After training, the policy network is solidified into an action decision table. The motion decision table is deployed to the robot's control center, and propulsion control commands are generated by combining the real-time collected posture information. Based on the grasping task requirements, the control center independently or collaboratively controls the drive mechanism of the grasping tentacles to execute the grasping action.

2. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 1, characterized in that, The dynamics and kinematics model includes the physical geometry parameters, mass distribution parameters, hydrodynamic parameters, and driving characteristic parameters of the biomimetic jellyfish, which are used to simulate the robot's motion response and force state in a simulation environment.

3. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 1, characterized in that, The construction of the dynamics and kinematics model includes: Each propelling tentacle is divided into several segments, and multiple SMA filaments are embedded in each segment. The composite curvature vector of the segment is calculated by superimposing the bending angles of each SMA filament. Then, each segment is integrated segment by segment to obtain the spatial pose of the entire tentacle, thereby establishing a mapping relationship from the driving amount of the SMA filaments to the position and attitude of the tentacle end. The driving torque acting on the robot body is calculated, and the rotational dynamics equation of the robot body is established by combining the hydrodynamic damping torque and the gyroscopic torque. When solving the rotational dynamics equation, a semi-implicit integration method is adopted. First, the increment of the angular velocity of the driving term and the gyroscopic term is explicitly calculated, and then the hydrodynamic damping term is implicitly processed to maintain numerical stability. Finally, the rotation matrix and attitude of the robot are updated.

4. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 3, characterized in that, The steps for explicitly calculating the driving term are as follows: , , in, The angular acceleration vector generated by the drive, The matrix is ​​the inverse of the moment of inertia tensor. This is the sum of undamped moments. This is the intermediate angular velocity vector after explicit updating by the driving term. This represents the time step for numerical integration. The implicit processing of the hydrodynamic damping term is described in the following formula: , in, To account for hydrodynamic damping corrections, the final updated tentacle angular velocity vector at this time step is... This is the equivalent second-order damping coefficient. , For effective rotational inertia, It is the Euclidean norm.

5. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 1, characterized in that, The action space includes: each pushing tentacle independently selects from 3 discrete actions, namely closing, bending upward and bending downward; The observation space includes at least the current decision step index, the remaining cooling steps for each pushing tentacle, the previous action, the robot's center of mass displacement, center of mass velocity, rotation matrix elements, and training round information; The reward function includes a basic reward at each step and a final state reward. The basic reward at each step is used to suppress invalid activation and collision behavior, and the final state reward includes a roll angle Gaussian reward, a parasitic rotation penalty, a position offset penalty, an energy-saving reward, and a collision comprehensive evaluation value.

6. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 1, characterized in that, The step of training the policy network using the PPO algorithm with multiple discrete actions includes: The policy network collects experience samples, including state, action, reward, and next state, as the data basis for parameter updates. The Actor network takes the observation space as input and outputs the probability distribution of the action space, while the Critic network takes the observation space as input and outputs the state value estimate. Before the action is output, an action masking mechanism is used to mask actions that do not conform to physical constraints. Based on the collected experience samples, the generalized advantage estimation method is used to obtain the advantage function for each action; the temporal difference error is calculated based on the immediate reward, discount factor, and the Critic network's value estimation of the current state and the next state from the experience samples. The policy loss is calculated based on the advantage function, which uses the PPO-clip objective function and incorporates an entropy regularization term to encourage exploration, thereby updating the Actor network parameters; the Critic network parameters are updated based on the temporal difference error.

7. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 6, characterized in that, The training process employs a four-round progressive training strategy, opening and exploring each segment of the propelling tentacle in rounds from near to far: in each round, only one segment is opened for the policy network to learn; the optimal action learned in the previous round is locked in subsequent rounds by fixing the probability distribution in the action space corresponding to that segment or by using the action masking mechanism, and is no longer included in the training; after each round of training, the action sequence with the highest cumulative reward is selected and saved to guide the fixed segment action in the next round of training.

8. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 6, characterized in that, The action masking mechanism includes at least one of a cooling mask, a stage mask, and a direction prohibition mask; The cooling mask includes: the number of remaining cooling steps. The push-tent forces the action to be turned off only, and the cooling counter is set to zero the instant the power is switched from on to off. Then, decrease by 1 for each subsequent decision step; The phase mask includes: during the forced cooling phase, all pushing tentacles are prohibited from bending upwards and downwards; The direction prohibition mask includes: the push tentacles that performed an upward bend in the previous decision step prohibit the current step from selecting a downward bend, and vice versa.

9. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 1, characterized in that, The propulsion control includes: When the control center receives a rotation command, it initiates a rotation cycle that includes a power-on drive period and a forced cooling period. During the power-on drive period, the activation range of the push tentacles is cyclically mapped according to the target rotation direction, and the action decision table is assigned to the push tentacles with the corresponding numbers. During the forced cooling period, all SMA filaments are de-energized, and new activation requests can only be executed after the cooling time constraint is met.

10. The biomimetic jellyfish robot propulsion and grasping coordinated control method according to claim 1, characterized in that, The coordinated control includes: The grasping tentacles and the pushing tentacles are driven by separate drive links. The grasping tentacles are driven by a motor to retract and extend ropes to achieve bending and wrapping, while the pushing tentacles are driven by SMA filaments to achieve propulsion and attitude adjustment. The control center reduces or pauses the propulsion action of the tentacles when performing the grasping task, and resumes the propulsion function after the grasping is completed, so as to achieve the timing coordination of propulsion and grasping. During the grasping process, the grasping parameters are adjusted in real time based on sensor data.