Soft arm multi-arm cooperative control method and device, electronic equipment and storage medium
By acquiring the real-time composite observation state of the soft arm and combining it with the cooperative control model and B-spline curve interpolation, the problems of single perception dimension and poor real-time performance in traditional soft arm control are solved, and smooth, stable and highly reliable spatial continuous cooperative control of the soft multi-arm system is realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PENG CHENG LAB
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-09
Smart Images

Figure CN122165399A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, electronic device and storage medium for multi-arm collaborative control of software arms. Background Technology
[0002] With the deep integration of flexible materials and robotics, soft robots, with their unlimited degrees of freedom, high adaptability, and excellent mechanistic safety, have shown great application potential in complex and confined space operations and human-robot collaboration. In multi-arm collaborative operation scenarios, it is usually necessary to acquire the configuration data of each soft arm in real time and convert it into driving commands to achieve high-precision path planning and target operation. In related technologies, soft arm control usually relies on preset analytical geometric models or simplified offline kinematic lookup tables. Discrete control point commands are generated by matching real-time sensor data with fixed rules, thereby driving the soft arm to complete basic collaborative movements.
[0003] However, due to the physical characteristics of soft robots, such as high nonlinearity, large material deformation, and low damping and easy vibration, traditional rule-driven or simplified model control is difficult to accurately capture the high-dimensional complex state of the soft arm in dynamic operation, resulting in a single perception dimension and poor real-time performance in the multi-arm collaboration process. In addition, since the discrete control commands output in the existing technology ignore the continuous distribution characteristics of torque in physical space when mapped to the soft arm actuator, it is easy to cause abrupt action, end-effector overshoot or significant residual vibration, thus making it impossible to achieve smooth and highly reliable continuous spatial collaborative control while meeting the upper limit constraints of the drive hardware. Summary of the Invention
[0004] This application provides a software arm multi-arm cooperative control method, device, electronic device and storage medium, which can achieve smooth and highly reliable spatial continuous cooperative control while meeting the upper limit constraints of the driving hardware.
[0005] To achieve the above objectives, a first aspect of this application proposes a soft-arm multi-arm cooperative control method, the method comprising: The real-time composite observation state of each soft arm in its current state is obtained, and the real-time composite observation state includes the pose features, kinematic features and multi-arm collaborative sensing features of each soft arm. The real-time composite observation state is input into the collaborative control model for strategy deduction to obtain the continuous motion components corresponding to each soft arm. Based on the drive upper limit parameter corresponding to each of the soft arms, a torque scaling factor is obtained, and the continuous motion components are linearly mapped based on the torque scaling factor to obtain torque control parameters. The torque control parameters are used as the control points of the B-spline curve and input into the interpolation function to generate a spatially continuous torque distribution that is continuously distributed along the axial direction of the soft arm. Each corresponding soft arm is driven to perform coordinated actions based on multiple spatial continuous torque distributions.
[0006] In some embodiments, obtaining the real-time composite observation state of each soft arm in its current state includes: The pose features are obtained based on the body pose vector of each of the soft arms; Based on the order rate of change of the pose vector, the kinematic features used to characterize the corresponding motion trend of the soft arm are calculated. The multi-arm collaborative sensing features are obtained based on the spatial topological association attributes between each of the soft arms. The real-time composite observation state is obtained by performing feature aggregation and alignment based on the pose features, kinematic features, and multi-arm collaborative sensing features.
[0007] In some embodiments, the step of obtaining the collaborative control model includes: Obtain the material nonlinearity parameters, structural geometric parameters, and multi-arm operation space constraints for each soft arm, and construct a multi-arm collaborative simulation environment based on the material nonlinearity parameters, structural geometric parameters, and multi-arm operation space constraints; The observation space, action space, and cooperative reward model of the cooperative control model are determined based on the multi-arm cooperative simulation environment. Based on the observation space, the action space, and the cooperative reward model, an initial cooperative control model is generated. The initial cooperative control model is trained, and the cooperative control model is obtained based on the trained initial cooperative control model.
[0008] In some embodiments, constructing a multi-arm collaborative simulation environment based on the material nonlinear parameters, the structural geometric parameters, and the multi-arm operation space constraints includes: Based on the structural geometric parameters, each soft arm is discretized into a preset number of discrete nodes, and an initial potential energy balance equation between each discrete node is established based on the material nonlinear parameters. Based on the Cosserat rod theory, the motion differential equations of each discrete node in three-dimensional space are established to obtain the nonlinear mechanical model of the soft arm; Based on the multi-arm operation space constraints, an obstacle envelope and a relative coordinate matrix of the base of each soft arm are generated; Configure a position Verlet integral engine, which is used to perform dynamic calculations on the nonlinear mechanical model after the action is performed; The multi-arm collaborative simulation environment is generated based on the initial potential energy balance equation, the nonlinear mechanical model, the relative coordinate matrix, and the position Verlet integration engine.
[0009] In some embodiments, the generation process of the collaborative reward model includes: In the multi-arm collaborative simulation environment, the distance deviation of each soft arm relative to the collaborative target is obtained, and the task guidance reward is obtained based on the inverse proportional function of the distance deviation; Obtain the envelope distance between every two soft arms. When the envelope distance is less than a preset threshold, generate a negative feedback value. Obtain a cooperative obstacle avoidance reward based on the negative feedback value. The drive energy loss of each soft arm during the execution of the action is obtained, and an energy consumption constraint reward is obtained based on the drive energy loss. The collaborative reward model is obtained by weighted summing of the task guidance reward, the collaborative obstacle avoidance reward, and the energy consumption constraint reward.
[0010] In some embodiments, training the initial cooperative control model includes: In each iteration, the initial potential energy balance of the multi-arm collaborative simulation environment is established based on the Cosserat lever theory, and the initial pose of each soft arm is generated. Obtain the real-time composite state vector of each of the software arms in the multi-arm collaborative simulation environment, wherein the real-time composite state vector includes the local tangential basis vector of each discrete unit; The real-time composite state vector is mapped to B-spline muscle torque control points using an actor network, and the position Verlet integrator is invoked to perform dynamic stepping in the multi-arm cooperative simulation environment to obtain the state vector at the next moment and the real-time reward value fed back by the cooperative reward model. A dual-commenter network is used to evaluate the current real-time composite state vector, the B-spline muscle torque control point, and the real-time reward value to obtain the evaluation result. The action network parameters of the actor network and the comment network parameters of the dual commentator network in the initial collaborative control model are updated based on the evaluation results.
[0011] In some embodiments, updating the action network parameters of the actor network and the comment network parameters of the dual commentator network in the initial collaborative control model based on the evaluation results includes: After the comment network parameters of the dual commenter network are updated a preset number of times, the action network parameters of the actor network are updated. The updated action network parameters and comment network parameters are processed by a moving average using target network smoothing technology.
[0012] To achieve the above objectives, a second aspect of this application provides a soft-arm multi-arm collaborative control device, the device comprising: The acquisition module is used to acquire the real-time composite observation state of each soft arm in its current state. The real-time composite observation state includes the pose features, kinematic features, and multi-arm collaborative sensing features of each soft arm. The strategy deduction module is used to input the real-time composite observation state into the collaborative control model to perform strategy deduction and obtain the continuous motion components corresponding to each soft arm. The torque parameter determination module is used to obtain a torque scaling factor based on the drive upper limit parameter corresponding to each of the soft arms, and to perform linear mapping processing on the continuous motion components based on the torque scaling factor to obtain torque control parameters. The torque distribution determination module is used to input the torque control parameters as control points of the B-spline curve into the interpolation function to generate a spatially continuous torque distribution that is continuously distributed along the axial direction of the soft arm. The collaborative control module is used to drive each corresponding soft arm to perform collaborative actions based on the multiple spatial continuous torque distributions.
[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the soft arm multi-arm cooperative control method as described in the first aspect.
[0014] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the software arm multi-arm cooperative control method described in the first aspect.
[0015] The soft arm multi-arm cooperative control method, device, electronic device, and storage medium proposed in this application include: First, acquiring the real-time composite observation state of each soft arm in its current state, including the pose features, kinematic features, and multi-arm cooperative sensing features of each soft arm; then, inputting the real-time composite observation state into a cooperative control model for strategy deduction to obtain the continuous motion components corresponding to each soft arm; second, obtaining a torque scaling factor based on the driving upper limit parameter corresponding to each soft arm, and performing linear mapping processing on the continuous motion components based on the torque scaling factor to obtain torque control parameters; next, using the torque control parameters as control points of a B-spline curve as input to an interpolation function to generate a spatially continuous torque distribution continuously distributed along the axial direction of the soft arm; finally, driving each corresponding soft arm to perform cooperative actions according to the multiple spatially continuous torque distributions. This application's embodiments acquire real-time composite observation states including pose features, kinematic features, and multi-arm collaborative sensing features. This enables real-time and accurate capture of the high-dimensional nonlinear composite states of the soft arm during dynamic operations, effectively solving the problems of single-dimensional sensing and poor real-time performance in traditional rule-driven systems. Simultaneously, by combining torque scaling processing of the driving upper limit parameter with B-spline curve interpolation mechanism, the continuous motion components derived from the model are transformed into a spatially continuous torque distribution continuously distributed along the axis of the soft arm. This not only ensures that the driving commands strictly meet the hardware physical constraints but also overcomes the defects of motion abruptness, end-effector overshoot, and residual vibration caused by the neglect of physical spatial continuity in discrete control commands in the prior art. Thus, smooth, stable, and highly reliable spatially continuous collaborative control of the soft multi-arm system in complex collaborative environments is achieved.
[0016] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description
[0017] Figure 1 This is a flowchart of a soft arm multi-arm collaborative control method provided in an embodiment of this application.
[0018] Figure 2 yes Figure 1 The flowchart for step 101.
[0019] Figure 3 This is a flowchart illustrating the steps for obtaining a collaborative control model according to an embodiment of this application.
[0020] Figure 4 yes Figure 3 The flowchart for step 301.
[0021] Figure 5 yes Figure 3 The flowchart for step 302.
[0022] Figure 6 yes Figure 3 The flowchart for step 304.
[0023] Figure 7 yes Figure 6 The flowchart for step 605.
[0024] Figure 8 This is a performance simulation diagram of a soft arm multi-arm cooperative control method provided in another embodiment of this application.
[0025] Figure 9 This is a performance simulation diagram of another soft arm multi-arm cooperative control method provided in another embodiment of this application.
[0026] Figure 10 This is a schematic diagram of the structure of a soft arm multi-arm collaborative control device provided in an embodiment of this application.
[0027] Figure 11 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0029] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0031] With the deep integration of flexible materials and robotics, soft robots, with their unlimited degrees of freedom, high adaptability, and excellent mechanistic safety, have shown great application potential in complex and confined space operations and human-robot collaboration. In multi-arm collaborative operation scenarios, it is usually necessary to acquire the configuration data of each soft arm in real time and convert it into driving commands to achieve high-precision path planning and target operation. In related technologies, soft arm control usually relies on preset analytical geometric models or simplified offline kinematic lookup tables. Discrete control point commands are generated by matching real-time sensor data with fixed rules, thereby driving the soft arm to complete basic collaborative movements.
[0032] However, due to the physical characteristics of soft robots, such as high nonlinearity, large material deformation, and low damping and easy vibration, traditional rule-driven or simplified model control is difficult to accurately capture the high-dimensional complex state of the soft arm in dynamic operation, resulting in a single perception dimension and poor real-time performance in the multi-arm collaboration process. In addition, since the discrete control commands output in the existing technology ignore the continuous distribution characteristics of torque in physical space when mapped to the soft arm actuator, it is easy to cause abrupt action, end-effector overshoot or significant residual vibration, thus making it impossible to achieve smooth and highly reliable continuous spatial collaborative control while meeting the upper limit constraints of the drive hardware.
[0033] To achieve smooth and highly reliable spatial continuous cooperative control while meeting the upper limit constraints of the driving hardware, this application's embodiments acquire real-time composite observation states including pose features, kinematic features, and multi-arm cooperative sensing features. This enables real-time and accurate capture of the high-dimensional nonlinear composite states of the soft arm during dynamic operations, effectively solving the problems of single sensing dimension and poor real-time performance in traditional rule-driven systems. Simultaneously, by combining torque scaling processing of the driving upper limit parameters with B-spline curve interpolation, the continuous motion components derived from the model are transformed into a spatially continuous torque distribution continuously distributed along the soft arm's axis. This not only ensures that the driving commands strictly meet the hardware physical constraints but also overcomes the defects of sudden motion, end-effector overshoot, and residual vibration caused by the neglect of physical spatial continuity in discrete control commands in existing technologies. Thus, smooth, stable, and highly reliable spatial continuous cooperative control of the soft multi-arm system in complex cooperative environments is achieved.
[0034] The following will first describe in detail the soft-arm multi-arm cooperative control method in the embodiments of this application. (Refer to...) Figure 1 This is an optional flowchart of the soft arm multi-arm cooperative control method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps 101 to 105. It is also understood that this embodiment... Figure 1The order of steps 101 to 105 is not specifically limited; the order of steps can be adjusted or certain steps can be added or removed according to actual needs. The soft arm multi-arm cooperative control method provided in this application embodiment can be applied to any control system (such as a smart terminal, server, control processor, etc.).
[0035] Step 101: Obtain the real-time composite observation state of each soft arm in its current state. The real-time composite observation state includes the pose features, kinematic features, and multi-arm collaborative sensing features of each soft arm.
[0036] Step 101 will be described in detail below.
[0037] In step 101 of some embodiments, the processing system first performs a comprehensive perception and data acquisition of the current physical state of the multi-soft-arm system. This process is not a simple reading of data from a single sensor, but rather the construction of a real-time composite observation state with high-dimensional characteristics. Specifically, the processing system acquires the pose features of each soft arm through embedded sensors (such as fiber Bragg grating sensors) or external vision capture devices. These pose features typically include the Cartesian coordinate positions and tangential direction vectors of multiple discrete nodes distributed along the axis of the soft arm, describing its static geometric configuration in three-dimensional space. Simultaneously, the processing system calculates the first or second derivatives of the pose features based on time-series data, thereby extracting kinematic features reflecting changes in the soft arm's velocity and acceleration. Furthermore, multi-arm cooperative sensing features are obtained by calculating parameters such as the relative distances, relative azimuth angles, and shortest distances to obstacle envelopes between the soft arms, aiming to quantitatively describe the spatial topological relationships and potential collision risks within the multi-arm system. After data alignment and normalization, these three types of features are stitched together to form a high-dimensional real-time composite observation state, providing a precise digital foundation for subsequent control decisions. The specific acquisition is described below.
[0038] Reference Figure 2 To obtain the real-time composite observation status of each soft arm in the current state, the steps 201 to 204 are included.
[0039] Step 201: Obtain pose features based on the body pose vector of each soft arm.
[0040] Step 202: Calculate the kinematic features used to characterize the motion trend of the corresponding soft arm based on the order rate of change of the pose vector.
[0041] Step 203: Based on the spatial topological association attributes between each soft arm, obtain the multi-arm collaborative perception features.
[0042] Step 204: Based on pose features, kinematic features, and multi-arm collaborative sensing features, feature aggregation and alignment are performed to obtain the real-time composite observation state.
[0043] In step 201 of some embodiments, the system performs precise extraction of the "global-local hybrid deformation features" of the soft arm. Specifically, the system extracts the features of the soft arm in real time from physical simulator or sensor data. Cartesian position vectors of discrete nodes To eliminate the interference of base position fluctuations on the policy network, the system not only uses absolute coordinates, but also transforms the positions of all nodes relative to the base nodes. relative coordinate matrix More importantly, to overcome the observation challenges posed by the infinite degrees of freedom of the soft arm, this step specifically extracts the local tangential basis vector from the Cosserat rod theory. This basis vector can sensitively perceive the microscopic bending and torsional posture of each soft unit in space. Through this combination of "relative position + tangential field" features, the system can accurately reconstruct the continuous envelope shape of the soft arm in three-dimensional space in digital space, providing a detailed static configuration description for the control algorithm.
[0044] In step 202 of some embodiments, the system constructs "multi-order temporal dynamics features" to characterize the motion trends of the soft arm. In addition to the static configuration, the observation space must also contain dynamic information. The system obtains the linear velocity of each node by calculating the partial derivatives of the discrete node positions with respect to time. Simultaneously obtain the angular velocity of the local element. For nonlinear systems like soft arms, which exhibit low damping and susceptibility to vibration, introducing these velocity terms is crucial. This allows the backend Critic network to learn the changes in the system's kinetic energy, thereby assisting the Actor network in pre-compensating for inertial terms during decision-making. This kinematic characteristic, incorporating both linear and angular velocities, effectively suppresses end-effector overshoot and residual vibration during high-speed motion or sudden stops, ensuring the dynamic stability of the control process.
[0045] In step 203 of some embodiments, the system extracts "multi-arm implicit cooperative sensing features" as the core increment for cooperative control. This step abandons complex explicit communication protocols and instead directly embeds spatial correlation parameters between the two arms into the observation vector. Specifically, this includes: (a) the Euclidean displacement vector between the end effectors of the two arms. (a) The closest distance of the opposing arm to the centerline of the main arm is used to directly quantify the proximity of the collaborative targets; (b) The closest distance of the opposing arm to the envelope surface of the main arm is used to establish safe spatial constraints. By observing the physical position of the opposing arm and its disturbance to the shared fluid environment (such as fluid loads), the system enables the two arms to achieve "implicit collaboration" like natural organisms. This design not only reduces the communication bandwidth requirements, but also gives the system the ability to adaptively adjust its strategy to avoid collisions and complete collaborative tasks in complex dynamic environments.
[0046] In step 204 of some embodiments, the system executes a high-dimensional composite observation vector. The system employs a refined digital construction method. It aggregates and aligns the global-local hybrid deformation features (relative coordinate matrix and tangential basis vector), multi-order temporal dynamic features (linear velocity and angular velocity), and multi-arm implicit cooperative sensing features acquired in the preceding steps. This process maps the raw sensor data at the code level into composite feature vectors with explicit physical meaning. Through this mathematically refined construction, the final real-time composite observation state is obtained. It not only includes the geometric and dynamic information of the soft arm ontology, but also deeply integrates environmental interaction and cooperative constraint information, providing a fully informative, low-noise and physically interpretable input space for deep reinforcement learning models based on the Actor-Critic architecture.
[0047] Through steps 201 to 204 above, the above steps, by introducing the tangential basis vector and relative coordinate matrix in Cosserat theory, achieve dimensionality reduction and accurate reconstruction of the infinite-degree-of-freedom deformation of the soft arm; by introducing the differential terms of linear velocity and angular velocity, the vibration suppression problem of low-damped systems is solved; by constructing implicit cooperative features such as Euclidean displacement and envelope distance, biological-level adaptive cooperation without communication burden is realized. This composite state space construction based on physical meaning greatly improves the algorithm's perception depth of the nonlinear dynamic characteristics of the soft arm, ensuring high precision, high stability and strong robustness of multi-arm cooperative control in complex environments.
[0048] Step 102: Input the real-time composite observation state into the collaborative control model to perform strategy deduction and obtain the continuous motion components corresponding to each soft arm.
[0049] Step 102 is described in detail below.
[0050] In step 102 of some embodiments, the real-time composite observation state obtained above is used as an input vector and input into a pre-trained cooperative control model for policy inference. This cooperative control model is usually based on a deep neural network architecture (such as an Actor-Critic network), which internally performs feature extraction and feature decoupling on the high-dimensional observation state through multi-layer nonlinear transformations. The policy inference process is that the model calculates the action instruction that maximizes the expected reward based on the currently perceived environmental state and a preset policy gradient or value function. The continuous action components corresponding to each soft arm output in this step are usually a set of dimensionless numerical sequences within a specific normalized numerical range (e.g., [-1, 1]). These components represent the decision intention of the control policy in the abstract action space, such as the activation degree of different bending modes of the soft arm, but at this time they have not yet been converted into driving torques with actual physical units.
[0051] The following section will further describe how to obtain this collaborative control model.
[0052] Reference Figure 3 The steps for obtaining the collaborative control model include steps 301 to 304.
[0053] Step 301: Obtain the material nonlinear parameters, structural geometric parameters, and multi-arm operation space constraints for each soft arm, and construct a multi-arm collaborative simulation environment based on the material nonlinear parameters, structural geometric parameters, and multi-arm operation space constraints.
[0054] Step 301 will be described in detail below.
[0055] In step 301 of some embodiments, the processing system first performs the digital construction of a high-fidelity physical simulation environment. This process begins with the precise parameterization of the physical properties of the soft arms, specifically including acquiring the material nonlinear parameters of each soft arm (e.g., constitutive coefficients, damping coefficients, etc. of a hyperelastic material model), structural geometric parameters (e.g., arm length, cross-sectional radius, chamber distribution, etc.), and multi-arm working space constraints (e.g., obstacle positions, working boundaries, base layout, etc.). Soft arms are typically made of silicone or flexible fabric, and their deformation exhibits highly nonlinear and infinite degrees of freedom characteristics, making simple rigid body models unsuitable. Based on the parameters acquired above, the system reconstructs the dynamic model of the soft arms in the physics engine and sets the corresponding multi-arm working space constraints, thereby constructing a multi-arm collaborative simulation environment that can realistically reflect the mechanical response of the soft arms under large deformations and the spatial interaction relationships between the arms. This simulation environment serves as a "digital twin" field for subsequent algorithm training, enabling the simulation of various extreme working conditions in a low-cost, zero-risk manner, providing physically consistent interactive feedback for the learning of the control model. The specific construction of the multi-arm collaborative simulation environment will be further described below.
[0056] Reference Figure 4 The multi-arm collaborative simulation environment is constructed based on the nonlinear parameters of materials, the geometric parameters of structures, and the spatial constraints of multi-arm operation, including the following steps 401 to 405.
[0057] Step 401: Based on the structural geometric parameters, discretize each soft arm into a preset number of discrete nodes, and establish the initial potential energy balance equation between each discrete node based on the material nonlinear parameters.
[0058] Step 402: Based on the Cosserat rod theory, establish the motion differential equations of each discrete node in three-dimensional space to obtain the nonlinear mechanical model of the soft arm.
[0059] Step 403: Based on the multi-arm operation space constraints, generate the obstacle envelope and the relative coordinate matrix of the base of each soft arm.
[0060] Step 404: Configure the position Verlet integral engine, which is used to perform dynamic calculations on the nonlinear mechanical model after the action is performed.
[0061] Step 405: Generate a multi-arm collaborative simulation environment based on the initial potential energy balance equation, nonlinear mechanical model, relative coordinate matrix, and position Verlet integration engine.
[0062] Steps 401 to 405 are described in detail below.
[0063] In step 401 of some embodiments, the system first performs geometric and physical initialization based on Discrete Cosserat Rod Theory. The system discretizes the continuous soft arm centerline into a structure composed of… Node positions ( )and A chain structure consisting of edge elements connecting nodes. To describe the torsional properties of the soft arm in three-dimensional space, each element is also bound to a local orthogonal basis (Frame). . ( Based on the obtained material nonlinear parameters (such as tensile stiffness) and bending stiffness The system calculates the axial strain of each element. and discrete curvature (through logarithmic mapping) The definition reflects the relationship between two adjacent orthogonal bases. and The rotational differences between the nodes are used to establish the initial potential energy balance equation between each discrete node. This equation ensures that at the initial moment without external force, the soft arm is in a natural drooping state where gravity and internal elastic potential energy are in balance, rather than a non-physical straight line state.
[0064] In step 402 of some embodiments, the system constructs a nonlinear mechanical model describing the dynamic response of the soft arm in a complex environment. This model establishes the motion differential equations for each discrete node based on Cosserat rod theory, specifically following the Newton-Euler equations as shown in the following formula.
[0065]
[0066]
[0067] in, and These are the nodal mass and element rotational inertia matrices, respectively. For the internal forces of the cross section, The internal moment of the cross section, and Here is the material stiffness matrix; The collision force between the arms or between the arm and an obstacle is calculated based on the penalty function method: ( (for penetration depth). The core control term is the driving torque generated by the reinforcement learning strategy. This set of equations precisely couples inertial force, elastic internal force, external environmental force, and control driving force, forming the core of high-fidelity soft arm dynamics.
[0068] In step 403 of some embodiments, the system performs digital constraint modeling of the physical space of the multi-arm collaborative operation. First, based on the spatial constraints of the multi-arm operation, an obstacle envelope is generated, and the contact force is calculated based on the penalty function method. (in (For penetration depth), used to simulate physical collisions between the soft arm and environmental obstacles or another soft arm. More importantly, to eliminate the interference of the absolute position of the base on policy training during simulation, the system generates a relative coordinate matrix of the base for each soft arm. This matrix transforms all node positions relative to the base node. The coordinates were normalized, enabling the constructed simulation environment to focus on the relative deformation and cooperative relationships of the soft arm itself.
[0069] In step 404 of some embodiments, the system is configured with a high-precision time integral solver—the Position Verlet Integrator—to perform explicit evolution of the aforementioned nonlinear mechanical model with second-order accuracy. This integral engine performs dynamic stepping using the following discretization formula.
[0070]
[0071]
[0072] To capture the high-frequency oscillation details of the soft arm during rapid movements, the integration engine typically performs a high number of internal substep operations. This specific integration scheme ensures the energy conservation properties (symplectic geometry) and computational stability of the soft arm during large deformation processes, avoiding numerical divergence.
[0073] In step 405 of some embodiments, the system organically integrates the above components to generate the final multi-arm collaborative simulation environment. This environment starts with the natural form determined by the initial potential energy balance equation and, within a spatial framework defined by the relative coordinate matrix, uses a position Verlet integral engine to drive the evolution of a nonlinear mechanical model over time. In this environment, the soft arm is not only subjected to internal elastic and inertial forces but also responds in real time to the continuous driving torque generated by B-spline interpolation and output by a reinforcement learning algorithm. This creates a closed-loop virtual experimental field with high physical realism and capable of large-scale adaptive training.
[0074] Understandably, in this application, the subsequent reinforcement learning algorithm does not directly control the torque of each discrete unit, but instead outputs a set of control points. The continuous moment distribution of the entire rod can be obtained by interpolation using B-spline curves, as shown in the following formula.
[0075]
[0076] in, for order basis functions, This is the torque scaling factor. This design ensures the smoothness of the action space and reduces the search difficulty for reinforcement learning.
[0077] Through steps 401 to 405 above, by combining Cosserat link theory and Verlet numerical integration algorithm, a high-precision simulation environment was successfully constructed that incorporates both the complex nonlinear mechanical properties of soft arms and the spatial cooperative constraints of multi-arm systems. This environment utilizes discretized nodes and differential equations to solve the technical problem of modeling continuum robots, and reproduces realistic cooperative operation scenarios through precise relative coordinate matrices and obstacle envelopes. This simulation environment construction based on physical mechanisms not only ensures a high degree of consistency between virtual and reality, but also provides a safe, efficient, and physically sound data generation platform for the subsequent large-scale training of cooperative control models.
[0078] Step 302: Determine the observation space, action space, and cooperative reward model of the cooperative control model based on the multi-arm cooperative simulation environment.
[0079] Step 302 will be described in detail below.
[0080] In step 302 of some embodiments, the system further formally defines the three core elements of reinforcement learning based on the constructed multi-arm cooperative simulation environment: observation space, action space, and cooperative reward model. The observation space defines the dimension of state data that the agent can perceive in the simulation environment, typically corresponding to data that real sensors can collect, such as pose, velocity, and contact force. The action space defines the range of control commands that the agent can output, such as the air pressure values of each drive chamber or the tension range of the drive ropes. The cooperative reward model is a quantified mathematical function designed to guide the multi-arm system to complete a specific task. This model typically includes task completion rewards (such as the distance between the end effector and the target), cooperative obstacle avoidance penalties (such as collisions between arms or collisions with the environment), and energy consumption constraints. By clarifying these three definitions, the system transforms the complex physical control problem into a standard Markov Decision Process (MDP), providing the algorithm with a clear optimization objective and interaction interface.
[0081] Specifically, the system, based on a multi-arm cooperative simulation environment, constructs the observation space of the cooperative control model by mapping continuous physical properties to discrete feature vectors that can be processed by neural networks. It is a high-dimensional composite vector designed to comprehensively characterize the instantaneous pose, deformation rate, and environmental interaction features of the two arms. Specifically, the system extracts the distance from the Cosserat arm... Node position vectors And in order to eliminate global displacement interference, through The coordinates are mapped to the base's relative coordinate system. More importantly, in order to sense the bending and torsional strain inside the soft arm, the system also extracts the tangential vector of each unit. and curvature vector This constitutes the deformation sensing component. In addition, the observation space also includes the kinematic sensing component, namely the linear velocity of each node. and local angular velocity This provides dynamic prediction support for the Actor-Critic architecture, assisting the Actor network in compensating for the inherent inertial oscillations of soft materials. As a core feature, the observation vector also specifically embeds multi-arm cooperative sensing components, including the relative displacement vector between the end effectors of the first and second soft arms. The distance between the opposing arm and its nearest neighbor within the collision detection envelope is also considered. This dimension is directly related to the spatial occupancy information of the multi-arm system and is key to achieving implicit cooperation without explicit communication. In addition, there is a task guidance component, which includes the Euclidean distance between the end node and the target point. Its direction vector provides a clear convergence target for the strategy.
[0082] Furthermore, the system defines the action space of the cooperative control model and establishes a direct relationship between the algorithm parameters and the Cosserat link mechanics model. The algorithm outputs the action vector. It does not directly represent joint forces, but is defined as the weight of B-spline control points acting on the centerline of the entire rod. This design establishes a bridge between control signals and physical actuations. When received from the physical environment... Then, a spatially continuous moment distribution function is generated using the B-spline interpolation function. Through this mechanism, the system can drive the Cosserat lever to deform, ensuring that the motion output can effectively drive the soft arm to complete large-amplitude bending tracking. The dimensionless output of the neural network is then mapped to a driving torque in actual physical units. The selection of this parameter is highly coupled with the Young's modulus of the soft material, ensuring that the motion output can effectively drive the soft arm to complete large-amplitude bending tracking, while maintaining the numerical computation stability of the Actor-Critic framework under complex physical constraints.
[0083] Furthermore, the system defines a collaborative reward model, which is a multi-objective weighted function designed to guide the multi-arm system to efficiently complete tasks while satisfying physical constraints. The reward function comprehensively considers the task guidance component (including the Euclidean distance between the end node and the target point). The design includes a direction vector, a cooperative obstacle avoidance component (a penalty term based on the nearest distance to the envelope surface), and an energy consumption constraint component. This design not only provides a clear convergence objective for the policy but also suppresses unsafe collision behavior and high-energy-consuming violent vibrations through the penalty term, thereby achieving a dynamic balance between task accuracy, system safety, and energy efficiency during reinforcement learning, as described below.
[0084] Reference Figure 5The process of generating the collaborative reward model includes the following steps 501 to 504.
[0085] Step 501: In a multi-arm collaborative simulation environment, obtain the distance deviation of each soft arm relative to the collaborative target, and obtain the task guidance reward based on the inverse proportional function of the distance deviation.
[0086] Step 502: Obtain the envelope distance between every two soft arms. When the envelope distance is less than a preset threshold, generate a negative feedback value and obtain a cooperative obstacle avoidance reward based on the negative feedback value.
[0087] Step 503: Obtain the driving energy loss of each soft arm during the execution of the action, and obtain the energy consumption constraint reward based on the driving energy loss.
[0088] Step 504: Weight the task guidance reward, cooperative obstacle avoidance reward, and energy consumption constraint reward to obtain the cooperative reward model.
[0089] Steps 501 to 504 are described in detail below.
[0090] In step 501 of some embodiments, the system constructs a "target approach reward" based on an exponential mapping. This system aims to address the gradient balance problem between far-field navigation and near-field precise positioning for soft arms. In a multi-arm cooperative simulation environment, the system calculates the Euclidean distance between each soft arm's end effector and the cooperative target point in real time. To overcome the shortcomings of traditional linear distance rewards, such as small gradients at long distances and slow convergence at short distances, the system constructs a nonlinear reward function based on distance deviation. .in To adjust the coefficients, this exponential form ensures a smooth guiding gradient when the soft arm is far from the target, while the function value increases sharply when the end effector enters the centimeter-level error tolerance region near the target. This gives the algorithm a significant sparsity bonus, enhancing its micro-manipulation positioning accuracy at the end of the task. Furthermore, the system also incorporates direction vectors. This provides a clear direction of convergence for the policy network.
[0091] In step 502 of some embodiments, the system is designed with a dual security mechanism, including "multi-arm implicit cooperative reward". "and physical collision penalty" First, to address the issue of inconsistent progress in multi-arm tasks, the system calculates the absolute value of the distance difference between each arm and its respective target. As a synchronization progress factor, and based on this, a collaborative reward is constructed. Dynamic penalties are applied to actions that disrupt progress, forcing the Actor network to learn an implicit cooperative strategy that allows it to perceive the state of other actors and adjust its own speed accordingly. Secondly, regarding spatial safety, the system obtains the envelope distance between every two soft arms and the penetration depth of the soft arms through environmental obstacles. When the penetration depth is detected Or, when the internal Cosserat strain exceeds the material's elastic limit, the system triggers high-intensity negative feedback. This design combines the penalty function method in the physics engine with material mechanics constraints, which strictly constrains the exploration space of the strategy from the perspective of physical safety, preventing the soft arm from becoming self-entangled or structurally damaged.
[0092] In step 503 of some embodiments, the system introduces a "control smoothness reward". This system is specifically optimized for the low-damping and vibration-prone characteristics of soft robots. Because soft arms lack the damping dissipation of rigid joints, high-frequency motion switching easily excites oscillations at the structure's natural frequencies. Therefore, the system acquires the driving motion vector for each soft arm during the execution of its actions. And calculate its second rate of change relative to the action vector at the previous moment. An energy-constrained reward is constructed based on this second-order rate of change, aiming to punish drastic changes in motion and ensure the continuity of the driving torque over time. This smooth constraint at the source effectively suppresses high-order modal vibrations excited by high-frequency movements, ensuring the dynamic stability and energy efficiency of the soft arm during high-speed motion.
[0093] In step 504 of some embodiments, the system performs multi-objective fusion to generate the final collaborative reward model. The model is defined as a weighted sum of the aforementioned physical constraints: .in to These are the weight coefficients for each component. The system dynamically adjusts these weights according to the training phase (e.g., increasing them in the initial stage). Weighting is used to accelerate exploration, and is increased in later stages. (Weights are used to optimize control quality). Through this composite reward evaluation system, the system guides the two arms to autonomously learn the optimal collaborative strategy that balances task accuracy, collaborative efficiency, physical safety, and motion smoothness in a complex continuous state space.
[0094] Through steps 501 to 504 above, a composite reward system encompassing three dimensions—target approach, spatial obstacle avoidance, and energy efficiency optimization—effectively solves the problem of multi-target conflict in soft arm control. A clear navigation gradient is provided through the inverse distance function, a strict safety boundary is established through envelope spacing negative feedback, and nonlinear oscillations of the system are suppressed through energy consumption constraints. This multi-dimensional cooperative reward model not only ensures that the multi-arm system can converge to the target state quickly and accurately, but also embeds physical safety and energy efficiency constraints at the algorithm level, thereby training a robust control strategy that possesses both high-precision operational capabilities and meets the long-term stable operation requirements of the physical system.
[0095] Step 303: Generate an initial cooperative control model based on the observation space, action space, and cooperative reward model.
[0096] Step 304: Train the initial cooperative control model and obtain the cooperative control model based on the trained initial cooperative control model.
[0097] Steps 303 to 304 are described in detail below.
[0098] In step 303 of some embodiments, the system constructs an initial cooperative control model based on an improved Actor-Critic architecture (specifically, the dual-delay deep deterministic policy gradient algorithm, TD3) based on a determined observation space and action space. This model mainly includes a policy network (Actor Network). ) and two independent value networks (DualCritic Networks, , The input layer dimension of the Actor network is adapted to high-dimensional composite observation vectors. (Including relative coordinate matrix) Tangential basis vector (and features such as multi-arm collaboration), its output layer dimension corresponds to the dimension of the action space. The Critic network is configured for receive mode. and actions The concatenated vector is used as input to evaluate the expected value of taking a specific cooperative action in the current state. The system processes the network parameters... and Perform random initialization (such as using Xavier initialization) to generate an initial cooperative control model that has not yet converged. At this time, the policy output by the model mainly manifests as exploratory random actions.
[0099] In mathematical description, the Actor network in the initial cooperative control model defines a deterministic mapping from state to action. The output action vector. It is not directly used as a torque value, but is mathematically defined as a set of normalized B-spline control point weights. These weights are then substituted into the B-spline interpolation formula. In this way, the soft arm axis is generated. Continuously distributed driving torque. Meanwhile, the goal of the dual-Critic network is based on a collaborative reward model. To minimize the Bellman error, the value function is estimated. This guides the Actor network to update its parameters to maximize the cumulative expected return.
[0100] In step 304 of some embodiments, the system iteratively trains the initial cooperative control model in a multi-arm cooperative simulation environment. During training, the initial model controls the virtual soft arm to interact with the environment, continuously trying and collecting trajectory data of state, action, reward, and next state. Using deep reinforcement learning algorithms (such as PPO or SAC), the system calculates gradients based on the collected data and backpropagates to update the model's network parameters to maximize the accumulated expected reward. As the number of training rounds increases, the model gradually learns the optimal cooperative strategy to be adopted under various observation states until the reward curve tends to converge smoothly. Finally, based on the network parameters after training convergence, the system obtains a cooperative control model with high robustness and generalization ability, which can be directly deployed or transferred to a real soft arm system for application, as described below.
[0101] Reference Figure 6 The initial collaborative control model is trained, including the following steps 601 to 605.
[0102] Step 601: In each iteration, establish the initial potential energy balance of the multi-arm collaborative simulation environment based on the Cosserat rod theory, and generate the initial pose of each soft arm.
[0103] Step 602: Obtain the real-time composite state vector of each software arm in the multi-arm collaborative simulation environment. The real-time composite state vector includes the local tangential basis vector of each discrete unit.
[0104] Step 603: Use the actor network to map the real-time composite state vector to B-spline muscle torque control points, and call the position Verlet integrator to perform dynamic stepping in the multi-arm cooperative simulation environment to obtain the state vector at the next moment and the real-time reward value fed back by the cooperative reward model.
[0105] Step 604: Use a dual-commenter network to evaluate the current real-time composite state vector, B-spline muscle torque control points, and real-time reward value to obtain the evaluation results.
[0106] Step 605: Update the action network parameters of the actor network and the comment network parameters of the dual commentator network in the initial collaborative control model based on the evaluation results.
[0107] Steps 601 to 605 are described in detail below.
[0108] In step 601 of some embodiments, the system performs a physics-based scene initialization to eliminate dynamic oscillations that may be caused by random initialization. At the beginning of each episode, the system not only generates the obstacle envelope and the base positions of the arms according to a random distribution, but also establishes the initial potential energy balance of the system using discrete Cosserat rod theory. The system calculates the minimum potential energy state of the soft arms under gravity alone, and solves for the node configuration that minimizes the sum of the internal elastic potential energy and the gravitational potential energy. This step ensures the initial pose of the soft arms (including node positions). and cross-sectional orientation It conforms to the real physical characteristics of hanging or stillness, avoiding numerical explosion in the simulator in the first time step caused by non-physical geometric initialization (such as forced stretching of a straight line).
[0109] In step 602 of some embodiments, the system executes a "high-dimensional composite state vector". The algorithm refines the extraction of geometric information. In the perception phase, it goes beyond simple Euclidean coordinates, aggregating deep geometric information that reflects the three-dimensional curvature characteristics of the soft arm. Specifically, the system extracts the local tangential basis vectors of each discrete unit. and curvature vector These vectors can accurately characterize the microscopic bending and torsional strain inside the soft arm, in conjunction with the relative coordinate matrix. and nodal linear velocity Construct the full information state vector This vector serves as the input to the neural network, enabling the algorithm to accurately capture the dynamic configuration of the soft arm under complex deformations, much like perceiving "proprioception."
[0110] In step 603 of some embodiments, the system performs a high-fidelity evolution from neural signals to physical actions. The Actor network receives the state vector. Instead of directly outputting discrete joint torques, it maps them to a set of B-spline muscle torque control points. These control points determine the smooth spatial distribution function of the total moment of the rod. This causes the Cosserat rod to deform. Upon receiving this continuous drive signal, the physical environment responds with a 10... -4The ultra-high sampling step size on the order of seconds invokes the Position Verlet Integrator for dynamic stepping. This explicit evolution scheme with second-order precision realistically simulates the nonlinear mechanical response of the soft arm under large deformation, ensuring energy conservation and providing feedback on the state at the next moment. And the real-time reward value $R_t$ calculated from the multidimensional physical constraints.
[0111] In step 604 of some embodiments, the system introduces a dual-Q estimation mechanism for value assessment. This involves a dual-critic network (i.e., a dual-critic network). , They operate in parallel, simultaneously receiving the current real-time composite state vector. and the motion of B-spline control points generated by the Actor The two networks each output a value prediction (Q-value) for the current state-action pair. To suppress the Q-value overestimation problem common in deep reinforcement learning, the system typically selects the smaller of the two network outputs. This serves as the final evaluation result. This mechanism provides a more conservative and robust value orientation for policy learning, preventing the policy from falling into local optima due to erroneous optimistic estimates.
[0112] In step 605 of some embodiments, the system performs parameter optimization based on delayed policy updates and target network smoothing. The system stores trajectory tuples (State, Action, Reward, Next State) containing complex physical interaction information into a large-capacity experience pool. During the parameter update phase, the system does not update all networks synchronously each time; instead, it updates the Critic network more frequently than the Actor network (e.g., updating the Actor only once every d updates to the Critic network), thus ensuring that the policy is improved based on relatively accurate value assessment. Simultaneously, a soft update formula is utilized. The target network parameters are processed by a moving average. This process is repeated until the multidimensional reward function, which includes task guidance, cooperative obstacle avoidance, and energy consumption constraints, tends to converge. Finally, a control model with autonomous obstacle avoidance and high-precision cooperative capabilities in complex environments is derived, as described below.
[0113] Reference Figure 7 The action network parameters of the actor network and the comment network parameters of the dual commentator network in the initial collaborative control model are updated based on the evaluation results, including the following steps 701 to 702.
[0114] Step 701: After updating the comment network parameters of the dual commenter network a preset number of times, update the action network parameters of the actor network.
[0115] Step 702: Use target network smoothing technology to perform a moving average on the updated action network parameters and comment network parameters.
[0116] Steps 701 to 702 are described in detail below.
[0117] In step 701 of some embodiments, the system employs a delayed policy update mechanism to coordinate the training pace between the actor network and the dual reviewer network. Specifically, during training iterations, the system does not synchronously update the action network parameters of the actor network at every time step. Instead, it sets a preset update frequency or number threshold (e.g., updating the actor network only once every two or three updates to the reviewer network). During this period, the system prioritizes using the latest empirical data to frequently optimize the reviewer network parameters of the dual reviewer network, enabling its value evaluation function (Q-function) to converge quickly to a low-error state. Only when the value evaluation of the dual reviewer network becomes stable and accurately reflects the true value of the current state-action pair does the system update the action network parameters of the actor network based on this more accurate evaluation result. This asynchronous update strategy effectively avoids the actor network optimizing based on incorrect gradients before the reviewer network converges, thereby suppressing policy oscillations and divergences in the early stages of training and ensuring the correctness of the policy improvement direction.
[0118] In step 702 of some embodiments, to further enhance the stability of the training process and suppress the variance of value estimation, the system uses Target Network Smoothing (often also called Soft Update) to post-process the updated network parameters. This technique involves two parameter systems: the main network (Online Network) and the target network (Target Network). In this step, the system does not directly hard-copy the parameters of the target network. Instead, it introduces a very small smoothing coefficient (e.g., 0.005) to perform a weighted moving average of the newly updated action network parameters and comment network parameters with the corresponding historical parameters of the target network. Specifically, the new target network parameters are equal to the sum of "smoothing coefficient multiplied by main network parameters" and "(1 - smoothing coefficient) multiplied by old target network parameters". Through this moving average processing, the magnitude of parameter changes in the target network is strictly limited, making its follow-up trajectory with the main network update smooth and lagging. This can effectively cut off the positive feedback loop in action value estimation, prevent training collapse due to rapid changes in Q-values, and ensure that the collaborative control model maintains a convergence trend during long-term training.
[0119] Through steps 701 and 702 above, a highly robust parameter optimization closed loop is constructed by introducing delayed policy updates and target network smoothing techniques. The delayed update mechanism ensures that "evaluation accuracy precedes decision optimization," eliminating policy misguidance caused by value assessment bias. Meanwhile, the target network smoothing technique, through low-pass filtering-style parameter iteration, suppresses the Q-value overestimation and training oscillation problems common in deep reinforcement learning from the algorithm's underlying layer. The combination of these two techniques enables the soft arm multi-arm cooperative control model to achieve stable convergence in high-dimensional, nonlinear, and complex state spaces, ultimately generating a cooperative control strategy with high accuracy and robustness.
[0120] Through steps 601 to 605 above, a high-fidelity physics-driven training closed loop is constructed by deeply coupling the Cosserat rod mechanics theory with the Actor-Critic deep reinforcement learning framework. The scheme in this application enhances the model's ability to perceive large soft deformations by introducing local tangential basis vectors; the combination of a position Verlet integrator and B-spline control points ensures that the dynamic evolution during training conforms to the law of energy conservation and possesses smooth motion; and the application of a dual-critic network effectively solves the problems of training instability and convergence difficulties caused by the high-dimensional state space in soft robot control, ultimately generating a control model with high robustness and precise collaborative capabilities.
[0121] Through steps 301 to 304 above, a systematic construction process from physical parameters to simulation environment and then to the definition of reinforcement learning elements effectively solves the control challenges of soft robots caused by strong nonlinearity and difficult modeling. By constructing a high-fidelity simulation environment that includes material nonlinearity and geometric constraints, this application ensures the physical realism of the training scenario; through a scientifically defined observation and action space and a multi-objective cooperative reward model, the model is guided to autonomously learn the optimal cooperative strategy in complex environments. This training paradigm based on simulation-reality transfer not only significantly reduces the time cost and hardware wear risk of direct training on physical robots, but also enables the generated control model to achieve high-precision and high-safety multi-arm cooperative operation capabilities in dynamic unstructured environments.
[0122] Step 103: Based on the upper limit parameter of the drive corresponding to each soft arm, obtain the torque scaling factor, and perform linear mapping processing on the continuous motion components based on the torque scaling factor to obtain the torque control parameters.
[0123] Step 103 will be described in detail below.
[0124] In step 103 of some embodiments, to convert the abstract numerical values output by the model into actual instructions capable of driving the physical hardware, the system performs a physical parameter mapping operation. First, the system obtains the upper limit parameter of the drive actuator (such as a pneumatic tendon or motor drive line) equipped with each soft arm. This parameter defines the maximum torque or maximum pressure value that the hardware can provide within its safe operating range. Based on this upper limit parameter, the system determines the torque scaling factor. This factor is essentially a physical dimension conversion coefficient. Subsequently, the system uses this torque scaling factor to perform linear mapping (e.g., multiplication) on the dimensionless continuous motion components obtained in step 102, thereby obtaining torque control parameters with definite physical units (such as Newton-meters or Pascals). This step ensures that the generated control commands are strictly within the executable domain of the hardware, preventing driver saturation or soft arm damage due to excessively large commands.
[0125] Step 104: Input the torque control parameters as the control points of the B-spline curve into the interpolation function to generate a spatially continuous torque distribution that is continuously distributed along the axis of the soft arm.
[0126] Step 104 is described in detail below.
[0127] In step 104 of some embodiments, the system introduces a spatial interpolation algorithm to adapt to the continuum characteristics of the soft arm. Since the soft arm is a continuous medium with infinite degrees of freedom in mechanics, discrete control points are difficult to use for smooth deformation control. Therefore, in this step, the system uses the aforementioned torque control parameters with physical dimensions as control points for a B-spline curve. A B-spline curve is a parametric curve fitting tool based on basis functions, possessing local support and continuity. The system inputs these control points into a preset interpolation function, using the arc length of the soft arm's axis as the independent variable, to fit and generate a spatially continuous torque distribution curve continuously distributed along the axis of the soft arm. This interpolation function is shown below.
[0128]
[0129] This distribution curve describes the variation of the driving torque along the entire length of the soft arm, thus extending the finite-dimensional control parameters into an infinite-dimensional spatial continuous field.
[0130] Step 105: Drive each corresponding soft arm to perform coordinated actions based on multiple spatial continuous torque distributions.
[0131] Step 105 is described in detail below.
[0132] In step 105 of some embodiments, the system drives each corresponding soft arm to perform cooperative actions based on multiple calculated spatial continuous torque distributions. In actual physical systems, this typically involves discretizing the spatial continuous torque distributions into specific control signals (such as pressure values in each air chamber or tension values in each cable) corresponding to each drive chamber or drive unit of the soft arm. Since the drive signals originate from smooth B-spline distribution curves, each soft arm can exhibit compliant motion characteristics with uniform energy distribution and continuous curvature changes during execution. Under the drive of their respective continuous torques, multiple soft arms can autonomously adjust their posture according to real-time perceived cooperative characteristics, efficiently completing cooperative tasks such as multi-point grasping, cooperative handling, or complex path tracking while avoiding collisions.
[0133] Through steps 101 to 105 above, by constructing a real-time composite observation state encompassing multi-dimensional physical information, the problem of the single perception dimension of soft robots is solved; through linear mapping based on the upper limit of the drive, it is ensured that the strategy instructions generated by the artificial intelligence algorithm can be safely implemented in the physical hardware; more importantly, by transforming discrete motion components into a spatially continuous torque distribution parameterized by B-spline, this invention overcomes the stress concentration and abrupt motion problems that are prone to occur in soft arms under discrete control, realizes smooth drive that conforms to the mechanical properties of continuous media, and significantly improves the control accuracy, response speed and system stability of multi-soft-arm systems in complex collaborative operations.
[0134] To further verify the reliability of the soft arm multi-arm cooperative control method provided in this application, the algorithm was validated using a high-fidelity physics simulation platform based on the Elastica engine. The simulation environment discretized each soft arm into... Each force-bearing unit was designed, and its Young's modulus and density were set according to the characteristics of silicone rubber-like materials. The experiment was set in a three-dimensional space containing dynamic obstacles, requiring both arms to collaboratively track an asymmetric spatial trajectory to test the algorithm's ability to perceive continuous deformation and spatial occupancy.
[0135] Reference Figure 8 This is a performance simulation diagram of a soft arm multi-arm cooperative control method provided in an embodiment of this application. Figure 8 The figure illustrates the real-time motion modes and cooperative tracking dynamics of the two soft arms driven by a B-spline moment field. The solid colored line represents the centerline configuration of the soft arms, and its smooth curvature distribution is highly consistent with the established Cosserat rod mechanical model, verifying the high fidelity of the simulation environment in modeling the nonlinear deformation of the soft arms. The figure specifically captures a keyframe of the two arms bypassing the central blue obstacle, vividly demonstrating the proposed "implicit cooperation" mechanism: when one soft arm (e.g., the orange arm) actively changes its bending plane to avoid the obstacle, the other soft arm (e.g., the blue arm) can capture the spatial positional shift of the other in real time through the "relative displacement vector" in the state space. Based on this perception, the blue arm synchronously adjusts its own moment distribution, maintaining extremely high tracking accuracy of the target trajectory while completely avoiding physical interference between the two arms, proving the effectiveness of the algorithm in dynamically constrained spaces.
[0136] Reference Figure 9 This is a performance simulation diagram of another soft-arm multi-arm cooperative control method provided in the embodiments of this application. Figure 9 The diagram shows the evolutionary process and error convergence curve of the cooperative control algorithm over millions of physical steps. The horizontal axis represents the training time step (i.e., training steps × 10). 5 The vertical axis represents the end-point tracking error. In the early stages of training (approximately the first 2×10⁻⁶), the vertical axis represents the end-point tracking error. 5 With each step (up to 8×10), the error curve fluctuates dramatically, reflecting the algorithm's exploration of complex physical boundaries and attempts to avoid collision penalties. 5After the first step, the curve enters a stable high-gain range, and the average tracking error at the end point decreases significantly and converges to the millimeter level. This convergence trend strongly demonstrates that the "synchronization progress factor" proposed in this invention can effectively compress the high-dimensional search space and solve the common time delay problem in soft robot control; while the extremely small variance after steady state further illustrates that the framework based on the delayed update strategy can provide a highly stable control signal for soft collaboration.
[0137] This application also provides a soft arm multi-arm collaborative control device, which can realize the above-mentioned soft arm multi-arm collaborative control method, see reference. Figure 10 The device 1000 includes: The acquisition module 1010 is used to acquire the real-time composite observation state of each soft arm in the current state. The real-time composite observation state includes the pose features, kinematic features and multi-arm collaborative perception features of each soft arm. The strategy deduction module 1020 is used to input the real-time composite observation state into the collaborative control model to perform strategy deduction and obtain the continuous motion components corresponding to each soft arm. The torque parameter determination module 1030 is used to obtain the torque scaling factor based on the drive upper limit parameter corresponding to each soft arm, and to perform linear mapping processing on the continuous motion components based on the torque scaling factor to obtain the torque control parameters. The torque distribution determination module 1040 is used to input the torque control parameters as the control points of the B-spline curve into the interpolation function to generate a spatially continuous torque distribution that is continuously distributed along the axial direction of the soft arm. The collaborative control module 1050 is used to drive each corresponding soft arm to perform collaborative actions based on multiple spatial continuous torque distributions.
[0138] In some embodiments, the acquisition module 1010 is further configured to: Based on the body pose vector of each soft arm, pose features are obtained; Based on the order rate of change of the pose vector, the kinematic features used to characterize the motion trend of the corresponding soft arm are calculated. Based on the spatial topological association attributes between each soft arm, multi-arm collaborative perception features are obtained; Based on pose features, kinematic features, and multi-arm collaborative sensing features, feature aggregation and alignment are performed to obtain the real-time composite observation state.
[0139] In some embodiments, the strategy deduction module 1020 is further configured to: Obtain the material nonlinearity parameters, structural geometric parameters, and multi-arm operation space constraints for each soft arm, and construct a multi-arm collaborative simulation environment based on the material nonlinearity parameters, structural geometric parameters, and multi-arm operation space constraints; The observation space, action space, and cooperative reward model of the cooperative control model are determined based on the multi-arm cooperative simulation environment. An initial cooperative control model is generated based on the observation space, action space, and cooperative reward model. The initial cooperative control model is trained, and the cooperative control model is obtained based on the trained initial cooperative control model.
[0140] In some embodiments, the strategy deduction module 1020 is further configured to: Based on the structural geometric parameters, each soft arm is discretized into a preset number of discrete nodes, and an initial potential energy balance equation between each discrete node is established based on the material nonlinear parameters. Based on the Cosserat rod theory, the motion differential equations of each discrete node in three-dimensional space are established to obtain the nonlinear mechanical model of the soft arm; Based on the multi-arm operation space constraints, generate the obstacle envelope and the relative coordinate matrix of the base of each soft arm; Configure the position Verlet integral engine, which is used to solve the dynamics of the nonlinear mechanical model after the action is performed; A multi-arm collaborative simulation environment is generated based on the initial potential energy balance equation, nonlinear mechanical model, relative coordinate matrix, and position Verlet integration engine.
[0141] In some embodiments, the strategy deduction module 1020 is further configured to: In a multi-arm cooperative simulation environment, the distance deviation of each soft arm relative to the cooperative target is obtained, and the task guidance reward is obtained based on the inverse proportional function of the distance deviation. Obtain the envelope distance between every two soft arms. When the envelope distance is less than a preset threshold, generate a negative feedback value. Obtain a cooperative obstacle avoidance reward based on the negative feedback value. Obtain the driving energy loss of each soft arm during the execution of the action, and obtain the energy consumption constraint reward based on the driving energy loss; The collaborative reward model is obtained by weighting and summing the task guidance reward, collaborative obstacle avoidance reward, and energy consumption constraint reward.
[0142] In some embodiments, the strategy deduction module 1020 is further configured to: In each iteration, the initial potential energy balance of the multi-arm collaborative simulation environment is established based on the Cosserat rod theory, and the initial pose of each soft arm is generated. Obtain the real-time composite state vector of each soft arm in the multi-arm collaborative simulation environment. The real-time composite state vector includes the local tangential basis vector of each discrete unit. The real-time composite state vector is mapped to B-spline muscle torque control points using an actor network, and the position Verlet integrator is invoked to perform dynamic stepping in a multi-arm cooperative simulation environment to obtain the state vector at the next moment and the real-time reward value fed back by the cooperative reward model. A dual-commenter network is used to evaluate the current real-time composite state vector, B-spline muscle torque control points, and real-time reward value, and the evaluation results are obtained. The action network parameters of the actor network and the comment network parameters of the dual commentator network in the initial collaborative control model are updated based on the evaluation results.
[0143] In some embodiments, the strategy deduction module 1020 is further configured to: After updating the comment network parameters of the dual commenter network a preset number of times, the action network parameters of the actor network are updated. The target network smoothing technique is used to perform a moving average on the updated action network parameters and comment network parameters.
[0144] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, the specific implementation of the soft arm multi-arm collaborative control device is basically the same as the specific implementation of the soft arm multi-arm collaborative control method described above, and will not be repeated here.
[0145] This application also provides an electronic device, including: At least one memory; At least one processor; At least one program; The program is stored in a memory, and the processor executes the at least one program to implement the software arm multi-arm collaborative control method described above. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.
[0146] Please see Figure 11 , Figure 11 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1101 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1102 can be implemented in the form of ROM (Read-Only Memory), static storage device, dynamic storage device, or RAM (Random Access Memory). The memory 1102 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1102 and is called and executed by the processor 1101 to execute the software arm multi-arm cooperative control method of the embodiments of this application. Input / output interface 1103 is used to implement information input and output; The communication interface 1104 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1105 transmits information between various components of the device (e.g., processor 1101, memory 1102, input / output interface 1103, and communication interface 1104); The processor 1101, memory 1102, input / output interface 1103 and communication interface 1104 are connected to each other within the device via bus 1105.
[0147] This application embodiment also provides a storage medium, which is a computer-readable storage medium, storing a computer program that, when executed by a processor, implements the above-described software arm multi-arm cooperative control method.
[0148] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0149] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0150] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0152] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0153] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0154] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0155] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.
[0156] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0158] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0159] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for multi-arm collaborative control of a soft arm, characterized in that, The method includes: The real-time composite observation state of each soft arm in its current state is obtained, and the real-time composite observation state includes the pose features, kinematic features and multi-arm collaborative sensing features of each soft arm. The real-time composite observation state is input into the collaborative control model for strategy deduction to obtain the continuous motion components corresponding to each soft arm. Based on the drive upper limit parameter corresponding to each of the soft arms, a torque scaling factor is obtained, and the continuous motion components are linearly mapped based on the torque scaling factor to obtain torque control parameters. The torque control parameters are used as the control points of the B-spline curve and input into the interpolation function to generate a spatially continuous torque distribution that is continuously distributed along the axial direction of the soft arm. Each corresponding soft arm is driven to perform coordinated actions based on multiple spatial continuous torque distributions.
2. The soft arm multi-arm cooperative control method according to claim 1, characterized in that, The process of obtaining the real-time composite observation state of each soft arm in its current state includes: The pose features are obtained based on the body pose vector of each of the soft arms; Based on the order rate of change of the pose vector, the kinematic features used to characterize the corresponding motion trend of the soft arm are calculated. The multi-arm collaborative sensing features are obtained based on the spatial topological association attributes between each of the soft arms. The real-time composite observation state is obtained by performing feature aggregation and alignment based on the pose features, kinematic features, and multi-arm collaborative sensing features.
3. The soft arm multi-arm cooperative control method according to claim 1, characterized in that, The steps for obtaining the collaborative control model include: Obtain the material nonlinearity parameters, structural geometric parameters, and multi-arm operation space constraints for each soft arm, and construct a multi-arm collaborative simulation environment based on the material nonlinearity parameters, structural geometric parameters, and multi-arm operation space constraints; The observation space, action space, and cooperative reward model of the cooperative control model are determined based on the multi-arm cooperative simulation environment. Based on the observation space, the action space, and the cooperative reward model, an initial cooperative control model is generated. The initial cooperative control model is trained, and the cooperative control model is obtained based on the trained initial cooperative control model.
4. The soft arm multi-arm cooperative control method according to claim 3, characterized in that, The construction of a multi-arm collaborative simulation environment based on the material nonlinear parameters, the structural geometric parameters, and the multi-arm operation space constraints includes: Based on the structural geometric parameters, each soft arm is discretized into a preset number of discrete nodes, and an initial potential energy balance equation between each discrete node is established based on the material nonlinear parameters. Based on the Cosserat rod theory, the motion differential equations of each discrete node in three-dimensional space are established to obtain the nonlinear mechanical model of the soft arm; Based on the multi-arm operation space constraints, an obstacle envelope and a relative coordinate matrix of the base of each soft arm are generated; Configure a position Verlet integral engine, which is used to perform dynamic calculations on the nonlinear mechanical model after the action is performed; The multi-arm collaborative simulation environment is generated based on the initial potential energy balance equation, the nonlinear mechanical model, the relative coordinate matrix, and the position Verlet integration engine.
5. The soft arm multi-arm cooperative control method according to claim 3, characterized in that, The generation process of the collaborative reward model includes: In the multi-arm collaborative simulation environment, the distance deviation of each soft arm relative to the collaborative target is obtained, and the task guidance reward is obtained based on the inverse proportional function of the distance deviation; Obtain the envelope distance between every two soft arms. When the envelope distance is less than a preset threshold, generate a negative feedback value. Obtain a cooperative obstacle avoidance reward based on the negative feedback value. The drive energy loss of each soft arm during the execution of the action is obtained, and an energy consumption constraint reward is obtained based on the drive energy loss. The collaborative reward model is obtained by weighted summing of the task guidance reward, the collaborative obstacle avoidance reward, and the energy consumption constraint reward.
6. The soft arm multi-arm cooperative control method according to claim 3, characterized in that, The step of training the initial cooperative control model includes: In each iteration, the initial potential energy balance of the multi-arm collaborative simulation environment is established based on the Cosserat lever theory, and the initial pose of each soft arm is generated. Obtain the real-time composite state vector of each of the software arms in the multi-arm collaborative simulation environment, wherein the real-time composite state vector includes the local tangential basis vector of each discrete unit; The real-time composite state vector is mapped to B-spline muscle torque control points using an actor network, and the position Verlet integrator is invoked to perform dynamic stepping in the multi-arm cooperative simulation environment to obtain the state vector at the next moment and the real-time reward value fed back by the cooperative reward model. A dual-commenter network is used to evaluate the current real-time composite state vector, the B-spline muscle torque control point, and the real-time reward value to obtain the evaluation result. The action network parameters of the actor network and the comment network parameters of the dual commentator network in the initial collaborative control model are updated based on the evaluation results.
7. The soft arm multi-arm cooperative control method according to claim 6, characterized in that, The process of updating the action network parameters of the actor network and the comment network parameters of the dual commentator network in the initial collaborative control model based on the evaluation results includes: After the comment network parameters of the dual commenter network are updated a preset number of times, the action network parameters of the actor network are updated. The updated action network parameters and comment network parameters are processed by a moving average using target network smoothing technology.
8. A soft-arm multi-arm collaborative control device, characterized in that, The device includes: The acquisition module is used to acquire the real-time composite observation state of each soft arm in its current state. The real-time composite observation state includes the pose features, kinematic features, and multi-arm collaborative sensing features of each soft arm. The strategy deduction module is used to input the real-time composite observation state into the collaborative control model to perform strategy deduction and obtain the continuous motion components corresponding to each soft arm. The torque parameter determination module is used to obtain a torque scaling factor based on the drive upper limit parameter corresponding to each of the soft arms, and to perform linear mapping processing on the continuous motion components based on the torque scaling factor to obtain torque control parameters. The torque distribution determination module is used to input the torque control parameters as control points of the B-spline curve into the interpolation function to generate a spatially continuous torque distribution that is continuously distributed along the axial direction of the soft arm. The collaborative control module is used to drive each corresponding soft arm to perform collaborative actions based on the multiple spatial continuous torque distributions.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the soft arm multi-arm cooperative control method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the soft arm multi-arm cooperative control method according to any one of claims 1 to 7.