Method and device for optimizing grabbing capacity of musculoskeletal mechanical arm and storage medium

By collaboratively optimizing muscle morphology and control strategies, and employing Bayesian optimization and Gaussian process models, the problems of high computational cost and poor morphological adaptability of musculoskeletal robots in grasping complex tasks were solved, achieving efficient and robust grasping capabilities.

CN121374582APending Publication Date: 2026-01-23TSINGHUA UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511652910.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing musculoskeletal robots suffer from high computational costs, poor shape adaptability, and insufficient generalization ability when grasping complex tasks, making it difficult to cope with objects of different weights and shapes.

Method used

By co-optimizing muscle morphology and control strategy, Bayesian optimization method is used to adjust muscle stiffness parameters, combined with Gaussian process modeling reward function, and performance indicators are collected in real time for feedback, thereby achieving co-optimization of the morphology and control strategy of the musculoskeletal robotic arm.

Benefits of technology

It improves the success rate and computational efficiency of grasping tasks, reduces the number of iterations, enhances adaptability to objects of different weights and shapes, and has efficient generalization ability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374582A_ABST
    Figure CN121374582A_ABST
Patent Text Reader

Abstract

The invention discloses a method for optimizing the grabbing capacity of a musculoskeletal mechanical arm. The method comprises the steps that a rigidity parameter space of the musculoskeletal mechanical arm is constructed; a plurality of initial rigidity parameter sets are generated based on the space, reward functions of the rigidity parameter sets are modeled into a Gaussian process, preliminary training is conducted on a grabbing control strategy based on the initial rigidity parameter sets and energy constraint conditions, initial reward values are obtained, and the rigidity parameter sets and the initial reward values of the rigidity parameter sets are stored in an observation data set; storing the rigidity parameter set into a rigidity parameter pool; selecting a to-be-optimized rigidity parameter set from the rigidity parameter pool through the expectation improvement function, training the grabbing control strategy based on the to-be-optimized rigidity parameter set and the energy constraint condition, updating the Gaussian process, the observation data set and the rigidity parameter pool, and continuously repeating the step until the iteration termination condition is reached, so as to obtain the optimal grabbing control strategy. And an optimal rigidity parameter and a grabbing control strategy are obtained. According to the method, the success rate and generalization ability of the grabbing task can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot control and biomechanical optimization, specifically relating to a method, device and storage medium for optimizing the grasping ability of a musculoskeletal robotic arm. Background Technology

[0002] In the field of musculoskeletal robotics, existing technologies mainly focus on optimizing control strategies (such as reinforcement learning algorithms). However, when faced with complex tasks (such as grasping heavy objects or irregular objects), relying solely on control strategies has the following problems:

[0003] 1. High computational cost: Traditional methods require repeated training of control strategies to adapt to different tasks, which is time-consuming and resource-intensive;

[0004] 2. Poor shape adaptability: Fixed muscle parameters are unable to cope with drastic changes in the weight or shape of objects, leading to grasping failure;

[0005] 3. Insufficient generalization ability: Existing research is mostly aimed at standardized task scenarios (such as objects with fixed weight), which cannot meet the needs of practical applications.

[0006] Furthermore, evolutionary methods such as genetic algorithms are often used in the field of morphological optimization, but these rely on random search, resulting in low efficiency and difficulty in handling high-dimensional parameter spaces. For example, genetic algorithms require a large number of iterations to evaluate candidate parameters, leading to a surge in computational costs. Therefore, there is an urgent need for an efficient and low-cost optimization method to achieve coordinated optimization of morphology and control in musculoskeletal systems. Summary of the Invention

[0007] The present invention aims to at least partially solve one of the technical problems in the related art.

[0008] Therefore, the purpose of this invention is to provide a method, device and storage medium for optimizing the grasping ability of a musculoskeletal robotic arm. By synergistically optimizing muscle morphology (such as stiffness parameters) and control strategies, the grasping ability of the bionic robotic arm to grasp objects of different weights and shapes can be improved, breaking through the performance bottleneck of traditional single optimization dimension, and significantly improving task success rate and computational efficiency.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] The first aspect of this invention provides a method for optimizing the grasping ability of a musculoskeletal robotic arm, comprising:

[0011] Step S100: The shape of the musculoskeletal robotic arm is represented by the stiffness parameter space of the musculoskeletal robotic arm, with the total energy consumption of all muscles in the entire grasping process not exceeding a preset threshold as the energy constraint condition.

[0012] Step S200: Generate several initial stiffness parameter sets based on the stiffness parameter space, model the reward function of the stiffness parameter sets as a Gaussian process, perform preliminary training on the grasping control strategy of the musculoskeletal robotic arm based on each initial stiffness parameter set and under the energy constraint condition, obtain the initial reward value of each stiffness parameter set, store each stiffness parameter set and its initial reward value as a pair of observation data in the observation dataset, and store the stiffness parameter sets in the stiffness parameter pool.

[0013] Step S300: Select a set of stiffness parameters to be optimized from the stiffness parameter pool through the expected improvement function, train the grasping control strategy based on the set of stiffness parameters to be optimized under energy constraints, and update the Gaussian process, the observation dataset and the stiffness parameter pool.

[0014] Step S400: Repeat step S300 several times until the iteration termination condition is met, and obtain the optimal stiffness parameter set and the optimal grasping control strategy.

[0015] In some embodiments, in step S100, the stiffness parameter space is: , This represents the stiffness parameter of the i-th muscle in the robotic arm, set as follows: The range of values ​​for;

[0016] The expression for the energy constraint is:

[0017]

[0018] In the formula, This represents the total energy consumption of all muscles in the robotic arm during the entire grasping process; N represents the number of muscles contained in the robotic arm. This represents the equivalent damping force of the i-th muscle during the grasping process; Indicates the duration of the grabbing action; This indicates the maximum energy consumption threshold that is set.

[0019] In some embodiments, step S200 includes:

[0020] Step S210: Perform layered sampling on the stiffness parameter space to generate M initial stiffness parameter sets. The dimensions of each stiffness parameter set are the same as the dimensions of the stiffness parameter space;

[0021] Step S220: The reward function of the stiffness parameter set based on the parameters set in step S100. Modeled as a Gaussian process, and using a zero-mean function. and kernel function Initialize the Gaussian process GP, i.e., satisfy: The kernel function for the Gaussian process GP uses radial basis functions, which are defined as follows:

[0022]

[0023] in, They represent the first A set of stiffness parameters; Indicates the signal variance; Describe a free parameter, Noise level; For the Kroneck delta function, when the set of stiffness parameters With stiffness parameter set When they are the same, take When the set of stiffness parameters With stiffness parameter set When they are different, take ;

[0024] Step S230: Based on the set of stiffness parameters The grasping control strategy was initially trained using reinforcement learning methods, and grasping tasks were performed. After the initial training was completed, the set of stiffness parameters was recorded. Reward value ;

[0025] Step S240: Transfer M pairs of observation data Store the observation dataset D, and simultaneously transfer the M stiffness parameter sets { Store it in the stiffness parameter pool P.

[0026] In some embodiments, the stratified sampling employs the Latin hypercube sampling method, where M should be greater than N.

[0027] In some embodiments, the grasping control strategy is trained using a proximal policy optimization algorithm or a soft actor-critic algorithm.

[0028] The reward function for training the grasping control strategy is set as a weighted combination of three factors: robotic arm hand tracking accuracy, robotic arm grasping object stability, and lifting efficiency.

[0029] In some embodiments, step S300 includes:

[0030] Step S310: Perform a posteriori update on the Gaussian process GP using the observation dataset D, taking each stiffness parameter in the stiffness parameter space as a candidate point x, and calculating the predicted mean of the Gaussian distribution corresponding to each candidate point x. and the predicted standard deviation Based on the predicted mean and predicted standard deviation, the stiffness parameter space is sampled to regenerate M sets of stiffness parameters. , and respectively serve as a set of candidate stiffness parameters;

[0031] Step S320: Evaluate the reward value of all candidate stiffness parameter sets based on the expected improvement function EI, and sort the reward values ​​from largest to smallest;

[0032] Step S330: Sequentially determine whether each candidate stiffness parameter set satisfies the energy constraint condition, select the maximum reward value that satisfies the energy constraint condition, and use the corresponding candidate stiffness parameter set as the stiffness parameter set to be optimized in this round, denoted as... ,Will Add to the stiffness parameter pool P;

[0033] Step S340, based on filtering The grasping control strategy is retrained using reinforcement learning methods, the grasping task is executed, and a new set of stiffness parameters and reward values ​​are recorded. ; to use new observation data ( , Add the observation dataset D.

[0034] In some embodiments, the desired improvement function EI is defined as:

[0035]

[0036] in, The maximum reward value in the observed dataset D; For expectation operators.

[0037] In some embodiments, the iteration termination condition is that the reward value calculated for the grasping control strategy converges or the maximum number of iterations is reached.

[0038] A second aspect of the present invention provides an optimization device for the grasping ability of a musculoskeletal robotic arm, comprising:

[0039] The parameter setting module is configured to represent the shape of the musculoskeletal robotic arm using the stiffness parameter space of the musculoskeletal robotic arm, with the total energy consumption of all muscles during the entire grasping process not exceeding a preset threshold as an energy constraint condition.

[0040] The initial training module is configured to generate several initial stiffness parameter sets based on the stiffness parameter space, model the reward function of the stiffness parameter sets as a Gaussian process, perform preliminary training on the grasping control strategy of the musculoskeletal robotic arm based on each initial stiffness parameter set and under the energy constraint condition, obtain the initial reward value of each stiffness parameter set, store each stiffness parameter set and its initial reward value as observation data in the observation dataset, and store the stiffness parameter sets in the stiffness parameter pool.

[0041] The iterative training module is configured to perform several rounds of iterative training until the iteration termination condition is met, thereby obtaining the optimal set of stiffness parameters and the optimal grasping control strategy. During each round of iterative training, a set of stiffness parameters to be optimized is selected from the stiffness parameter pool through the expected improvement function. Based on the set of stiffness parameters to be optimized, the grasping control strategy is trained under energy constraints, and the Gaussian process, the observation dataset, and the stiffness parameter pool are updated.

[0042] A third aspect of the present invention provides a computer-readable storage medium storing computer instructions for causing the computer to perform an optimized method according to any embodiment of the first aspect of the present invention.

[0043] This invention has the following characteristics and beneficial effects:

[0044] Dynamic parameter optimization: The Bayesian optimization (BO) method is used to actively adjust the muscle stiffness parameters. The reward function is modeled through a Gaussian process to find the optimal muscle stiffness parameters with the fewest iterations.

[0045] Co-optimization of morphology and control strategy: In each grasping process, this invention collects performance indicators such as the displacement of the robotic arm hand, object stability, and lifting efficiency in real time, and feeds them back to the Bayesian optimization module as reward values. The Bayesian optimization module updates the muscle stiffness parameters used to characterize the morphology of the robotic arm based on the input reward values, and the reinforcement learning algorithm (such as PPO or SAC) trains the control strategy based on the updated muscle stiffness parameters. This process is repeated cyclically to build a closed-loop feedback mechanism for the co-iterative optimization of muscle stiffness parameters and the training of the strategy.

[0046] High generalization capability: The Bayesian optimization stage of this invention generates and evaluates candidate stiffness parameters in parallel for different weights (≤0.1 kg to ≥1.0 kg) and different geometries (such as cup-shaped, banana-shaped and irregular shapes); the candidate parameters are jointly optimized by a cross-task collaborative objective function, and the obtained parameter set can achieve the expected performance in the full-coverage weight-shape combination space; therefore, without additional training, the parameter set can be directly applied to various grasping scenarios.

[0047] Computational efficiency: Compared with genetic algorithms, Bayesian optimization reduces the number of iterations by 20%. Bayesian optimization can efficiently optimize stiffness parameters with the fewest function evaluations, ensuring optimal performance and saving computational resources.

[0048] Task performance: Success rate of grabbing heavy objects increased by 10%, cumulative rewards increased by 15%;

[0049] Robustness: The mission success rate remains ≥80% under external disturbances or dynamic targets. Attached Figure Description

[0050] Figure 1 This is a flowchart of a method for optimizing the grasping ability of a musculoskeletal robotic arm according to a first aspect embodiment of the present invention;

[0051] Figure 2 This is a schematic diagram of the structure of an electronic device provided in a third aspect embodiment of the present invention. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of this application clearer, the application will be described in further detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application.

[0053] Conversely, this application covers any alternatives, modifications, equivalent methods, and schemes made within the spirit and scope of this application as defined by the claims. Furthermore, to provide the public with a better understanding of this application, certain specific details are described in detail below. However, this application can be fully understood by those skilled in the art even without these detailed descriptions.

[0054] See Figure 1 The first aspect of this invention provides a method for optimizing the grasping ability of a musculoskeletal robotic arm, comprising the following steps:

[0055] Step S100: The shape of the musculoskeletal robotic arm is represented by the stiffness parameter space of the musculoskeletal robotic arm, and the stiffness parameter space is defined as a multi-dimensional vector. The total energy consumption of all muscles in the entire grasping process does not exceed a preset threshold as an energy constraint condition.

[0056] Step S200: Generate several initial stiffness parameter sets based on the stiffness parameter space set in step S100, construct an observation dataset and initialize it as an empty set, model the reward function of the stiffness parameter set as a Gaussian process (GP), perform preliminary training on the grasping control strategy of the musculoskeletal robotic arm based on each initial stiffness parameter set under energy constraints, obtain the initial reward value of each stiffness parameter set, store each stiffness parameter set and its initial reward value as observation data in the observation dataset, and store the stiffness parameter set in the stiffness parameter pool.

[0057] Step S300: Select a set of stiffness parameters to be optimized from the stiffness parameter pool using the expected improvement (EI) function, train the grasping control strategy under energy constraints based on the set of stiffness parameters to be optimized, and update the Gaussian process, the observation dataset and the stiffness parameter pool.

[0058] Step S400: Repeat step S300 continuously, that is, alternately perform morphological optimization and control strategy training to iteratively update the GP model, observation dataset and stiffness parameter pool until the iteration termination condition is reached, to obtain the optimal stiffness parameter set and the optimal grasping control strategy, and deploy them on the musculoskeletal robotic arm.

[0059] In some embodiments, step S100 specifically includes:

[0060] First, based on the structure of the musculoskeletal robotic arm, determine the total number of muscles N it contains, and then represent the stiffness parameters of all muscles within the robotic arm as an N-dimensional stiffness parameter space. ,in, This represents the stiffness parameter of the i-th muscle. The value range is from 0.1 N / mm to 10 N / mm; N is a positive integer, and the general principle is to take a value between 10 and 50.

[0061] Then, to ensure that energy consumption is controlled during the grasping process and to prevent excessive energy use, the following energy constraints are defined:

[0062]

[0063] In the formula, This represents the total energy consumption of all muscles in the robotic arm during the entire grasping process; This represents the equivalent damping force of the i-th muscle during the grasping process; Indicates the duration of the grabbing action; This indicates the maximum energy consumption threshold that is set.

[0064] It is understood that step S100 of this embodiment of the invention uniformly represents the morphology of the musculoskeletal robotic arm using a stiffness parameter space, and introduces energy constraints on this basis, ensuring that the search space of candidate parameters in the subsequent optimization process is both physically reasonable and avoids excessive energy consumption. This step provides a structured and constrained input basis for subsequent Gaussian process modeling and parameter optimization. Compared with the prior art's method of directly using fixed parameters or random search under unconstrained conditions, this embodiment of the invention, through parameter definition and energy constraint design in step S100, has the following advantages:

[0065] 1. Improved efficiency: Reduced the exploration range of invalid or unreasonable parameter combinations, thereby reducing computational costs;

[0066] 2. Enhance controllability: Ensure that the optimization results meet energy consumption constraints and avoid generating parameter configurations that are practically infeasible;

[0067] 3. Improve generalization: By defining the parameter space within a reasonable range, the optimization results can be better adapted to grasping tasks of objects with different weights and shapes.

[0068] In some embodiments, step S200 specifically includes:

[0069] Step S210: Perform layered sampling on the stiffness parameter space K to generate M initial stiffness parameter sets. Each initial stiffness parameter set is an N-dimensional vector, where M should be greater than N. Furthermore, the value of M should cover the high-dimensional parameter space to avoid the locality of pure random sampling. Preferably, M is at least 10 times N.

[0070] In one specific embodiment of the present invention, the hierarchical sampling method adopts the Latin hypercube sampling (LHS) method, which ensures the global distribution of the initial samples by hierarchically covering the high-dimensional parameter space.

[0071] Step S220: Construct the observation dataset D and initialize D as an empty set, which will serve as the input for subsequent steps; based on the settings in step S100... The range of values ​​and energy constraints of the stiffness parameter set, and the reward function. Modeled as a Gaussian process (GP) and using a zero-mean function. and kernel function Initialize the Gaussian process GP, i.e., satisfy: The kernel function for the Gaussian process GP uses the radial basis function RBF, which is defined as:

[0072]

[0073] in, They represent the first A set of stiffness parameters; In this embodiment, the variance of the signal is represented. =1.0; This represents a free parameter, in this embodiment =1.0, noise level is , For the Kronecker delta function, when the set of stiffness parameters... With stiffness parameter set When they are the same, take When the set of stiffness parameters With stiffness parameter set When they are different, take ;

[0074] Step S230: The initial stiffness parameter sets generated in step S210 are... ( The inputs are respectively fed into a reinforcement learning-based simulation platform to perform initial training of the control strategy, execute the grasping task, and record the corresponding initial reward value after the initial training is completed. ;

[0075] Furthermore, the training of the grasping control strategy employs either the Proximal Policy Optimization (PPO) algorithm or the Soft Actor-Critic (SAC) algorithm. Compared to existing methods such as Deep Deterministic Policy Gradient (DDPG) and traditional Actor-Critic, the PPO algorithm can accelerate convergence while ensuring update stability, while the SAC algorithm enhances the exploratory nature and robustness of the proximal policy optimization strategy through maximum entropy optimization, making it more suitable for handling complex grasping tasks in high-dimensional continuous action spaces. The reward function for training the grasping control strategy is set as a weighted combination of three factors: robotic arm hand tracking accuracy, robotic arm grasping object stability, and lifting efficiency.

[0076] Step S240: Set stiffness parameters and the corresponding reward value Constitute a pair of observation data Each pair of observation data is sequentially written into the observation dataset D, and the M stiffness parameter sets { are simultaneously written to the dataset D. Store the data into the stiffness parameter pool P; this completes the construction of the initial observation dataset D and the stiffness parameter pool P.

[0077] Understandably, by generating an initial set of stiffness parameters, constructing an observation dataset, and modeling the reward function as a Gaussian process, initial training samples and a statistical model foundation are provided for subsequent iterative optimization. This step ensures that the optimization process has global distribution and scalability from the outset, enabling subsequent Bayesian optimization to achieve better convergence results with a smaller sample size. Compared to existing technologies, step S200 has the following advantages:

[0078] 1. Improved initial sampling method: The Latin hypercube sampling (LHS) method is used to cover the high-dimensional parameter space in a hierarchical manner, which avoids the locality problem of traditional random sampling and improves the sample diversity and representativeness;

[0079] 2. Reduce computational costs: By incorporating energy constraints during the parameter initialization phase, physically infeasible parameter combinations are eliminated, reducing redundant overhead in subsequent simulation training.

[0080] 3. Improve modeling accuracy: By modeling the reward function as a Gaussian process and introducing the observation dataset in the initial stage, it is possible to provide an uncertainty measure under limited data conditions, laying the foundation for the effective acquisition of the expected improvement (EI) function;

[0081] 4. Enhanced convergence efficiency: Compared with existing evolutionary algorithms or heuristic methods, the embodiments of the present invention achieve convergence in fewer iterations through GP modeling and initial sample distribution optimization, thereby improving the overall training efficiency.

[0082] In some embodiments, step S300 specifically includes:

[0083] Step S310: Perform a posteriori update on the Gaussian process GP using the observation dataset D. Take each set of stiffness parameters in the stiffness parameter pool P as a candidate point x, perform a Gaussian distribution on each candidate point, and calculate the predicted mean of the Gaussian distribution. and the predicted standard deviation Based on the predicted mean and predicted standard deviation, the stiffness parameter space is sampled to regenerate M sets of stiffness parameters. , and respectively serve as a set of candidate stiffness parameters;

[0084] Step S320: Evaluate the reward value of all candidate points in the stiffness parameter pool P based on the expected improvement function EI, and sort the reward values ​​of all candidate points from largest to smallest;

[0085] Furthermore, in this embodiment, the desired improvement function EI is defined as:

[0086]

[0087] in, This is the historical best reward value, i.e., the maximum reward value in the observation dataset D; For expectation operators;

[0088] Step S330: Sequentially (from largest to smallest), determine whether the set of stiffness parameters corresponding to the reward value of each candidate point satisfies the energy constraint conditions set in step S100. Select the largest reward value that satisfies the energy constraint conditions, and use its corresponding set of stiffness parameters as the set of stiffness parameters to be optimized in this round, denoted as . This allows for adjustments to the shape of the robotic arm;

[0089] Step S340, based on filtering The grasping control strategy is trained using reinforcement learning methods, the grasping task is executed, and new reward values ​​are recorded. ; to use new observation data ( , Add the observation dataset D.

[0090] Understandably, by performing a posteriori updates on the Gaussian process (GP) model and using the expected improvement function (EI) to sort and filter candidate stiffness parameters, the set of stiffness parameters with the highest improvement potential is prioritized as the optimization target while ensuring energy constraints. This step achieves "purposeful exploration" in the parameter space, avoiding blind search and enabling the optimization process to converge quickly to the optimal solution. Compared to existing technologies, step S300 has the following advantages:

[0091] 1. Introducing probabilistic modeling: Through the posterior update of GP, the predicted mean and uncertainty of candidate points can be given simultaneously, thus taking into account both "exploration" (high uncertainty region) and "utilization" (high reward region) in the optimization process.

[0092] 2. Improved sampling strategy: The EI function is used to select candidate parameters, ensuring that each iteration can achieve performance improvement based on the historical best value, thus avoiding invalid iterations in traditional genetic algorithms or random sampling;

[0093] 3. Save computational resources: Introducing energy constraints when selecting parameters and filtering out infeasible solutions in advance reduces unnecessary simulation training and lowers computational costs;

[0094] 4. Improved convergence speed: By sorting and filtering mechanisms, the most valuable set of stiffness parameters is optimized first, enabling the optimization process to converge with fewer iterations, which is more efficient and practical than existing methods.

[0095] In some embodiments, step S400 specifically includes:

[0096] Step S410: Repeat step S300;

[0097] Step S420: Determine the iteration termination condition: If the reward value of the stiffness parameter set converges (e.g., the reward fluctuation is less than the threshold for T consecutive iterations) or the number of iterations reaches the maximum value, then terminate the iteration, obtain the optimal stiffness parameter set and grasping control strategy, and deploy them on the musculoskeletal robotic arm to complete the optimization of the musculoskeletal robotic arm's grasping ability; otherwise, use the historical optimal reward value. Updated to max( , Then return to step S300 to continue execution.

[0098] 1. System Architecture

[0099] The system of this invention includes the following modules:

[0100] Input module: Receives the task target (object weight, shape) and the initial range of muscle parameters;

[0101] Bayesian optimization module: Based on Gaussian process modeling, iteratively selects the optimal stiffness parameters;

[0102] Control strategy training module: Uses reinforcement learning algorithms to generate control commands that match the current parameters;

[0103] Evaluation module: Calculates metrics such as success rate and cumulative rewards using a simulation environment, and feeds these metrics back to the optimization module;

[0104] Output module: Deploys optimal parameters and control strategies to physical actuators (such as bionic robotic arms).

[0105] The effectiveness of the optimization method provided in the embodiments of the present invention is verified below. Taking the MyoSuite simulation platform as an example, the implementation steps are as follows:

[0106] 1. Task configuration: Define the target to be grasped as a cup weighing 1.0 kg, and set the initial muscle stiffness range to 0.5-8 N / mm;

[0107] 2. Bayesian optimization initialization: A Gaussian process model is constructed using the RBF kernel, and 10 sets of initial parameters are generated through Latin hypercube sampling;

[0108] 3. Iterative optimization: Select one set of candidate parameters in each round, train the PPO policy until convergence (approximately 5000 steps), and evaluate the success rate;

[0109] 4. Result Verification: In 20 tests, the optimized system achieved an 85% success rate in grasping heavy objects and a 95% success rate in grasping light objects.

[0110] 5. Parameter migration: The optimized stiffness parameters were directly applied to the banana grabbing task, and the success rate remained at 82%.

[0111] The advantages and effects of the embodiments of the present invention are reflected in:

[0112] 1. Computational efficiency: Compared to genetic algorithms, Bayesian optimization reduces redundant evaluations through active sampling, reducing the number of iterations by 20%;

[0113] 2. Task Performance: Success rate of grabbing heavy objects increased by 10%, and cumulative rewards increased by 15%;

[0114] 3. Generalization ability: Optimized parameters can be transferred to objects of different shapes to adapt to dynamic task requirements;

[0115] 4. Robustness: The success rate remains ≥80% under external disturbances (such as random external forces).

[0116] A second aspect of the present invention provides an optimization device for the grasping ability of a musculoskeletal robotic arm, comprising:

[0117] The parameter setting module is configured to represent the shape of the musculoskeletal robotic arm using the stiffness parameter space of the musculoskeletal robotic arm, with the total energy consumption of all muscles during the entire grasping process not exceeding a preset threshold as an energy constraint condition.

[0118] The initial training module is configured to generate several initial stiffness parameter sets based on the stiffness parameter space, model the reward function of the stiffness parameter sets as a Gaussian process, perform preliminary training on the grasping control strategy of the musculoskeletal robotic arm based on each initial stiffness parameter set and under the energy constraint condition, obtain the initial reward value of each stiffness parameter set, store each stiffness parameter set and its initial reward value as observation data in the observation dataset, and store the stiffness parameter sets in the stiffness parameter pool.

[0119] The iterative training module is configured to perform several rounds of iterative training until the iteration termination condition is met, thereby obtaining the optimal set of stiffness parameters and the optimal grasping control strategy. During each round of iterative training, a set of stiffness parameters to be optimized is selected from the stiffness parameter pool through the expected improvement function. Based on the set of stiffness parameters to be optimized, the grasping control strategy is trained under energy constraints, and the Gaussian process, the observation dataset, and the stiffness parameter pool are updated.

[0120] It should be noted that the aforementioned explanation of the method for optimizing the grasping ability of the musculoskeletal robotic arm also applies to the device for optimizing the grasping ability of the musculoskeletal robotic arm in this embodiment, and will not be repeated here.

[0121] To implement the above embodiments, this invention also proposes a computer-readable storage medium storing a computer program that is executed by a processor to perform the optimization method for the grasping ability of the musculoskeletal robotic arm described above.

[0122] The following is for reference. Figure 2 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present invention. It should be noted that the electronic device in the embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs, desktop computers, and servers. Figure 2 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0123] like Figure 2As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 102 or a program loaded from a storage device 108 into a random access memory (RAM) 103. The RAM 103 also stores various programs and data required for the operation of the electronic device. The processing unit 101, ROM 102, and RAM 103 are interconnected via a bus 104. An input / output (I / O) interface 105 is also connected to the bus 104.

[0124] Typically, the following devices can be connected to I / O interface 105: input devices 106 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 107 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 108 including, for example, magnetic tapes, hard disks, etc.; and communication devices 109. Communication device 109 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 2 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0125] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, this embodiment includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via communication device 109, or installed from storage device 108, or installed from ROM 102. When the computer program is executed by processing device 101, it performs the functions defined above in the methods of embodiments of this disclosure.

[0126] It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0127] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0128] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the aforementioned method for optimizing the grasping ability of the musculoskeletal robotic arm.

[0129] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and Python, as well as conventional procedural programming languages ​​such as the "C-" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0130] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0131] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0132] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.

[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0134] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0135] Those skilled in the art will understand that implementing all or part of the steps of the methods in the above embodiments can be accomplished by instructing related hardware through a program. The developed program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0136] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0137] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for optimizing the grasping ability of a musculoskeletal robotic arm, characterized in that, include: Step S100: The shape of the musculoskeletal robotic arm is represented by the stiffness parameter space of the musculoskeletal robotic arm, with the total energy consumption of all muscles in the entire grasping process not exceeding a preset threshold as the energy constraint condition. Step S200: Generate several initial stiffness parameter sets based on the stiffness parameter space, model the reward function of the stiffness parameter sets as a Gaussian process, perform preliminary training on the grasping control strategy of the musculoskeletal robotic arm based on each initial stiffness parameter set and under the energy constraint condition, obtain the initial reward value of each stiffness parameter set, store each stiffness parameter set and its initial reward value as a pair of observation data in the observation dataset, and store the stiffness parameter sets in the stiffness parameter pool. Step S300: Select a set of stiffness parameters to be optimized from the stiffness parameter pool through the expected improvement function, train the grasping control strategy based on the set of stiffness parameters to be optimized under energy constraints, and update the Gaussian process, the observation dataset and the stiffness parameter pool. Step S400: Repeat step S300 several times until the iteration termination condition is met, and obtain the optimal stiffness parameter set and the optimal grasping control strategy.

2. The optimization method according to claim 1, characterized in that, In step S100, the stiffness parameter space is: , This represents the stiffness parameter of the i-th muscle in the robotic arm, set as follows: The range of values ​​for ; The expression for the energy constraint is: In the formula, This represents the total energy consumption of all muscles in the robotic arm during the entire grasping process; N represents the number of muscles contained in the robotic arm. This represents the equivalent damping force of the i-th muscle during the grasping process; Indicates the duration of the grabbing action; This indicates the maximum energy consumption threshold that is set.

3. The optimization method according to claim 1, characterized in that, Step S200 includes: Step S210: Perform layered sampling on the stiffness parameter space to generate M initial stiffness parameter sets. The dimensions of each stiffness parameter set are the same as the dimensions of the stiffness parameter space; Step S220: The reward function of the stiffness parameter set based on the parameters set in step S100. Modeled as a Gaussian process, and using a zero-mean function. and kernel function Initialize the Gaussian process GP, i.e., satisfy: The kernel function for the Gaussian process GP uses radial basis functions, which are defined as follows: in, They represent the first A set of stiffness parameters; Indicates the signal variance; Describe a free parameter, Noise level; For the Kroneck delta function, when the set of stiffness parameters With stiffness parameter set When they are the same, take When the set of stiffness parameters With stiffness parameter set When they are different, take ; Step S230: Based on the set of stiffness parameters The grasping control strategy was initially trained using reinforcement learning methods, and grasping tasks were performed. After the initial training was completed, the set of stiffness parameters was recorded. Reward value ; Step S240: Transfer M pairs of observation data Store the observation dataset D, and simultaneously transfer the M stiffness parameter sets { Store it in the stiffness parameter pool P.

4. The optimization method according to claim 3, characterized in that, The stratified sampling uses the Latin hypercube sampling method, and M should be greater than N.

5. The optimization method according to claim 1, characterized in that, The grasping control strategy is trained using either a proximal policy optimization algorithm or a soft actor-critic algorithm. The reward function for training the grasping control strategy is set as a weighted combination of three factors: robotic arm hand tracking accuracy, robotic arm grasping object stability, and lifting efficiency.

6. The optimization method according to claim 1, characterized in that, Step S300 includes: Step S310: Perform a posteriori update on the Gaussian process GP using the observation dataset D, taking each stiffness parameter in the stiffness parameter space as a candidate point x, and calculating the predicted mean of the Gaussian distribution corresponding to each candidate point x. and the predicted standard deviation Based on the predicted mean and predicted standard deviation, the stiffness parameter space is sampled to regenerate M sets of stiffness parameters. , and respectively serve as a set of candidate stiffness parameters; Step S320: Evaluate the reward value of all candidate stiffness parameter sets based on the expected improvement function EI, and sort the reward values ​​from largest to smallest; Step S330: Sequentially determine whether each candidate stiffness parameter set satisfies the energy constraint condition, select the maximum reward value that satisfies the energy constraint condition, and use the corresponding candidate stiffness parameter set as the stiffness parameter set to be optimized in this round, denoted as... ,Will Add to the stiffness parameter pool P; Step S340, based on filtering The grasping control strategy is retrained using reinforcement learning methods, the grasping task is executed, and a new set of stiffness parameters and reward values ​​are recorded. ; to use new observation data ( , Add the observation dataset D.

7. The optimization method according to claim 6, characterized in that, The definition of the expected improvement function EI is: in, The maximum reward value in the observed dataset D; For expectation operators.

8. The optimization method according to claim 1, characterized in that, The iteration termination condition is that the reward value calculated for the grasping control strategy converges or the maximum number of iterations is reached.

9. A device for optimizing the grasping ability of a musculoskeletal robotic arm, characterized in that, include: The parameter setting module is configured to represent the shape of the musculoskeletal robotic arm using the stiffness parameter space of the musculoskeletal robotic arm, with the total energy consumption of all muscles during the entire grasping process not exceeding a preset threshold as an energy constraint condition. The initial training module is configured to generate several initial stiffness parameter sets based on the stiffness parameter space, model the reward function of the stiffness parameter sets as a Gaussian process, perform preliminary training on the grasping control strategy of the musculoskeletal robotic arm based on each initial stiffness parameter set and under the energy constraint condition, obtain the initial reward value of each stiffness parameter set, store each stiffness parameter set and its initial reward value as observation data in the observation dataset, and store the stiffness parameter sets in the stiffness parameter pool. The iterative training module is configured to perform several rounds of iterative training until the iteration termination condition is met, thereby obtaining the optimal set of stiffness parameters and the optimal grasping control strategy. During each round of iterative training, a set of stiffness parameters to be optimized is selected from the stiffness parameter pool through the expected improvement function. Based on the set of stiffness parameters to be optimized, the grasping control strategy is trained under energy constraints, and the Gaussian process, the observation dataset, and the stiffness parameter pool are updated.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the optimized method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Impedance controller parameter adjusting method, device, equipment and medium

    CN121832305A