Method, device and equipment for in-hand operation of robot and storage medium
By using deep reinforcement learning and tactile sensor information in the physics simulator, combined with robot mechanical characteristics, training and transfer of in-hand operation strategies, the challenges of operating slender objects are solved, and efficient and safe transfer of dexterous in-hand operation is achieved.
Patent Information
- Application Number
- CN202410181719.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-18
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art is difficult to achieve agile continuous in-hand operation of elongated cylindrical objects, especially adjusting the position and posture of the object while maintaining grip, and the existing methods are dangerous and costly to train in real environments.
In-hand operation strategies are trained in-hand using deep reinforcement learning in physics simulators, using contact position information of tactile sensors, combining the robot's mechanical backlash and self-locking characteristics, calibrating the knuckle model, and transferring the strategy to the real robot through domain randomization, reducing the time and danger of training data collection.
The dexterous continuous in-hand operation of slender cylindrical objects is achieved. The strategy performs well in simulation and real world, has good generalization ability, and avoids the time and manpower consumption of real training.
Smart Images

Figure CN120503185A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and robotics, and more particularly, to a method, apparatus, device, and storage medium for in-hand operation of a robot. Background Art
[0002] In-hand manipulation refers to the use of a single robotic hand to change the relative position and posture of an object within the hand by moving its fingers, palm, and other parts. Continuous in-hand manipulation requires the robotic hand to be able to manipulate an object deftly and continuously while grasping it, without having to release and re-grasp it. This capability enables the robot to adjust the position, orientation, and shape of an object while maintaining a stable grasp, similar to how the human hand can perform complex manipulation tasks.
[0003] Continuous in-hand manipulation involves a combination of fine control, tactile perception, and adaptive grasping strategies, which enables a robot to reposition an object, rotate it, or change its orientation without losing contact with the object. Developing effective algorithms and control strategies for continuous in-hand manipulation is an important research area in the field of robotic manipulation.
[0004] Therefore, there is a need for a method for in-hand manipulation of a robot that enables dexterous continuous in-hand manipulation. Summary of the Invention
[0005] To this end, this paper uses deep reinforcement learning for policy training in simulation, learns policies by extracting contact position information from signals from tactile sensors through robot interaction with the environment, and transfers the learned policies to real-world experiments using coordinated model calibration and domain randomization, thereby achieving tactile-based dexterous in-hand manipulation of slender cylindrical objects.
[0006] Embodiments of the present disclosure provide a method, apparatus, device, and computer-readable storage medium for in-hand manipulation of a robot.
[0007] An embodiment of the present disclosure provides a method for in-hand manipulation of a robot, wherein the robot has at least one finger, wherein each finger has at least one knuckle, and the fingertip of each finger is configured with a tactile sensor, the method comprising: simulating and modeling the robot in a physical simulator to obtain a simulated robot, wherein the simulated robot has the same features as the robot, wherein the features include tooth clearance and self-locking features of the knuckles; calibrating a joint model of each knuckle of each finger of the simulated robot; generating a training data set based on the calibrated joint model of each knuckle of the simulated robot, wherein the training data set includes a plurality of candidate knuckles; Selecting an initial state, wherein each candidate initial state includes finger joint positions of the simulated robot and a posture of a target object; initializing a training environment based on sampling from the training data set, and training an in-hand operation strategy through interaction between the simulated robot and the training environment, wherein the in-hand operation strategy takes observations of the simulated robot at a current time step as input and takes actions of the simulated robot at a next time step as output, and the observations of the simulated robot at the current time step include the contact center position of each finger of the simulated robot with the target object at the current time step; and applying the in-hand operation strategy to the in-hand operation of the robot after training is completed.
[0008] An embodiment of the present disclosure provides a device for in-hand operation of a robot, wherein the robot has at least one finger, wherein each finger has at least one knuckle, and the fingertip of each finger is configured with a tactile sensor, and the device comprises: a simulation modeling module, configured to simulate and model the robot in a physical simulator to obtain a simulated robot, wherein the simulated robot has the same features as the robot, wherein the features include tooth clearance and self-locking features of the knuckles; a model calibration module, configured to calibrate the joint model of each knuckle of each finger of the simulated robot; a data preparation module, configured to generate a training data set based on the calibrated joint model of each knuckle of the simulated robot, wherein the training data set The method comprises a plurality of candidate initial states, wherein each candidate initial state comprises the finger joint positions of the simulated robot and the posture of the target object; a strategy training module, configured to initialize the training environment based on sampling from the training data set, and train the in-hand operation strategy through the interaction between the simulated robot and the training environment, wherein the in-hand operation strategy takes the observation of the current time step of the simulated robot as input and the action of the simulated robot in the next time step as output, and the observation of the current time step of the simulated robot comprises the contact center position of each finger of the simulated robot with the target object in the current time step; and a strategy application module, configured to apply the in-hand operation strategy to the in-hand operation of the robot after the training is completed.
[0009] An embodiment of the present disclosure provides a device for in-hand operation of a robot, comprising: one or more processors; and one or more memories, wherein a computer executable program is stored in the one or more memories, and when the computer executable program is executed by the processor, the method for in-hand operation of the robot as described above is performed.
[0010] An embodiment of the present disclosure provides a computer-readable storage medium having computer-executable instructions stored thereon. When the instructions are executed by a processor, the instructions are used to implement the method for in-hand operation of a robot as described above.
[0011] Embodiments of the present disclosure provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a method for in-hand operation of a robot according to an embodiment of the present disclosure.
[0012] The method provided by the embodiments of the present disclosure trains an in-hand manipulation policy for a three-fingered robotic hand equipped with tactile sensors in a physical simulator and applies the learned in-hand manipulation policy to a real robot. A deep reinforcement learning method is used to train the policy using the contact center position between the robot hand and the target object as a partial observation. The finger joint models are calibrated by considering the robot's mechanical backlash and self-locking characteristics, and the initial training states are randomized to ensure successful transfer of the trained policy to the real robot, thereby enabling the robot to perform dexterous continuous in-hand manipulation of the target object. The method of the embodiments of the present disclosure simulates the robot using a physical simulator and trains the policy for direct transfer to the real robot, avoiding the time and manpower required to collect training data for the real robot and the potential risks involved. Deep reinforcement learning enables dexterous in-hand manipulation of the target object based on tactile perception, wherein a more optimized policy is trained by using the estimated contact center position as an observation. Furthermore, calibrating the finger joint model and randomizing the initial states mitigate the gap between simulation and real-world experiments, allowing the learned policy to perform well in both simulation and real-world experiments and generalize well to new target objects and tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some exemplary embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0014] Figure 1 1 is a schematic diagram showing the structure of a hand of a robot according to an embodiment of the present disclosure;
[0015] Figure 2 is a flow chart illustrating a method for in-hand operation of a robot according to an embodiment of the present disclosure;
[0016] Figure 3 is a schematic diagram illustrating a simulated robot and a target object according to an embodiment of the present disclosure;
[0017] Figure 4 is a schematic diagram illustrating calibration of proportional gain and differential gain of a finger joint according to an embodiment of the present disclosure;
[0018] Figure 5 is a schematic diagram illustrating modeling of a backlash joint of a robot according to an embodiment of the present disclosure;
[0019] Figure 6 is a schematic diagram illustrating example candidate initial states according to an embodiment of the present disclosure;
[0020] Figure 7 is a result diagram showing performance verification of strategies for training different predetermined finger motion tasks according to an embodiment of the present disclosure;
[0021] Figure 8 is a graph illustrating changes in the positions of the finger joints of the fingers of the simulated robot during the execution of a predetermined finger motion task according to an embodiment of the present disclosure;
[0022] Figure 9 is a graph showing results of errors in performing different predetermined finger motion tasks according to an embodiment of the present disclosure;
[0023] Figure 10 is a diagram showing the results of training a policy using different observations of a target object according to an embodiment of the present disclosure;
[0024] Figure 11 is a result diagram showing contact position trajectories of one finger of a robot in different directions and activation of a tactile sensor during execution of a predetermined finger motion task according to an embodiment of the present disclosure;
[0025] Figure 12 is a schematic diagram illustrating an apparatus for in-hand operation of a robot according to an embodiment of the present disclosure;
[0026] Figure 13 A schematic diagram illustrating an apparatus for in-hand manipulation of a robot according to an embodiment of the present disclosure; and
[0027] Figure 14 A schematic diagram illustrating the architecture of an exemplary computing device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present disclosure more apparent, the exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0029] In this specification and the accompanying drawings, substantially the same or similar steps and elements are denoted by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, the terms "first", "second", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance or ranking.
[0030] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present disclosure pertains. The terms used herein are for the purpose of describing embodiments of the present invention only and are not intended to limit the present invention.
[0032] To facilitate description of the present disclosure, concepts related to the present disclosure are introduced below.
[0033] The disclosed method for in-hand manipulation of a robot can be based on artificial intelligence (AI). Artificial intelligence (AI) refers to theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a manner similar to human intelligence. For example, an AI-based method for in-hand manipulation of a robot can control the movement of the knuckles of a robot hand in a manner similar to how humans change the posture of an object by simply moving their fingers without excessive wrist or arm movement. This enables the movement and posture of a target object to be manipulated through multiple touches with the fingertips of a dexterous robotic hand equipped with a tactile sensor array. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware- and software-level technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or foundational models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0034] The method for in-hand operation of a robot disclosed in the present invention can be based on reinforcement learning (RL). Reinforcement learning is a branch of machine learning that aims to enable an agent to learn how to make decisions through interaction with the environment so that it can obtain the maximum cumulative reward in the future. In reinforcement learning, the agent observes the feedback of the environment by trying different behaviors and gradually learns the optimal behavior strategy. The core concepts of reinforcement learning may include: (1) Agent: The agent is an entity that performs learning tasks. It influences the environment by observing the state of the environment and choosing behaviors. In the large language model, the model itself can be regarded as an agent; (2) Environment: The environment is the external world in which the agent is located. It provides feedback on the agent's behavior and determines the agent's next state. In the large language model, the process of generating text can be regarded as an interaction process between the agent and the environment; (3) Reward: The agent receives rewards or penalties based on the feedback from the environment. Reward refers to the numerical feedback that an agent receives after performing a certain action in a certain state. The goal is to enable the agent to learn the optimal behavior strategy by maximizing the cumulative reward. (4) Action: Different environments allow different types of actions. In a given environment, the set of valid actions is often called the action space, which includes discrete action spaces and continuous action spaces. (5) Policy: The policy defines the rules for the agent to choose behavior in a specific state. The goal of reinforcement learning is to learn the optimal policy so that the behavior selected by the agent in different states can maximize the cumulative reward.
[0035] Specifically, the method for in-hand operation of a robot disclosed in the present invention can be based on deep reinforcement learning (DRL). Deep reinforcement learning is a machine learning method that combines deep learning and reinforcement learning, wherein deep learning is a machine learning method that uses artificial neural networks for feature learning and representation learning, while reinforcement learning is a method of learning how to make decisions to obtain the maximum cumulative reward through the interaction between an intelligent agent and an environment. Deep reinforcement learning uses deep learning to process high-dimensional perceptual data, such as images, sounds, etc., and combines reinforcement learning methods to enable intelligent agents to make decisions and learn in complex environments. The core of deep reinforcement learning is to use deep neural networks to approximate value functions or policy functions to achieve learning and decision-making for complex environments and tasks. In an embodiment of the present disclosure, deep reinforcement learning can be used to learn in-hand operation strategies for the interaction of a robot with a target object.
[0036] The method for in-hand manipulation of a robot disclosed in the present invention can be based on the Proximal Policy Optimization (PPO) algorithm. The core idea of the PPO algorithm is to improve the stability of training by limiting the amplitude of policy updates to ensure that each update is not too large. The PPO algorithm uses two important concepts at the same time: clipping and surrogate objective. Clipping refers to limiting the ratio between the new policy and the old policy to prevent the update amplitude from being too large; the surrogate objective is an alternative optimization objective used to measure the improvement of the policy when updating.
[0037] The method for in-hand operation of a robot disclosed herein can be based on tactile sensing. Tactile sensing refers to the ability of humans to perceive and understand the shape, texture, temperature and other characteristics of objects through the skin and nervous system. Robots can imitate and apply tactile sensing technology to achieve perception and understanding of objects. The tactile perception of a robot refers to the ability of a robot system to perceive and understand the external environment using sensors and information processing technology. Through tactile perception, robots can grasp, assemble, and operate objects more accurately, and can better adapt to complex and uncertain environments, thereby improving the autonomy and flexibility of the robot. With the continuous development of sensor technology and information processing technology, the tactile perception ability of robots will be further improved, thereby better adapting to various complex tasks and environments.
[0038] In summary, the solutions provided by the embodiments of the present disclosure involve technologies such as artificial intelligence, deep reinforcement learning, and tactile perception. The embodiments of the present disclosure will be further described below in conjunction with the accompanying drawings.
[0039] Dexterous manipulation of objects in-hand is an important skill that humans practice every day. This manipulation involves changing the pose of an object by moving only the fingers without excessive wrist or arm movements. For robots, this locomotion capability is highly desirable, especially when working in confined spaces or using certain tools (e.g., stirring sticks). However, in-hand manipulation remains challenging for robots because it involves maintaining and securing a firm grip while continuously changing the pose of the object, especially when the objects being manipulated are small or slender in shape.
[0040] Many studies have been conducted on in-hand manipulation of robots, and most of them involve reorienting blocky objects (such as cuboids and spheres). For example, model-based trajectory optimization has achieved good performance on object reorientation tasks in both underactuated and fully actuated hands. However, the high-dimensional search space presented by multi-fingered robotic hands makes the optimization problem difficult to solve in real time, and the errors and uncertainties in the hand dynamics and contact models also limit the planning performance in real-world experiments. In addition, unlike the manipulation of blocky objects mentioned above, the manipulation of small or slender objects (e.g., slender cylindrical objects) requires highly coordinated finger motions and precise contact perception. Because the feasible contact area of such objects is small and narrow, in-hand manipulation is more sensitive to errors and uncertainties. In addition, the manipulation of small or slender objects often requires sliding and rotation between the fingers and the object, which makes it difficult to achieve precise control of such objects.
[0041] Therefore, in this application, the in-hand manipulation of such slender objects by a three-fingered robotic hand equipped with tactile sensors will be explored. Figure 1 is a schematic diagram illustrating the structure of a robot hand according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, the robot may have at least one finger, each finger having at least one knuckle, and the tip of each finger may be configured with a tactile sensor. Optionally, the tactile sensor may be covered by a layer of material that provides mechanical compliance.
[0042] like Figure 1 As shown, the robot hand of the present disclosure may be, for example, a three-fingered robot hand, which may have three fingers (e.g., Figure 1 , which can simulate the human thumb, index finger and middle finger respectively) and 8 fully driven joints (e.g., Figure 1 (e.g., J0-J7 in
[15] ). Optionally, each finger can have two joints, and the bases of two of the fingers (e.g., fingers 2 and 3) can be mounted on additional rotational joints (e.g., joints J2 and J5) that allow them to rotate independently around the palm. Unlike many other robotic hand designs that use flat fingertips, the inner surfaces of the fingertips of the robotic hand of the present disclosure can be deformable and bendable to facilitate more dexterous in-hand manipulation, simulating humans.
[0043] As mentioned above, robots can imitate and apply human tactile perception to achieve perception and manipulation of objects. When humans manipulate objects, contact events are captured in fine detail by dense mechanoreceptors embedded in the skin, providing important contact information such as the time, position, and force experienced. In robots, similar information can only be effectively captured by tactile sensors because other sensors such as vision and proprioception may be blocked by fingers or drowned out by background noise. Currently, there are a variety of available tactile sensors, such as vision-based tactile sensors and distributed tactile sensor arrays, which enable the real-time spatial and dynamic relationship between the robot hand and the manipulated object to be derived from the signals collected by the tactile sensors, thereby providing a way to adjust the operational control loop. Optionally, each of the three fingertips of the robotic hand of the present disclosure may be configured with tactile sensors. These tactile sensors may be implanted inside the fingertips and may have an array of piezoresistive sensing elements (tactile pixels (taxel)) distributed on a continuous curved surface. For example, the tactile sensor on each fingertip may be an array consisting of a plurality (e.g., 128) of piezoresistive sensing elements, and each piezoresistive sensing element may return a value proportional to the normal force applied thereto. Optionally, these tactile sensors may be covered by a material layer (e.g., a silicone material layer) that can provide mechanical compliance.
[0044] Tactile information acquired from tactile sensors provides the most direct and indispensable sensory information for dexterous in-hand manipulation tasks. For learning-based in-hand manipulation algorithms, while joint position and torque provide appropriate feedback for adaptive manipulation, tactile information can improve the performance and sample efficiency of policy training. However, leveraging tactile information to achieve human-level dexterous in-hand manipulation remains a challenge.
[0045] Deep reinforcement learning has been successfully implemented in various dexterous in-hand manipulation tasks in robots. Deep reinforcement learning algorithms can be roughly divided into model-based reinforcement learning and model-free reinforcement learning, both of which have achieved remarkable success in in-hand manipulation of robots. Among them, model-based methods apply learned state transition models to guide exploration and policy search. However, for in-hand manipulation tasks with a large number of contacts, it is challenging to accurately model the state transition dynamics, especially when the physical properties of the target object are unknown or when the finger surface is deformable. In the present disclosure, model-free reinforcement learning can be used to achieve a more challenging task: continuous in-hand manipulation of a slender cylindrical object, where the contact area between the robot's fingertips and the target object is narrow and the target object can move on the finger surface.
[0046] Based on this, the present disclosure provides a method for in-hand manipulation of a robot, which uses deep reinforcement learning for policy training in simulation, learns the policy by extracting contact position information from signals from tactile sensors through interaction between the robot and the environment, and transfers the learned policy to real-world experiments using coordinated model calibration and domain randomization, thereby achieving tactile-based dexterous in-hand manipulation of slender cylindrical objects.
[0047] The method provided by the embodiments of the present disclosure trains an in-hand manipulation policy for a three-fingered robotic hand equipped with tactile sensors in a physical simulator and applies the learned in-hand manipulation policy to a real robot. A deep reinforcement learning method is used to train the policy using the contact center position between the robot hand and the target object as a partial observation. The finger joint models are calibrated by considering the robot's mechanical backlash and self-locking characteristics, and the initial training states are randomized to ensure successful transfer of the trained policy to the real robot, thereby enabling the robot to perform dexterous continuous in-hand manipulation of the target object. The method of the embodiments of the present disclosure simulates the robot using a physical simulator and trains the policy for direct transfer to the real robot, avoiding the time and manpower required to collect training data for the real robot and the potential risks involved. Deep reinforcement learning enables dexterous in-hand manipulation of the target object based on tactile perception, wherein a more optimized policy is trained by using the estimated contact center position as an observation. Furthermore, calibrating the finger joint model and randomizing the initial states mitigate the gap between simulation and real-world experiments, allowing the learned policy to perform well in both simulation and real-world experiments and generalize well to new target objects and tasks.
[0048] Figure 2 is a flow chart illustrating a method 200 for in-hand operation of a robot according to an embodiment of the present disclosure.
[0049] In step S201 , the robot may be simulated and modeled in a physical simulator to obtain a simulated robot, wherein the simulated robot has the same features as the robot, including tooth clearance and self-locking features of finger joints.
[0050] Considering that there may be potential dangers in robot training in a real environment, such as improper robot operation may cause personal injury or equipment damage, and the collection of training data for real robots often requires a lot of time and resource investment, such as but not limited to setting up an experimental environment, collecting data, and organizing data, in an embodiment of the present disclosure, a physical simulator can be used for robot training to learn in-hand operation strategies for target objects from simulated experiments and apply them to in-hand operations of real robots, thereby saving costs and time, speeding up the iteration and optimization process of the algorithm, and avoiding these potential risks to ensure the safety of the training process.
[0051] Alternatively, the physical simulator can be any of a variety of existing open-source robotics simulators, such as the MuJoCo (Multi-Joint Dynamics with Contact) physics engine. Of course, the aforementioned physics engine is used in this disclosure only as an example and not as a limitation, and the method of this disclosure can also be used with other physical simulators for policy training.
[0052] Alternatively, the simulation modeling of the robot in the physical simulator can be based on the replication of the structure and features of the real robot to minimize the gap between the simulated and real experiments of the robot and to fully utilize the potential of the tactile sensor.
[0053] Figure 3 is a schematic diagram illustrating a simulated robot and a target object according to an embodiment of the present disclosure.
[0054] like Figure 3 As shown, during the entire operation of the target object by the simulated robot, the wrist of the simulated robot's hand can be fixed in a certain posture in mid-air, and u can represent the unit vector of the main axis of the target object, P1 and P2 represent the positions of the key points (i.e., the two endpoints) of the target object, where the superscripts d and c represent the expected value and the current value (observed value), respectively. In addition, optionally, considering that the human index finger and middle finger usually do not rotate laterally when operating the target object, although the simulated robot has the same joints as the above-mentioned real robot, in the embodiment of the present disclosure, only some of the joints can be used, and the additional rotation joints can be fixed in zero posture. For example, for the above-mentioned Figure 1 The three-finger robotic hand shown can fix joints J2 and J5 at zero pose, while controlling the base joints (e.g., joints J0, J3, and J6) and end joints (e.g., joints J1, J4, and J7) of each of the three fingers to perform in-hand operations on a target object.
[0055] As a simulation of a robotic hand with a tactile sensor as described above, according to an embodiment of the present disclosure, the tactile sensor of each finger of the simulated robot may include a tactile pixel array, and the tactile pixel array may include multiple tactile pixels, each tactile pixel returning a value proportional to the normal force applied thereto.
[0056] Alternatively, the tactile sensor of each finger of the simulated robot can be simulated as an array of tactile pixels located on the inner surface of the finger, such as Figure 3 As shown in the dot array in , each dot can represent a tactile pixel, and the tactile pixel array can be modeled as a surface to fit the inner surface of the fingertip of a real robot.
[0057] During in-hand manipulation of a slender cylindrical object, in the tactile sensor, often only a small portion of the piezoresistive sensing elements are in contact with the target object simultaneously because the piezoresistive sensing elements are thin and the inner surface of the fingertip is curved. Therefore, the raw tactile data collected from the tactile sensor is sparse, and directly using such tactile data for end-to-end policy learning requires a large training dataset or a complex network structure (e.g., a graph convolutional network (GCN)). Therefore, in an embodiment of the present disclosure, the tactile information can be simplified by extracting the contact information between the target object and each fingertip from the high-dimensional tactile information. The processed tactile information is low-dimensional and more intuitive, thereby avoiding the high-dimensional search space presented by the multi-fingered robotic hand, which makes the optimization problem difficult to solve in real time. Considering that the nonlinear mapping from the received contact force to the returned tactile signal is unique for each tactile sensor, modeling of this nonlinear mapping is very important in both simulation experiments and real experiments. At the same time, in in-hand manipulation tasks where the target object is rigid and solid, measuring the normal force of the contact is not as critical as measuring the contact position. Therefore, according to an embodiment of the present disclosure, a Boolean value can be used to represent the contact information between the robot's tactile sensor and the target object. In other words, the complex tactile signal can be binarized into a Boolean value, that is, each piezoresistive sensing element can return a Boolean value to indicate whether the piezoresistive sensing element is in contact with the target object, thereby indicating the contact position between the robot's tactile sensor and the target object based on the return value of these piezoresistive sensing elements.
[0058] As described above, a simulated robot can be modeled in a physical simulator based on the replication of the structure and features of a real robot. Therefore, for the base joints of each finger transmitted by the worm gear mechanism in the real robot (for example, joints J0, J3 and J6), considering that they have backlash and self-locking features, in order to avoid the difference between the simulated experiment and the real experiment due to inaccuracies in joint modeling, differences in sensory feedback and errors in physical parameter estimation, in an embodiment of the present disclosure, the backlash and self-locking features of the finger joints of the real robot can also be replicated in the modeling of the simulated robot to reduce the gap between the simulated experiment and the real experiment, and to improve the generalization ability of the training strategy used for direct transfer from the simulated experiment to the real experiment.
[0059] Based on the simulated robot modeled in the physical simulator described above, it is possible to learn in-hand manipulation policies for, for example, slender cylindrical objects through policy training. However, when transferring the learned policy from simulation experiments to real-world experiments, the differences between the simulation and real-world environments often lead to significant policy performance degradation. Therefore, in the embodiments of the present disclosure, the gap between simulation and real-world experiments can be reduced by modeling the dynamics of the finger joints and calibrating the parameters. Furthermore, by randomizing the dynamics of the simulation environment and the perception space, policy optimality can be traded for policy generalization.
[0060] In step S202 , the joint model of each joint of each finger of the simulated robot may be calibrated.
[0061] According to an embodiment of the present disclosure, for each knuckle of each finger of the simulated robot, calibrating the joint model of the knuckle may include: calibrating the proportional gain and differential gain of the knuckle, wherein the proportional gain and the differential gain may be used to control the torque applied to the knuckle; modeling the gap joint of the robot and calibrating the range of the gap joint, wherein the gap joint corresponds to a joint with a gap; and modeling the self-locking feature of the proximal knuckle of the robot. Optionally, the dynamic features of the knuckle may be calibrated, and these dynamic features may include but are not limited to the kinematic model of the knuckle, the gap and self-locking features of the knuckle, etc.
[0062] According to an embodiment of the present disclosure, calibrating the proportional gain and differential gain of the finger joints may include: controlling the finger joints of the simulated robot and the corresponding finger joints of the robot to move along the same reference trajectory; and optimizing the proportional gain and differential gain of the finger joints by minimizing the error between the actual trajectory of the finger joints of the simulated robot and the corresponding finger joints of the robot.
[0063] Optionally, the robot hand of the present disclosure may be position-controlled. For example, in the simulation experiment of the present disclosure, the torque τ applied to each finger joint may be calculated by proportional-derivative (PD) control as follows:
[0064]
[0065] Among them, K P represents the proportional gain, and K D represents the differential gain, q d represents the target joint position, q represents the measured joint position, represents the derivative of the measured joint position.
[0066] Optionally, in order to calibrate the parameters of the dynamic model of the finger joints, in the simulation experiment, the proportional gain K of each finger joint can be adjusted. P and differential gain K D Calibration is performed. In an embodiment of the present disclosure, parameter calibration can be performed for each finger joint. For each finger joint, the finger joint can be controlled to draw the same reference trajectory in both the simulation experiment and the real experiment. Parameter calibration is performed by comparing the finger joint position trajectories of the simulated robot and the real robot. For example, parameters can be optimized by minimizing the error between the finger joint position trajectories of the simulated robot and the real robot. Figure 4 FIG is a schematic diagram showing the calibration of the proportional gain and the differential gain of the finger joint according to an embodiment of the present disclosure. Figure 4 As shown, by controlling the finger joints to draw the same reference trajectory in simulation experiments and real experiments, the calibrated simulated finger joints have similar dynamic responses to step and continuous signals, and can perform task execution results close to those of real finger joints.
[0067] As an example, the proportional gain K for each knuckle is P and differential gain K D The optimization problem can be expressed as follows:
[0068]
[0069] in, denotes the knuckle positions of the simulated robot and the real robot at time step t, and T denotes the time length of the trajectory. Alternatively, a CMA-ES (Covariance Matrix Adaptation Evolution Strategy) algorithm can be used to solve the above optimization problem (2).
[0070] Of course, the above-mentioned optimization problem and its solution algorithm are only used as examples and not limitations in this disclosure. This disclosure can also establish other optimization problems for the dynamic model parameter calibration problem of the finger joints, and can also use other optimization algorithms to solve the above-mentioned optimization problem or other optimization problems.
[0071] According to an embodiment of the present disclosure, modeling the gap joint of the robot and calibrating the range of the gap joint may include: determining the static joint position of the gap joint of the robot and the extreme joint position under the application of external torque; and determining the range of the gap joint based on the static joint position and the extreme joint position of the gap joint.
[0072] Backlash is the physical gap between mechanical gears. When gear motion reverses and contact is reestablished, backlash causes a certain amount of lost motion due to play or slack. For gear-driven joints, backlash is an unavoidable characteristic caused by the small gaps between the gears and the gearbox. In the presence of backlash, the joint can lose control within the gaps between the gears at certain positions or during reverse motion. Therefore, modeling this backlash characteristic is essential for setting up simulation experiments and thereby learning the correct strategy.
[0073] Therefore, in the embodiments of the present disclosure, the effect of backlash on the relationship between the parent and child links can be considered in the simulation experiment of the physical simulator. Because the backlash between the parent and child links will cause uncontrollable changes in the positions of the finger joints, it is necessary to consider such uncontrolled backlash joints when modeling the joint model in order to more accurately simulate the behavior of the real mechanical system. Specifically, a joint with backlash (i.e., a backlash joint) can be modeled alongside the driven joint, and the actual rotation between the parent link and the child link is the sum of the positions of these two joints.
[0074] Alternatively, the range of the backlash joint can be calibrated experimentally. For example, the target joint (i.e., the backlash joint) can be controlled to remain at a specific position, and then an external torque is applied to rotate the target joint as much as possible to determine the static joint position and extreme joint positions of the backlash joint. Figure 5 : is a schematic diagram showing the modeling of the tooth gap joint of the robot according to an embodiment of the present disclosure. Figure 5 As shown, the backlash joint can be controlled to maintain the target position q d , and record the static joint position q of the tooth gap joint under control a and the extreme joint positions q under external torque c and q b , therefore, [q c –q a ,q b-q a ] can be used as the target position q d Therefore, by repeating the above experiment at different finger joint positions, the average value of these ranges can be used as the tooth gap joint range of the simulated robot.
[0075] In addition to the above-mentioned dynamic characteristics and tooth clearance characteristics, in an embodiment of the present disclosure, the self-locking characteristics of the proximal finger joints of the robot can also be modeled. Since the proximal finger joints use a worm gear mechanism as a motion transmission device, reverse drive is not allowed, which means that the proximal finger joints cannot move against the driving direction. Therefore, according to an embodiment of the present disclosure, modeling the self-locking characteristics of the proximal finger joints of the robot may include: simulating the self-locking characteristics of the proximal finger joints by adjusting the position and speed of the corresponding finger joints of the simulated robot. That is, the above-mentioned self-locking characteristics can be simulated by adjusting the finger joint position and speed of the proximal finger joints in the simulated robot.
[0076] Through the finger joint model calibration as described above, the joint model of the simulated robot's finger joints can be prepared for policy training, and the gap between the simulation experiment and the real experiment can be minimized as much as possible. Next, before conducting policy training, it is necessary to construct a training dataset for policy training.
[0077] In step S203, a training data set may be generated based on a calibrated joint model of each finger joint of the simulated robot, and the training data set may include multiple candidate initial states, wherein each candidate initial state includes a finger joint position of the simulated robot and a posture of a target object.
[0078] Optionally, the method disclosed herein can diversify the data used for policy training through domain randomization, so that more changes and noise can be exposed during the policy training process, thereby improving the adaptability of the learned policy to uncertainty and changes. Optionally, these data may include but are not limited to the above-mentioned gain parameters, the passive stiffness / damping of the backlash joint, and physical property parameters that are difficult to model (for example, the weight and radius of the target object and the friction coefficient between the target object and the finger, etc.). By randomizing these parameters, the generalization ability of the learned policy can be enhanced.
[0079] According to an embodiment of the present disclosure, generating a training data set based on a calibrated joint model of each finger joint of the simulated robot may include: generating the multiple candidate initial states; and generating one or more of the following categories of data: the distribution of proportional gain and differential gain of each finger joint of the simulated robot; the distribution of passive stiffness or damping of the finger joints of the simulated robot corresponding to the tooth gap joints of the robot; and attribute parameters of the target object, the attribute parameters including one or more of the weight, radius and friction coefficient of the target object with the finger joints.
[0080] As mentioned above, by applying the proportional gain K to each knuckle P and differential gain K D By optimizing the solution, the probability distribution (e.g., a normal distribution) of the proportional gain and differential gain of each finger joint of the simulated robot can be determined. Furthermore, based on the range of the gap joints, the distribution of the passive stiffness or damping of the finger joints corresponding to the gap joints of the simulated robot can also be determined. Furthermore, other domain parameters can be randomly generated to increase the generalization capability of the policy, including, but not limited to, the weight and radius of the target object and the friction coefficient between the finger joints and the target object.
[0081] In addition, when using reinforcement learning to train dexterous in-hand manipulation strategies, there is often a problem of insufficient exploration of the state space because the target object may fall in the early stages of training, and as the strategy changes, the exploration becomes more limited. Therefore, strategies trained with insufficient data space are often more sensitive to interference and uncertainty. To address these problems, in an embodiment of the present disclosure, the initial states of the manipulated target object and the fingers of the simulated robot can be randomized to increase the ability of the strategy to handle interference and uncertainty. Optionally, a dataset of random finger joint positions and object poses (e.g., multiple candidates for the initial states of the target object and the fingers of the simulated robot) can be generated, and the training environment can be initialized by sampling from the dataset at the beginning of each training period.
[0082] According to an embodiment of the present disclosure, generating the multiple candidate initial states may include: sampling the finger joint positions of the simulated robot and the posture of the target object from a uniform distribution, wherein the range of the uniform distribution ensures that the target object is located between the fingers of the simulated robot; fixing the target object in the sampled posture and closing the fingers of the simulated robot at a constant speed until the fingers of the simulated robot contact the target object; and taking the current finger joint positions of the simulated robot and the posture of the target object as a candidate initial state.
[0083] Optionally, for better policy convergence and operability, in the configuration of the candidate initial state, all fingers of the simulated robot can be set to be in contact with the target object, but the grasp of the target object formed does not have to be stable, because first of all, the state is a transient state during dynamic motion, and the policy will learn to recover from such an unstable state and continue the in-hand operation task.
[0084] Optionally, the exemplary method for generating candidate initial states disclosed herein may be as shown in Algorithm 1 below.
[0085]
[0086] As shown in Algorithm 1, the algorithm takes the finger joint position distribution and the target object posture distribution as input, and outputs multiple (for example, N) candidate initial states. When generating the kth candidate initial state, the finger joint position q can be first sampled from the uniform distribution (the finger joint position distribution and the target object posture distribution). s and the pose p of the target object s , where the uniform distribution range should ensure that the target object is between the fingers. Next, the starting knuckle position can be set to the sampled knuckle position q s , and fix the target object at the sampled posture p s , which cannot be moved by any external force, by closing the fingers at a constant speed until the fingers come into contact with the target object, at which point the finger joint positions and the posture of the target object can be used as a candidate initial state.
[0087] Figure 6 is a schematic diagram illustrating example candidate initial states according to an embodiment of the present disclosure. Figure 6 As shown, three example candidate initial states ((a), (b) and (c)) are shown, wherein the initial state that cannot be obtained by a flat fingertip can be obtained by contacting a fingertip with a deformable surface with a target object.
[0088] Therefore, by the processing described above with reference to steps S201-S203, the preparation for strategy training can be completed, including building a simulation robot and generating training data for strategy training. Next, a training environment can be built based on the simulation robot and training data, and strategy training can be performed.
[0089] In step S204, a training environment may be initialized based on sampling from the training data set, and an in-hand operation strategy may be trained through interaction between the simulated robot and the training environment, wherein the in-hand operation strategy may take observations of the simulated robot at a current time step as input and take actions of the simulated robot at a next time step as output, and the observations of the simulated robot at the current time step may include the contact center position of each finger of the simulated robot with the target object at the current time step.
[0090] According to an embodiment of the present disclosure, initializing the training environment based on sampling from the training dataset may include initializing the training environment based on sampling of data for each category from the training dataset. Optionally, as described above, parameters involved in the policy training process may be randomized based on the sampling from the training dataset, and the states of the finger joint positions of the simulated robot and the posture of the target object may be randomly initialized to establish the training environment.
[0091] Optionally, at the beginning of each training period, each finger joint of the simulated robot can be set to an initial finger joint position, and the target object can be placed in an initial posture, and the gravity applied to the target object can be compensated within a predetermined time to allow the fingers of the simulated robot to close and establish contact with the target object.
[0092] In an embodiment of the present disclosure, the in-hand operation strategy of the simulated robot can be learned by deep reinforcement learning. Specifically, the in-hand operation task can be modeled as a finite-horizon discounted Markov decision process, and the control strategy can be learned using a deep reinforcement learning algorithm. The process can be composed of an action space A, a state space S, and a state transition dynamics process T: S×A→S and a reward function r: S×A→R modeled by a physical simulator. According to an embodiment of the present disclosure, training the in-hand operation strategy through the interaction between the simulated robot and the training environment may include: in multiple training periods, using a deep reinforcement learning algorithm to train the in-hand operation strategy by causing the simulated robot to control the target object to perform a predetermined finger motion task. That is, in an embodiment of the present disclosure, the above-mentioned multiple training periods can be used to train the strategy, wherein, in each training period, the simulated robot is trained to control the target object to perform a predetermined finger motion task according to the in-hand operation strategy.
[0093] According to an embodiment of the present disclosure, each training period includes multiple time steps, and training the in-hand operation strategy through the interaction of the simulated robot with the training environment may also include: for each training period, in each time step of the training period, obtaining an observation of the current state of the simulated robot, using the in-hand operation strategy to generate the action of the simulated robot in the next time step, and determining the reward of the generated action based on the reward function; and training the in-hand operation strategy by maximizing the expected discounted reward sum, wherein the expected discounted reward sum is the sum of the rewards for each time step.
[0094] Optionally, for each time step, the observation of the current state of the simulated robot of the present disclosure may include three parts: (a) the measured finger joint positions, (b) the unit vector of the main axis direction representing the desired target object posture in Indicates the endpoint of the target object (such as Figure 3 (as shown), and (c) the contact center position of each finger with the target object at the current time step. Optionally, the contact center position can be calculated as the average position of all tactile pixels on the finger that are in contact with the target object.
[0095] According to an embodiment of the present disclosure, initializing a training environment based on sampling from the training data set and training an in-hand operation strategy through the interaction between the simulated robot and the training environment may include: in the training, at the current time step, binarizing the tactile signal returned from the tactile sensor of the simulated robot to determine the contact center position of each finger of the simulated robot with the target object at the current time step.
[0096] Optionally, as described above, since the present disclosure uses Boolean values to represent the contact information between the robot's tactile sensor and the target object, for a simulated robot, corresponding to a real robot, the tactile signal returned by the tactile sensor of the simulated robot can also be binarized to indicate the contact position between the tactile sensor of the simulated robot and the target object. Specifically, according to an embodiment of the present disclosure, binarizing the tactile signal returned from the tactile sensor of the simulated robot to determine the contact center position of each finger of the simulated robot with the target object at the current time step may include: for each finger of the simulated robot, binarizing the value returned by each tactile pixel in the tactile sensor of the finger according to a predetermined threshold to determine the tactile pixel in the tactile sensor of the finger that is in contact with the target object; and determining the contact center position of the finger with the target object at the current time step based on the determined tactile pixel in the tactile sensor of the finger that is in contact with the target object.
[0097] Unlike rigid body simulations, real tactile sensors can be wrapped in deformable fingertips. When the finger is in contact with the target object, the piezoresistive sensing elements around the actual contact area but not in the actual contact area typically return a positive signal due to the pressure from the deformation. Therefore, in an embodiment of the present disclosure, a predetermined threshold can be set for the tactile signal, so that tactile pixels that return a value greater than the predetermined threshold can be considered to be in contact with the target object, and tactile pixels that return a value lower than the predetermined threshold can be filtered. As an example, the predetermined threshold can be set to an empirical value that filters out noise and passes the signal from the tactile pixels in contact with the target object. Of course, the predetermined threshold can also be set in other ways, and the present disclosure is not limited to this.
[0098] Based on the above operations, a number of tactile pixels in contact with the target object can be determined from the simulated robot's tactile sensor. Therefore, in embodiments of the present disclosure, the contact position of the finger with the target object at the current time step can be determined based on the positions of these tactile pixels. Optionally, the contact position can be the center position of the finger's contact with the target object at the current time step, which can be the average position of the positions of the aforementioned tactile pixels.
[0099] Alternatively, for each time step, the action for the next time step output by the in-hand manipulation strategy may be the expected displacement (or velocity) of the finger joints of the simulated robot.
[0100] As an example, at each time step t, the in-hand manipulation policy π can obtain the observation s of the current time step of the simulated robot t And generate the action a of the simulated robot in the next time step t =π(s t ). In the simulation robot based on the action a t After interacting with the environment, the reward r(s t ,a t ). Therefore, the goal of this in-hand operation strategy can be to maximize the expected discounted reward sum It is the weighted sum of the rewards at each time step t, where γ is a discount factor. Optionally, Proximal Policy Optimization (PPO) can be used as the DRL algorithm to train the in-hand operation policy. Of course, the present disclosure can also adopt other reinforcement learning algorithms, and the present disclosure is not limited to this.
[0101] According to an embodiment of the present disclosure, the reward function may include a positive task reward term, an orientation error penalty term, a position error penalty term and a contact force control term; wherein, the positive task reward term is used to reward a longer training period length, the orientation error penalty term is used to penalize the orientation error of the target object, the position error penalty term is used to penalize the position error between the desired posture and the measured posture of the target object, and the contact force control term is used to adjust the sum of the contact forces detected by the tactile sensors of the simulated robot.
[0102] Optionally, to alleviate over-reliance on human prior knowledge, the designed reward function can be positively correlated with the performance of a predetermined finger motion task for the target object.
[0103] As an example, the reward function R used in this disclosure can be expressed as follows:
[0104]
[0105] Among them, the reward function R can include the four terms as described above, ω0, ω1, ω2 represent weights (for example, [C, ω0, ω1, ω2] can take the value of [0.5, 1.5, 2.0, 0.005]). Among them, the first term C can correspond to the positive task reward term, which represents the positive task reward. Considering that if the height of the geometric center of the target object is lower than the threshold, it can indicate that it has fallen out of the robot's hand, and the current training period is terminated, therefore, the positive task reward term C (usually a constant term) can promote the learning of stable in-hand operations by rewarding a longer training period length. The second term || u d -u c ||2 can correspond to the orientation error penalty term, which can represent the desired orientation u of the target object d With the actual orientation u c The error between (for example, L2 norm error). It can correspond to a position error penalty term, which can represent the position error (e.g., L2 norm error) between the desired pose and the actual pose of the target object. That is, the second and third terms guide policy learning by directly quantifying the pose error of the target object. As for the fourth term, in order to avoid excessive force causing damage to the target object or the robot hand, the fourth term can be used to adjust the sum of the contact forces detected by the tactile sensors (e.g., the maximum combined fingertip force of each finger is limited to 15N), thereby promoting gentle manipulation skills, where f i is the magnitude of the contact force on each tactile pixel.
[0106] Of course, it should be understood that the above reward function is only used as an example and not as a limitation in this disclosure, and the method of this disclosure can also adopt other forms of reward functions.
[0107] Optionally, the predetermined finger motion task may include but is not limited to manipulating a target object with the fingers of a simulated robot to move according to a real-time target posture trajectory. Figure 3 As shown, the fingers of a simulated robot manipulate a dark-colored target object (a slender rod) to move according to the reference posture shown by the light-colored target object. This in-hand manipulation only involves finger movements, and the robot hand is fixed in a certain posture, in which the palm is perpendicular to the ground during the manipulation.
[0108] Optionally, for different predetermined finger motion tasks, different target object trajectories can be selected as reference trajectories, that is, different trajectories (for example, straight lines, circles, spirals, and number 8, etc.) are drawn using the end of the target object, and corresponding strategies are trained to perform the predetermined finger motion task.
[0109] Based on the policy training as described above, an in-hand operation policy that can be directly applied to a real robot can be learned. Therefore, in step S205, the in-hand operation policy can be applied to the in-hand operation of the robot after the training is completed. By simulating the robot with a physical simulator and training the policy to transfer it directly to the real robot, the time and manpower consumption and potential dangers in the collection of training data for the real robot can be avoided, and tactile-based dexterous in-hand operation of the target object can be achieved through deep reinforcement learning, wherein a better policy can be trained by using the estimated contact center position as an observation. In addition, the gap between the simulation and the real experiment can be reduced by calibrating the model of the finger joints and randomizing the initial state, so that the learned policy performs well in both simulation and real-world experiments, and has good generalization ability for new target objects and tasks.
[0110] Below, we will refer to Figure 7-11 The performance of the disclosed method for in-hand manipulation of a robot is described. The performance of the strategy learned using the disclosed method for in-hand manipulation of a robot in both simulated and real environments is presented, and the effectiveness of the tactile feedback is verified by comparing strategies trained using different observations of the target object.
[0111] Figure 7 3 is a result diagram showing the performance verification of the strategy trained for different predetermined finger motion tasks according to an embodiment of the present disclosure.
[0112] Optionally, in order to verify the effectiveness of the strategies learned by the method for in-hand manipulation of a robot disclosed herein, in the present disclosure, four strategies may be trained for four predetermined finger motion tasks based on four reference trajectories (e.g., a straight line, a circle, a spiral, and a figure 8) as described above, i.e., controlling the endpoints of a target object to draw a straight line, a circle, a spiral, and a figure 8. Figure 7 The performance of the strategy trained on these four predetermined finger movement tasks on a simulated robot is shown. Figure 7 As shown, in Figure 7 Figures (a), (b), (c), and (d) show the reference trajectory (target position) and the actual trajectory (actual position) of the target object's lower endpoint position for each predetermined finger motion task, respectively. In other words, the trained policy is able to manipulate the target object to move along a continuous reference trajectory. The spiral trajectory is the most challenging for the robot because it requires more delicate manipulation, especially in the inner circle, resulting in a larger error in the resulting trajectory.
[0113] Figure 8 is a graph illustrating changes in the positions of the knuckles of the fingers of a simulated robot during the execution of a predetermined finger motion task according to an embodiment of the present disclosure.
[0114] Optionally, in order to analyze the strategy learned by the method for in-hand operation of a robot disclosed in the present invention, Figure 8 The finger joint positions during the target object manipulation in the simulation are plotted using a predetermined finger motion task based on a circular reference trajectory as an example. Figure 8 As shown, finger 1 (thumb) remains almost stationary during the operation, providing a fixed support surface for the target object, while fingers 2 and 3 are used to control the movement of the target object. Their joint motions are periodic and highly correlated, which also demonstrates the coordinated operation between different fingers and knuckles.
[0115] Figure 9 is a graph showing results of errors in performing different predetermined finger motion tasks according to an embodiment of the present disclosure. Figure 9 shows the average position error p of the endpoint of the target object during different predetermined finger motion tasks err and the orientation error q of the target object err Among them, Figure 9 The left part of shows the error under the training setting, which shows that the learned policy using tactile feedback can manipulate the target object to adjust to the reference pose in real time, and the error increases with the complexity of the reference trajectory (e.g., from a straight trajectory with random directions to a spiral trajectory).
[0116] In addition, in order to evaluate the generalization ability of the trained policy, Figure 9In the paper, a strategy trained for a predetermined finger motion task based on a circular reference trajectory is also selected as an example to show that the strategy can also achieve good performance on circular trajectories with new expected rotation speeds (e.g., a larger expected rotation speed v+ or a smaller expected rotation speed v-) and radii (e.g., a larger radius r+ or a smaller radius r-), which indicates the good generalization ability of the trained strategy.
[0117] Next, to evaluate the effectiveness of using the contact center position as a form of tactile feedback in learning dexterous in-hand manipulation strategies, this paper conducts a comparative study in which ablation experiments are performed to obtain the results of training strategies using different observations of the target object. Figure 10 is a diagram showing the results of training a policy using different observations of a target object according to an embodiment of the present disclosure.
[0118] Optionally, compared with using (a) the contact center position on each finger as tactile feedback to train the strategy, as an example, the present disclosure uses (b) object posture, (c) object posture and contact center position, (d) object posture and binary information of whether each finger is in contact with the target object, and (e) original tactile information as tactile feedback to train the strategy. Figure 10 In , the learning curves of the policy training with each tactile feedback are shown, which correspond to the change of rewards as time steps increase. Figure 10 As shown, the policy trained using contact center location as tactile feedback ((a)) achieves the best performance, and adding the ground truth object pose to the observation space does not provide further performance improvement. In addition, the policy trained without any tactile information cannot achieve comparable performance, indicating that tactile information is more important than the object pose provided by visual information in the in-hand manipulation task.
[0119] As shown in the curves corresponding to (a), (d), and (e), tactile information from binary contact checks of the robot's fingers (e.g., 1×3 dimensions for a three-fingered robot hand) is ineffective for learning appropriate manipulation strategies, while raw tactile data (e.g., 128×3 dimensions) requires a more complex network structure (e.g., a graph convolutional network) to extract spatial information. Compared to these two approaches, the disclosed approach of using contact center positions (e.g., 3x3 dimensions) as tactile feedback provides more direct and useful information and achieves better performance in the target object rotation task with the same amount of training data and network structure.
[0120] Through the above comparative studies, it can be seen that the strategy of using the center contact position as tactile feedback training is significantly better than the strategy of using other forms of tactile feedback training, which also highlights the effectiveness of low-dimensional tactile feedback for manipulating slender cylindrical objects by extracting contact information between the target object and each fingertip from high-dimensional tactile information.
[0121] Finally, to validate the trained policy in the real world, the present disclosure can deploy the trained policy to a real robotic hand with tactile sensors. Figure 11 is a result diagram illustrating contact position trajectories of one finger of a robot in different directions and activation of a tactile sensor during execution of a predetermined finger motion task according to an embodiment of the present disclosure.
[0122] Alternatively, the experiment can be performed from Figure 3 and Figure 6 The target object shown starts from an initial state where it is already grasped by three fingertips. During operation, the robot hand can be fixed in a specific posture by the robot arm.
[0123] Alternatively, due to deformation of the material wrapping the tactile sensor, some sensor readings may drift, resulting in a positive value being returned even when there is no contact. Since a fixed threshold is used in the present disclosure to binarize the tactile readings, such sensor drift may result in errors in distinguishing contact states. Therefore, the tactile sensor may be calibrated before each operation experiment. For example, the sensor readings without any contact may be recorded for a period of time, and then the average value may be used. As an offset, the calibrated tactile reading s(t) can therefore be calculated as follows:
[0124]
[0125] Among them, s * Represents the raw tactile readings. Negative values of s(t) can be considered invalid and set to 0.
[0126] Based on this, Figure 11 Figure 2 shows the contact position trajectories of a robot finger in different directions during a predetermined finger motion task (e.g., a predetermined finger motion task based on a circular reference trajectory) in both simulation and real experiments. Note that the contact position trajectories in simulation and real experiments are not expected to be identical because their initial states are different. The trajectories obtained from both experiments are presented in the same figure for qualitative analysis.
[0127] like Figure 11As shown on the left side of , both the real and simulated trajectories exhibit periodicity, especially in the Z direction, which indicates that the learned policy can utilize the curved finger surface to manipulate the target object to keep the contact position within a feasible range. Figure 11 The right side of the figure shows examples of tactile sensors activated in simulation experiments and real experiments. Due to the uncertainty of fingertip deformation and the inconsistency of the sensitivities of different piezoresistive sensing elements, the contact area detected in the real experiment has more irregular edges than that in the simulation experiment. However, the above-mentioned difference in tactile feedback is instantaneous, and the characteristics of the real tactile sensor are consistent with the simulation most of the time. Therefore, in the method of the present disclosure, although pure simulated tactile information is used for training, the trained strategy is able to adapt to the gap between simulation and reality and is able to perform well in real-world experiments. In addition, through real-time tactile perception, although trained with a single target object in simulation, the trained strategy can adapt to the diversity of radius, shape, weight and friction coefficient of the target object.
[0128] Figure 12 is a schematic diagram illustrating an apparatus 1200 for in-hand manipulation of a robot according to an embodiment of the present disclosure.
[0129] According to an embodiment of the present disclosure, the apparatus 1200 for in-hand operation of a robot may include a simulation modeling module 1201 , a model calibration module 1202 , a data preparation module 1203 , a strategy training module 1204 and a strategy application module 1205 .
[0130] The simulation modeling module 1201 can be configured to simulate and model the robot in a physical simulator to obtain a simulated robot having the same features as the robot, including the backlash and self-locking features of the finger joints. Optionally, the simulation modeling module 1201 can perform the operations described above with reference to step S201.
[0131] In an embodiment of the present disclosure, a physical simulator can be used for robot training to learn in-hand manipulation strategies for target objects from simulation experiments and apply them to in-hand manipulations of real robots, thereby saving costs and time, speeding up the iteration and optimization process of the algorithm, and avoiding these potential risks to ensure the safety of the training process.
[0132] Alternatively, the physical simulator can be any of a variety of existing open-source robotics simulators, such as the MuJoCo (Multi-Joint Dynamics with Contact) physics engine. Of course, the aforementioned physics engine is used in this disclosure only as an example and not as a limitation, and the method of this disclosure can also be used with other physical simulators for policy training.
[0133] Alternatively, the simulation modeling of the robot in the physical simulator can be based on the replication of the structure and features of the real robot to minimize the gap between the simulated and real experiments of the robot and to fully utilize the potential of the tactile sensor.
[0134] The model calibration module 1202 may be configured to calibrate the joint model of each joint of each finger of the simulated robot. Optionally, the model calibration module 1202 may perform the operations described above with reference to step S202.
[0135] Optionally, the dynamic characteristics of the finger joints may be calibrated. These dynamic characteristics may include but are not limited to the dynamic model of the finger joints, and the tooth clearance and self-locking characteristics of the finger joints.
[0136] Optionally, to calibrate the parameters of the knuckle dynamics model, the proportional gain and differential gain of each knuckle can be calibrated during the simulation. In embodiments of the present disclosure, parameter calibration can be performed for each knuckle. For each knuckle, the knuckle can be controlled to draw the same reference trajectory during both the simulation and the real experiment, allowing parameter calibration to be performed by comparing the knuckle position trajectories of the simulated and real robots. For example, parameters can be optimized by minimizing the error between the knuckle position trajectories of the simulated and real robots.
[0137] Alternatively, the range of the backlash joint can be calibrated experimentally. For example, the target joint (i.e., the backlash joint) can be controlled to remain at a specific position, and then an external torque is applied to rotate the target joint as much as possible to determine the static joint position and extreme joint positions of the backlash joint.
[0138] In addition to the aforementioned dynamic and backlash characteristics, embodiments of the present disclosure also model the self-locking characteristics of the robot's proximal finger joints. Because the proximal finger joints utilize a worm gear mechanism as their motion transmission device, reverse drive is not permitted, meaning they cannot move against the drive direction. Therefore, the aforementioned self-locking characteristics can be simulated by adjusting the joint position and velocity of the proximal finger joints in the simulated robot.
[0139] Through the finger joint model calibration as described above, the joint model of the simulated robot's finger joints can be prepared for policy training, and the gap between the simulation experiment and the real experiment can be minimized as much as possible. Next, before conducting policy training, it is necessary to construct a training dataset for policy training.
[0140] Data preparation module 1203 can be configured to generate a training dataset based on the calibrated joint model of each finger joint of the simulated robot, the training dataset comprising a plurality of candidate initial states, wherein each candidate initial state comprises a finger joint position of the simulated robot and a pose of a target object. Optionally, data preparation module 1203 can perform the operations described above with reference to step S203.
[0141] Optionally, domain randomization can be used to diversify the data used for policy training, so that more changes and noise can be exposed during policy training, thereby improving the adaptability of the learned policy to uncertainty and changes. Optionally, this data may include but is not limited to the aforementioned gain parameters, the passive stiffness / damping of the backlash joint, and physical property parameters that are difficult to model (for example, the weight and radius of the target object, and the friction coefficient between the target object and the finger). By randomizing these parameters, the generalization ability of the learned policy can be enhanced.
[0142] As described above, by optimizing the proportional gain and differential gain of each finger joint, the probability distribution (e.g., a normal distribution) of the proportional gain and differential gain of each finger joint of the simulated robot can be determined. Furthermore, based on the range of the gap joints, the distribution of the passive stiffness or damping of the finger joints corresponding to the gap joints of the simulated robot can also be determined. Furthermore, other domain parameters can be randomly generated to increase the generalization capability of the policy, including, but not limited to, the weight and radius of the target object and the friction coefficient between the finger joints and the target object.
[0143] In addition, when using reinforcement learning to train dexterous in-hand manipulation strategies, there is often a problem of insufficient exploration of the state space because the target object may fall in the early stages of training, and as the strategy changes, the exploration becomes more limited. Therefore, strategies trained with insufficient data space are often more sensitive to interference and uncertainty. To address these problems, in an embodiment of the present disclosure, the initial states of the manipulated target object and the fingers of the simulated robot can be randomized to increase the ability of the strategy to handle interference and uncertainty. Optionally, a dataset of random finger joint positions and object poses (e.g., multiple candidates for the initial states of the target object and the fingers of the simulated robot) can be generated, and the training environment can be initialized by sampling from the dataset at the beginning of each training period.
[0144] Optionally, for better policy convergence and operability, in the configuration of the candidate initial state, all fingers of the simulated robot can be set to be in contact with the target object, but the grasp of the target object formed does not have to be stable, because first of all, the state is a transient state during dynamic motion, and the policy will learn to recover from such an unstable state and continue the in-hand operation task.
[0145] Therefore, through the above process, the preparation work for strategy training can be completed, including building a simulation robot and generating training data for strategy training. Next, a training environment can be built based on the simulation robot and training data, and strategy training can be performed.
[0146] The policy training module 1204 can be configured to initialize a training environment based on samples from the training dataset and train an in-hand manipulation policy through interaction between the simulated robot and the training environment, wherein the in-hand manipulation policy takes the observation of the simulated robot at the current time step as input and outputs the action of the simulated robot at the next time step, wherein the observation of the simulated robot at the current time step includes the contact center position of each finger of the simulated robot with the target object at the current time step. Optionally, the policy training module 1204 can perform the operations described above with reference to step S204.
[0147] Optionally, the parameters involved in the strategy training process may be randomized based on sampling from the training data set, and the states of the finger joint positions of the simulated robot and the posture of the target object may be randomly initialized to build a training environment.
[0148] Optionally, at the beginning of each training period, each finger joint of the simulated robot can be set to an initial finger joint position, and the target object can be placed in an initial posture, and the gravity applied to the target object can be compensated within a predetermined time to allow the fingers of the simulated robot to close and establish contact with the target object.
[0149] In an embodiment of the present disclosure, a deep reinforcement learning algorithm can be used to learn the in-hand operation strategy of a simulated robot. Specifically, the in-hand operation task can be modeled as a finite-period discounted Markov decision process, and a deep reinforcement learning algorithm can be used to learn the control strategy.
[0150] The strategy application module 1205 may be configured to apply the in-hand operation strategy to the in-hand operation of the robot after the training is completed. Optionally, the strategy application module 1205 may perform the operations described above with reference to step S205.
[0151] Based on the policy training, a policy for in-hand manipulation can be learned that can be directly applied to the real robot. Therefore, after the training is completed, the in-hand manipulation policy can be applied to the in-hand manipulation of the robot.
[0152] According to yet another aspect of the present disclosure, a device for in-hand operation of a robot is also provided. Figure 13 A schematic diagram of a device 2000 for in-hand manipulation of a robot is shown according to an embodiment of the present disclosure.
[0153] like Figure 13 As shown, the apparatus 2000 for in-hand operation of a robot may include one or more processors 2010 and one or more memories 2020. The memories 2020 may store computer-readable codes, which, when executed by the one or more processors 2010, may execute the method for in-hand operation of a robot as described above.
[0154] The processor in the embodiments of the present disclosure may be an integrated circuit chip having signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor, and may be an X86 architecture or an ARM architecture.
[0155] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0156] For example, the method or apparatus according to the embodiment of the present disclosure may also be implemented by Figure 14 The architecture of the computing device 3000 shown in FIG. Figure 14 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method for in-hand operation of a robot provided by the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 14 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 14One or more components of a computing device are shown.
[0157] According to another aspect of the present disclosure, a computer-readable storage medium is also provided. The computer storage medium has computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the method for in-hand operation of a robot according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0158] Embodiments of the present disclosure also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method for in-hand operation of a robot according to an embodiment of the present disclosure.
[0159] Embodiments of the present disclosure provide a method, apparatus, device, and computer-readable storage medium for in-hand manipulation of a robot.
[0160] The method provided by the embodiments of the present disclosure trains an in-hand manipulation policy for a three-fingered robotic hand equipped with tactile sensors in a physical simulator and applies the learned in-hand manipulation policy to a real robot. A deep reinforcement learning method is used to train the policy using the contact center position between the robot hand and the target object as a partial observation. The finger joint models are calibrated by considering the robot's mechanical backlash and self-locking characteristics, and the initial training states are randomized to ensure successful transfer of the trained policy to the real robot, thereby enabling the robot to perform dexterous continuous in-hand manipulation of the target object. The method of the embodiments of the present disclosure simulates the robot using a physical simulator and trains the policy for direct transfer to the real robot, avoiding the time and manpower required to collect training data for the real robot and the potential risks involved. Deep reinforcement learning enables dexterous in-hand manipulation of the target object based on tactile perception, wherein a more optimized policy is trained by using the estimated contact center position as an observation. Furthermore, calibrating the finger joint model and randomizing the initial states mitigate the gap between simulation and real-world experiments, allowing the learned policy to perform well in both simulation and real-world experiments and generalize well to new target objects and tasks.
[0161] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of the code, and the module, program segment, or a part of the code contains at least one executable instruction for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0162] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0163] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will appreciate that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for in-hand manipulation of a robot, the robot having at least one finger, wherein each finger has at least one knuckle, and the tip of each finger is configured with a tactile sensor, the method comprising: In a physical simulator, the robot is simulated and modeled to obtain a simulated robot, wherein the simulated robot has the same features as the robot, including tooth clearance and self-locking features of finger joints; For each knuckle of each finger of the simulated robot, calibrate the joint model of the knuckle; generating a training data set based on a calibrated joint model of each finger joint of the simulated robot, the training data set comprising a plurality of candidate initial states, wherein each candidate initial state comprises a finger joint position of the simulated robot and a posture of a target object; Initializing a training environment based on samples from the training data set, and training an in-hand manipulation strategy through interaction between the simulated robot and the training environment, wherein the in-hand manipulation strategy takes observations of the simulated robot at a current time step as input and takes actions of the simulated robot at a next time step as output, wherein the observations of the simulated robot at the current time step include the contact center position of each finger of the simulated robot with the target object at the current time step; and After the training is completed, the in-hand operation strategy is applied to the in-hand operation of the robot.
2. The method according to claim 1, wherein For each knuckle of each finger of the simulated robot, calibrating the joint model of the knuckle includes: calibrating a proportional gain and a differential gain of the finger joint, wherein the proportional gain and the differential gain are used to control a torque applied to the finger joint; Modeling a backlash joint of the robot and calibrating a range of the backlash joint, wherein the backlash joint corresponds to a joint where backlash exists; and The self-locking feature of the proximal finger joints of the robot is modeled.
3. The method according to claim 2, wherein: Calibrating the proportional gain and the differential gain of the finger joint includes: controlling the finger joints of the simulated robot and the corresponding finger joints of the robot to move along the same reference trajectory; and Proportional gains and derivative gains of the finger joints are optimized by minimizing the error between the actual trajectories of the finger joints of the simulated robot and the corresponding finger joints of the robot.
4. The method according to claim 2, wherein: Modeling the gap joint of the robot and calibrating the gap joint scope include: determining static joint positions of a backlash joint of the robot and extreme joint positions under application of an external torque; and Based on the static joint position and the extreme joint positions of the backlash joint, a range of the backlash joint is determined.
5. The method according to claim 2, wherein: Modeling the self-locking feature of the proximal finger joint of the robot includes: The self-locking feature of the proximal finger joints is simulated by adjusting the position and velocity of the corresponding finger joints of the simulated robot.
6. The method of claim 1, wherein: Using a Boolean value to represent contact information between the robot's tactile sensor and the target object; Initializing a training environment based on sampling from the training data set, and training an in-hand operation strategy through interaction between the simulated robot and the training environment includes: During the training, at a current time step, the tactile signal returned from the tactile sensor of the simulated robot is binarized to determine the contact center position of each finger of the simulated robot with the target object at the current time step.
7. The method according to claim 6, wherein: The tactile sensor of each finger of the simulated robot includes a tactile pixel array, the tactile pixel array includes a plurality of tactile pixels, each tactile pixel returns a value proportional to a normal force applied thereto; Binarizing the tactile signal returned from the tactile sensor of the simulated robot to determine the contact center position of each finger of the simulated robot with the target object at the current time step includes: For each finger of the simulated robot, binarize the value returned by each tactile pixel in the tactile sensor of the finger according to a predetermined threshold value to determine the tactile pixel in the tactile sensor of the finger that is in contact with the target object; and Based on the determined tactile pixels in the tactile sensor of the finger that are in contact with the target object, a contact center position of the finger and the target object at a current time step is determined.
8. The method of claim 1, wherein: Generating a training dataset based on the calibrated joint model of each finger joint of the simulated robot includes: generating the plurality of candidate initial states; and Generate data in one or more of the following categories: distribution of proportional gain and derivative gain of each finger joint of the simulated robot; distribution of passive stiffness or damping of finger joints of the simulated robot corresponding to the backlash joints of the robot; and The attribute parameters of the target object include one or more of the weight, radius, and friction coefficient of the target object with the finger joint.
9. The method of claim 8, wherein: Generating the multiple candidate initial states includes: Sampling finger joint positions of the simulated robot and the posture of the target object from a uniform distribution, wherein the range of the uniform distribution ensures that the target object is located between the fingers of the simulated robot; Fixing the target object in the sampled posture, and closing the fingers of the simulated robot at a constant speed until the fingers of the simulated robot come into contact with the target object; and The current finger joint positions of the simulated robot and the posture of the target object are taken as a candidate initial state.
10. The method of claim 8, wherein: Initializing the training environment based on sampling from the training dataset includes: The training environment is initialized based on sampling of data for each category from the training dataset.
11. The method according to claim 10, wherein: Training the in-hand operation strategy through the interaction between the simulated robot and the training environment includes: During multiple training sessions, the in-hand manipulation strategy is trained using a deep reinforcement learning algorithm by causing the simulated robot to control the target object to perform a predetermined finger motion task.
12. The method of claim 11, wherein: Each training epoch consists of multiple time steps; Wherein, training the in-hand operation strategy through the interaction between the simulated robot and the training environment further includes: For each training period, in each time step of the training period, obtaining an observation of the current state of the simulated robot, generating an action of the simulated robot for the next time step using the in-hand manipulation strategy, and determining a reward for the generated action based on a reward function; and The in-hand manipulation policy is trained by maximizing the expected sum of discounted rewards, which is the sum of the rewards at each time step.
13. The method of claim 12, wherein: The reward function includes a positive task reward term, an orientation error penalty term, a position error penalty term, and a contact force control term; Among them, the positive task reward item is used to reward a longer training period length, the orientation error penalty item is used to penalize the orientation error of the target object, the position error penalty item is used to penalize the position error between the expected posture and the measured posture of the target object, and the contact force control item is used to adjust the sum of the contact forces detected by the tactile sensor of the simulated robot.
14. A device for in-hand manipulation of a robot, the robot having at least one finger, wherein each finger has at least one knuckle, and the tip of each finger is configured with a tactile sensor, the device comprising: a simulation modeling module configured to simulate and model the robot in a physical simulator to obtain a simulated robot, wherein the simulated robot has the same features as the robot, including tooth clearance and self-locking features of the finger joints; a model calibration module, configured to calibrate a joint model of each joint of each finger of the simulated robot; a data preparation module configured to generate a training data set based on a calibrated joint model of each finger joint of the simulated robot, the training data set comprising a plurality of candidate initial states, wherein each candidate initial state comprises a finger joint position of the simulated robot and a posture of a target object; a policy training module configured to initialize a training environment based on samples from the training dataset, and train an in-hand manipulation policy through interaction between the simulated robot and the training environment, wherein the in-hand manipulation policy takes observations of the simulated robot at a current time step as input and takes actions of the simulated robot at a next time step as output, wherein the observations of the simulated robot at the current time step include the contact center position of each finger of the simulated robot with the target object at the current time step; and The strategy application module is configured to apply the in-hand operation strategy to the in-hand operation of the robot after the training is completed.
15. A device for in-hand manipulation of a robot, comprising: one or more processors; as well as One or more memories storing a computer executable program, which, when executed by the processor, performs the method of any one of claims 1 to 13. 16 . A computer-readable storage medium having computer-executable instructions stored thereon, wherein the instructions are used to implement the method according to claim 1 when executed by a processor.