Training method, device and system for grabbing strategy model, and grabbing method, device and system for grabbing strategy model

Through the deep reinforcement learning training crawling strategy model, combined with multi-dimensional operation text and current state data, the problem of incomplete crawling evaluation in the existing technology is solved, and a more general and stable crawling strategy is achieved.

CN120155933APending Publication Date: 2025-06-17PACINI PERCEPTION TECH (ZHANGJIAGANG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311725125.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing robot crawling evaluation methods lack comprehensive evaluation, cannot effectively judge the global stability of crawling, and lack universality, making it difficult to train crawling strategies that adapt to different scenarios and methods.

Method used

The deep reinforcement learning method is used to train the crawling strategy model. Through multi-dimensional operation text prompt information and current crawling status data, combined with the preset expected crawling pose, the crawling strategy model is iteratively trained, and the reward function is optimized to improve the universality of the model.

Benefits of technology

A more comprehensive crawl quality evaluation and more general crawl strategies are achieved, improving the stability and operability of the robot crawl process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120155933A_ABST
    Figure CN120155933A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the technical field of artificial intelligence, and relates to a training method of a grabbing strategy model. The training method comprises the steps of obtaining operation text prompt information and current grabbing state data in a current grabbing state; the operation text prompt information and at least the current image data serve as input of a preset expected grabbing model, and a current expected grabbing posture is obtained through output; the operation text prompt information and the current grabbing state data serve as input of the grabbing strategy model, the current expected grabbing posture is combined, and the grabbing strategy model is iteratively trained on the basis of a deep reinforcement learning method till a termination condition is met. The invention further provides a grabbing method, a related device, a related system and the like. According to the technical scheme, the universality of the grabbing strategy model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a training method of a grasping strategy model, a grasping method, a device and a system. Background Art

[0002] In robot grasping, the grasping quality is usually used to evaluate the stability of grasping. Currently, the commonly used evaluation methods include using visual recognition to determine whether the grasping is successful, or using force sense and / or tactile sensing to detect the grasping force, the distribution and magnitude of the contact force, and the contact area between the finger and the object to determine whether the grasping is stable. However, this evaluation method has the following problems:

[0003] 1. Currently, many evaluation methods are based on a single sensor judgment, but these methods may not be able to provide a comprehensive evaluation of the grasping quality. For example, the vision-based method cannot give an accurate judgment in the case of occlusion, and the force and tactile sensors can usually only give local information and cannot give a good index for the global stability of grasping.

[0004] 2. In the process of robot grasping, stability is not the only consideration. The grasping evaluation criteria should give corresponding evaluations according to the operability of the object after grasping the object, because the quality of grasping will affect the stability of subsequent robot operations.

[0005] 3. Currently, these methods also lack good generality and cannot provide different evaluation criteria according to different grasping scenarios and methods. Therefore, in the process of training grasping, it is very difficult for these methods to help train a good and general grasping strategy. Summary of the Invention

[0006] The purpose of the embodiments of this application is to propose a training method of a grasping strategy model, a grasping method, a device and a system to improve the generality of the grasping strategy model.

[0007] In the first aspect, the embodiments of this application provide a training method of a grasping strategy model, including the following technical solutions:

[0008] A training method of a grasping strategy model, applied to a grasping training system, includes: a robot, a state data acquisition sensor and a controller; the robot includes: a robotic arm and a grasping actuator arranged at the grasping execution end of the robotic arm; the state data acquisition sensor includes an image sensor, a force / tactile sensor arranged at the grasping actuator and a joint encoder; the force / tactile sensor is arranged at the grasping actuator, and the method includes the following steps:

[0009] Acquire operation text prompt information and current grasping state data under the current grasping state; the current grasping state data includes: current image data output by the image sensor; current force data output by the force / tactile sensor; and the current grasping posture of the robot obtained based on the joint motion output by the robot's encoder;

[0010] Using the operation text prompt information and at least the current image data as inputs of a preset expected grasping model, and outputting a current expected grasping posture;

[0011] The operation text prompt information and the current grasping state data are used as inputs of the grasping strategy model, and combined with the current expected grasping posture, the grasping strategy model is iteratively trained based on a deep reinforcement learning method until a termination condition is met.

[0012] Furthermore, in one embodiment, the current expected grasping posture includes: expected force data in the grasping state and an expected grasping posture of the robot in the grasping state.

[0013] Furthermore, in one embodiment, the operation text prompt information and the current grasping state data are used as inputs of the grasping strategy model, combined with the current expected grasping posture, and the grasping strategy model is iteratively trained based on a deep reinforcement learning method until a termination condition is met, including the following steps:

[0014] The operation text prompt information and the current grasping state data are used as inputs of the grasping strategy model, and the grasping posture of the robot at the next moment is obtained as output;

[0015] Obtaining a posture difference between the grasping posture at the next moment and the current expected grasping posture; obtaining a force difference between the current force data and the expected force data; and calculating a corresponding reward function according to the posture difference and the force difference;

[0016] Based on the reward function, optimizing the grasping strategy model;

[0017] Generate a grasping instruction based on the grasping posture at the next moment, so as to instruct the robot to perform a grasping motion through the grasping instruction;

[0018] Repeat the operation of obtaining the operation text prompt information and the current grasping state data in the current grasping state; use the operation text prompt information and at least the current image data as the input of a preset expected grasping model; use the operation text prompt information and the current grasping state data as the input of the grasping strategy model, and output the next moment's grasping posture of the robot; calculate the posture difference between the next moment's grasping posture and the current expected grasping posture; calculate the force difference between the current force data and the expected force data; calculate the corresponding reward function according to the posture difference and the force difference; optimize the grasping strategy model based on the reward function; generate a grasping instruction based on the next moment's grasping posture, so as to instruct the robot to execute the grasping motion through the grasping instruction until the termination condition is satisfied.

[0019] Further, in one embodiment, the current grasping posture includes: the end pose information of the robot at the current moment and the pose data of the grasping actuator at the current moment; and / or,

[0020] The next moment's grasping posture includes: the end pose information of the robot at the next moment and the pose data of the grasping actuator at the next moment;

[0021] The expected grasping posture includes: the expected end pose information of the robot at the current moment and the expected pose data of the grasping actuator at the current moment.

[0022] Further, in one embodiment, before using the operation text prompt information and at least the current image data as the input of a preset expected grasping model and outputting the current expected grasping posture, the following steps are further included:

[0023] Obtain training samples; the training samples include: the operation text prompt information and the training target image data in the corresponding grasping state; use the training grasping posture data corresponding to the operation text prompt information as the annotation of the training samples; the training grasping posture data includes the training force data in the grasping state and the training grasping posture of the robot in the grasping state;

[0024] Train the expected grasping model based on the training samples and the corresponding annotations to obtain the preset expected grasping model.

[0025] Further, in one embodiment, before obtaining the training samples, the following steps are further included:

[0026] Generate the corresponding annotations and the training target image data in the grasping state based on the operation text prompt information.

[0027] In a second aspect, an embodiment of the present application provides a grasping method of the grasping strategy model obtained based on the training method of the grasping strategy model described above, the method comprising the following steps:

[0028] Acquire the operation text prompt information and current grasping state data; the current grasping state data includes: current image data output by an image sensor; current force data output by a force / tactile sensor; and the current grasping posture of the robot obtained based on the joint motion output by an encoder of the robot;

[0029] Obtaining the crawling strategy model;

[0030] The operation text prompt information and the current grasping posture are used as inputs of the grasping strategy model, and the grasping posture of the robot at the next moment is obtained as output;

[0031] Generate a grasping instruction based on the grasping posture at the next moment, so as to instruct the robot to grasp the target object through the grasping instruction;

[0032] Repeat the above steps until the target object is captured.

[0033] In a third aspect, an embodiment of the present application provides a training device for a grasping strategy model, the device comprising:

[0034] A data acquisition module is used to acquire operation text prompt information and current grasping state data under the current grasping state; the current grasping state data includes: current image data output by the image sensor; current force data output by the force / tactile sensor; and the current grasping posture of the robot obtained based on the joint motion output by the robot's encoder;

[0035] an expectation generation module, used to use the operation text prompt information and at least the current image data as inputs of a preset expected grasping model, and output a current expected grasping posture;

[0036] The model training module is used to use the operation text prompt information and the current grasping state data as inputs of the grasping strategy model, and iteratively train the grasping strategy model based on a deep reinforcement learning method in combination with the current expected grasping posture until a termination condition is met.

[0037] In a fourth aspect, an embodiment of the present application provides a grasping device based on the grasping strategy model trained by the training device of the grasping strategy model described above, the device comprising:

[0038] An information acquisition module, used for a text acquisition module to acquire the operation text prompt information and the current grasping state data; the current grasping state data includes: the current image data output by an image sensor; the current force data output by a force / tactile sensor; the current grasping posture of the robot obtained by calculating the joint movement amount based on the output of the robot's encoder.

[0039] A model acquisition module, used to acquire the grasping strategy model.

[0040] A posture generation module, used to use the operation text prompt information and the current grasping state data as inputs to the grasping strategy model, and output the next moment's grasping posture of the robot.

[0041] An instruction generation module, used to generate a grasping instruction based on the next moment's grasping posture, so as to instruct the robot to grasp the target object through the grasping instruction.

[0042] A step repetition module, used to repeat the above steps until the grasping of the target object is completed.

[0043] In a fifth aspect, an embodiment of the present application provides a grasping system, the system includes: a robot, a state data acquisition sensor, and a controller; the robot includes: a robotic arm and a grasping actuator arranged at the grasping execution end of the robotic arm.

[0044] The controller is respectively communicatively connected to the robot and the state data acquisition sensor.

[0045] The controller is used to implement the training method of the above-mentioned grasping strategy model and / or the above-mentioned grasping method.

[0046] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:

[0047] In the embodiments of the present application, multi-dimensional operation text prompt information and partial current grasping state data are used as inputs to a preset expected grasping model, and the current expected grasping posture can be output (the preset expected grasping model is a pre-trained expected grasping model, that is, the current expected grasping posture is obtained based on a learning method); then, the operation text prompt information and the current grasping state data are used as inputs to the grasping strategy model, combined with the current expected grasping posture, and the grasping strategy model is iteratively trained based on the deep reinforcement learning method. The present application obtains the corresponding expected grasping posture in a learning manner, and calculates the corresponding reward function based on this, and can optimize to obtain a more general grasping strategy model. Description of the Drawings

[0048] To more clearly illustrate the solutions in this application, the following provides a brief introduction to the drawings required for the description of the embodiments of this application. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 It is a system architecture diagram of an embodiment of the application system of this application;

[0050] Figure 2 It is a system architecture diagram of an embodiment of the training sample acquisition system based on teleoperation of this application;

[0051] Figure 3 It is a structural schematic diagram of an embodiment of the training device of the grasping strategy model of this application;

[0052] Figure 4 It is a flowchart of an embodiment of the training method of the grasping strategy model of this application;

[0053] Figure 5 It is a flowchart of an embodiment of the grasping method of this application;

[0054] Figure 6 It is a structural schematic diagram of an embodiment of the training device of the grasping strategy model of this application;

[0055] Figure 7 It is a structural schematic diagram of an embodiment of the grasping device of this application;

[0056] Figure 8 It is a structural schematic diagram of an embodiment of the computer device of this application. Detailed implementation manners

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used in the description of the embodiments of this application in this specification are only for the purpose of describing specific embodiments, and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects, rather than to describe a specific order.

[0058] References to "embodiments" in this specification mean that particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0059] To enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0060] As Figure 1 shown, Figure 1 is the system architecture diagram of an embodiment of the application system of the present application.

[0061] In one embodiment, the present application provides a grasping system, which includes: a robot 110, a state data sensor 120, and a controller 130.

[0062] Robot

[0063] Specifically, the robot 110 can be, but is not limited to: a humanoid robot; or a manipulator connected in series or parallel (such as: a Delta manipulator, a four-axis manipulator, or a six-axis manipulator).

[0064] The robot includes: a robotic arm and a grasping actuator 111 disposed at the grasping execution end of the robotic arm; the grasping actuator (such as: a dexterous hand or a gripper).

[0065] Status data sensor

[0066] In one embodiment, the state data sensor 120 can include, but is not limited to: an image sensor 121, a force / tactile sensor 122 disposed on the grasping actuator, and a joint encoder. In addition, other required sensors can also be used as needed.

[0067] By adopting the above state data sensor in the embodiments of the present application, multi-dimensional state data can be obtained from three aspects: vision, touch, and the robot, thereby improving the generality of the subsequent model.

[0068] The image sensor 121 can be fixed to the robot (such as: fixed to the end joint of the robotic arm) or fixed to a preset position outside the robot as needed.

[0069] Generally, there is a preset eye-hand calibration relationship between the image sensor and the robot, so that the coordinate transformation relationship between the image sensor and the robot can be obtained.

[0070] Specifically, the image sensor can be various existing or future-developed image acquisition devices, such as: RGB cameras, depth cameras, point cloud cameras, video cameras, or various devices including such cameras or video cameras.

[0071] The force / tactile sensor 122 is disposed on the grasping actuator 111.

[0072] Wherein, the above-mentioned force / tactile sensor refers to: a force sensor and / or a tactile sensor.

[0073] Specifically, the force sensor can be, but is not limited to, a two-dimensional or multi-dimensional force sensor for measuring two-dimensional or multi-dimensional force data.

[0074] Specifically, the tactile sensor is a type of sensing device that can be placed at the end of actuators such as robotic arms and / or robots, etc., and is used to measure tactile information on the premise of cooperating with the end effector to achieve the grasping function.

[0075] To grasp objects of different shapes and softness, the contact surface between the tactile sensor and the object is usually flexible and has good resilience. The tactile information includes, but is not limited to: array multi-dimensional force information, surface deformation information, temperature information, texture information, etc.

[0076] The implementation of the tactile sensor includes a flexible contact surface, a sensing circuit, a computing device, and a tactile signal parsing algorithm.

[0077] It should be noted that in the embodiments of the present application, the signals measured and output by the force / tactile sensor are collectively referred to as force data. In addition, for the convenience of understanding, the embodiments of the present application mainly take the tactile sensor as an example for description below.

[0078] Specifically, the force / tactile sensor is pre-calibrated, so that the pose conversion relationship between the force / tactile sensor and the end effector can be obtained.

[0079] Exemplarily, taking the humanoid robot as an example, the end of the dual robotic arms of the robot is equipped with a dexterous hand with tactile sensors, and force / tactile sensors are distributed on each finger and palm of the dexterous hand; or as Figure 1 As shown, a gripper 111 is provided at the end of the robotic arm of the robot 110, and force / tactile sensors are distributed on the contact surface between the gripper 111 and the object.

[0080] The joint encoder (not shown in the figure) is disposed on the moving joints of the robot (such as: robotic arm joints and finger joints) for collecting the motion amount data of the joints of the robot.

[0081] Controller

[0082] The controller 130 is communicatively connected to the robot 110 and the state data sensor 120 respectively in a wired or wireless manner. For the definition of the controller, refer to the description of the training method and / or the grasping method of the grasping strategy model in the following embodiments.

[0083] It should be noted that the above wireless connection methods may include but are not limited to 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.

[0084] The controller in the embodiments of the present application may be but is not limited to: a computer terminal (Personal Computer, PC); an industrial control computer terminal (Industrial Personal Computer, IPC); a mobile terminal; a server; a system including a terminal and a server, and implemented through the interaction between the terminal and the server; a programmable logic controller (Programmable Logic Controller, PLC); a field programmable gate array (Field-Programmable Gate Array, FPGA); a digital signal processor (Digital Signal Processer, DSP) or a microcontroller unit (Microcontroller unit, MCU). The controller generates program instructions according to a pre-fixed program in combination with the data output by the robot and the force / tactile sensor, etc. Exemplarily, it can be applied to a computer device as shown in Figure 8 shown.

[0085] The controller 130 described in the embodiments of the present application may be an independent controller, or may be wholly or partially integrated in the robot and the state data sensor, etc., and the present application does not make any limitations.

[0086] It should be noted that for the training of the grasping strategy model, the above robot and state data sensor may be real devices, or may be virtual devices in a virtual environment. When they are virtual devices, the controller may first complete the training of the grasping strategy model based on the virtual devices, and then transplant it to the real training environment.

[0087] As Figure 2 shown, Figure 2 is a system architecture diagram of an embodiment of the training sample acquisition system based on teleoperation of the present application.

[0088] In one embodiment, the present application provides a training sample acquisition system 200 based on teleoperation, and the system includes: a master end interactuator 210, a slave robot 220, a slave image sensor 230, a force / tactile sensor 240, and a controller 250.

[0089] Master end interactor

[0090] In one embodiment, the master end interactuator 210 is configured to collect motion data during an operator's execution of a target operation, and send the motion data to the controller 250.

[0091] The master end interactuator 210 may be, but is not limited to: a master end sensor, an actuator body provided with the master end sensor, and the like.

[0092] Specifically, the master end sensor may be any sensor capable of collecting motion data, such as: an IMU, an image sensor, a position encoder, a biochemical sensor (such as: an electromyogram slave sensor). Among them, the IMU is an inertial measurement unit, which is used to measure motion data such as acceleration and angular velocity related to the operator.

[0093] In one embodiment, the above master end interactuator may be directly fixed to a preset key part of the operator. The master end interactuator may also be preset at the execution end of the actuator body, and the motion of the actuator body is controlled based on the operator's subjective operation, so as to collect the operator's motion data through the master end sensor provided on the actuator body.

[0094] In another embodiment, the master end interactuator may also be pre-set on the actuator body (such as: a master robot, a wearable device (such as: an exoskeleton or a data glove)), and the motion of the actuator body is controlled based on the operator's subjective operation, so as to collect the operator's motion data through the master end sensor provided on the actuator body.

[0095] In one embodiment, there is a preset calibration relationship between the master end interactuator and the operator, so that the motion condition of the operator can be directly or indirectly reflected based on the motion data collected by the master end interactuator.

[0096] Exemplarily, taking an arm exoskeleton as an example, an actuator body is composed of multiple connecting rods and the like, and an IMU can be set at the position corresponding to the arm joint of the actuator body. The arm exoskeleton is worn on the operator's arm, so that the motion data of the corresponding joint can be collected by the IMU during the motion of the operator's arm.

[0097] It should be noted that the operator described in the embodiments of the present application is not limited to a human, and may also be other living bodies according to needs. For the convenience of understanding, the embodiments of the present application mainly take the operator as a human as an example for description.

[0098] Slave robot

[0099] Specifically, the slave robot can be, but is not limited to: a humanoid robot; or a robotic arm connected in series or in parallel (such as: a Delta robot, a four-axis robot or a six-axis robot).

[0100] A force / tactile sensor 240 is arranged at the execution end of the slave robot 220.

[0101] In one embodiment, the force / tactile sensor can be arranged on the gripper or the dexterous hand at the execution end of the robot. Exemplarily, as Figure 2 shown, taking the humanoid robot as an example, the dexterous hands 221 equipped with tactile sensors are mounted at the ends of the two arms of the robot (i.e., the execution ends), and force / tactile sensors 240 are distributed on each finger joint and the palm of the dexterous hand.

[0102] Slave image sensor

[0103] In one embodiment, the slave image sensor 230 is used to collect the object image. The pose of the object can be obtained through the object image, and the relative pose between the operation part of the robot (such as: finger joint) and the object can be reflected, etc.

[0104] Generally, there is a preset calibration relationship between the slave image sensor 230 and the robot 220, so that the object image collected by the slave image sensor can be mutually mapped with the robot.

[0105] In one embodiment, taking the robotic arm as an example of the robot, the image sensor and the robotic arm can be pre-calibrated by the eye-in-hand calibration method.

[0106] Controller

[0107] The controller 250 is communicatively connected to the master end interaction device 210, the slave robot 220, the slave image sensor 230, etc. by wired or wireless means.

[0108] For the definition of the controller, refer to the description of the training method of the grasping strategy model and / or the grasping method in the following embodiments, and details are not described herein again.

[0109] In another embodiment, the training sample acquisition system based on teleoperation further includes: Demonstrator 240.

[0110] The controller is further configured to obtain the object image collected by the slave image sensor; generate a spatial virtual image of the object based on the object image, and send the virtual image to the demonstrator to display the spatial virtual image of the object to the operator through the demonstrator.

[0111] The embodiment of the present application adopts a remote operation-based training sample acquisition system to use operation text information and corresponding image data and force data information as training sample samples of the robot operation availability definition model, thereby obtaining more comprehensive and accurate training samples.

[0112] It should be noted that the training method of the grasping strategy model provided in the embodiment of the present application is generally Figure 1 The controller 130 of the grasping system described in the above is executed, and accordingly, the device for training the grasping strategy model is generally arranged in the controller 130 of the grasping training system.

[0113] like Figure 3 and 4 As shown, Figure 3 It is a structural schematic diagram of an embodiment of a training device for a grasping strategy model of the present application; Figure 4 The flowchart of an embodiment of the training method of the grasping strategy model of the present application is shown in FIG. The present application embodiment provides a training method of the grasping strategy model, which may include the following method steps:

[0114] Step 310 obtains the operation text prompt information and the current grasping state data under the current grasping state; the current state data includes: the current image data output by the image sensor; the current force data output by the force / tactile sensor; the current grasping posture of the robot obtained based on the joint motion output by the robot's encoder.

[0115] Step 320 uses the operation text prompt information and at least the current image data as inputs of a preset expected grasping model, and outputs the current expected grasping posture.

[0116] Step 330 uses the operation text prompt information and the current grasping state data as inputs of the grasping strategy model, combines the current expected grasping posture, and iteratively trains the grasping strategy model based on a deep reinforcement learning method until the termination condition is met.

[0117] The embodiment of the present application uses multi-dimensional operation text prompt information and part of the current grasping state data as the input of the preset expected grasping model, and can output the current expected grasping posture (the preset expected grasping model is a pre-trained expected grasping model, that is, the current expected grasping posture is obtained based on a learning method); then the operation text prompt information and the current grasping state data are used as the input of the grasping strategy model, combined with the current expected grasping posture, the grasping strategy model is iteratively trained based on the deep reinforcement learning method. The present application uses a learning method to obtain the corresponding expected grasping posture, and based on this, calculates the corresponding reward function, which can optimize to obtain a more general grasping strategy model.

[0118] For ease of understanding, the above method steps are further described in detail below.

[0119] Step 310 obtains the operation text prompt information and the current grasping state data under the current grasping state; the current grasping state data includes: the current image data output by the image sensor; the current force data output by the force / tactile sensor; the current grasping posture of the robot obtained based on the joint motion data output by the robot's encoder.

[0120] The operation text prompt information may be instructions for performing different operation tasks on the target object. For example, the operation text prompt information for a cup may be, but is not limited to, text information prompts such as pouring water, filling water, placing and / or washing. The embodiment of the present application may express the specific content of prompting the operator to perform the corresponding operation task in a more intuitive manner through the operation text prompt information.

[0121] The object information may be a description of the characteristics of the object, such as the type, shape and / or material. For example, the object information may be the name of the object type, such as a cup.

[0122] In one embodiment, the controller responds to a manually or automatically generated grasping start instruction, obtains operation text prompt information from a memory or server, etc. according to a preset address (the operation text prompt information can be obtained based on operator input or selection in preset options, or in any other required manner), and obtains current grasping status data output by a status data sensor at the current moment from a memory or server, etc. according to a preset address, and the current grasping status data is the status data at the current moment during the grasping attempt process.

[0123] In one embodiment, based on the above embodiments, the state data sensor 120 may include but is not limited to: an image sensor 121, a force / tactile sensor 122 disposed on a grasping actuator, and a joint encoder.

[0124] In one embodiment, the controller obtains current image data output by the image sensor from the memory or the server according to a preset address;

[0125] In one embodiment, the controller obtains current force data on the force distribution of the gripping surface of the gripping actuator output by the force / tactile sensor from the memory or the server according to a preset address; and

[0126] In one embodiment, the controller obtains the current joint motion data output by the joint encoder from the memory or the server according to the preset address, and the current grasping posture of the robot can be obtained based on the joint motion data.

[0127] It should be noted that the current grasping posture of the robot may include but is not limited to: the position information of the end joint of the robot arm and the posture information of the grasping actuator. For example, the end posture of the robot arm can be obtained based on the movement of each joint and the kinematic equation; the posture information of the grasping actuator such as the contact point or angle of the finger relative to the target object can be obtained based on the movement of each joint of the finger.

[0128] By adopting the above-mentioned state data sensor, the embodiment of the present application can obtain multi-dimensional state data from three aspects of vision, touch and robotics, thereby improving the versatility of subsequent models.

[0129] It should be noted that the "current crawling status data" described in the embodiments of the present application is not limited to the various data listed in the above embodiments, and may also include other required status data as needed, all of which fall within the scope of protection of this application.

[0130] Step 320 uses the operation text prompt information and at least the current image data as inputs of a preset expected grasping model, and outputs the current expected grasping posture.

[0131] In one embodiment, the controller obtains a pre-trained preset expected grasping model from a memory or a server according to a preset address; and uses the operation text prompt information and at least the current image data as inputs of the preset expected grasping model, thereby outputting the current expected grasping posture.

[0132] In one embodiment, the above-mentioned current expected grasping posture may include, but is not limited to: expected force data in the grasping state and expected robot posture data in the grasping state.

[0133] It should be noted that the above-mentioned “current expected grasping posture of the robot” may include, but is not limited to, expected posture information of the end of the robot and expected posture data of the grasping actuator when grasping an object.

[0134] The training process of the preset desired grasping model will be further described in detail in the following embodiments.

[0135] It should be noted that in addition to the operation text prompt information and the current image data, other current state data, such as the current force data and the current grasping posture of the robot, can also be used as the input of the desired grasping posture model together with the operation text prompt information and the current image data as needed. The operation text prompt information of the embodiment of the present application is mainly to be able to more accurately describe the grasping target. The image data is mainly to enable the robot to understand the scene, including information such as the grasping target object, position, obstacles, etc., so as to better complete the grasping of the target object.

[0136] Step 330 uses the operation text prompt information and the current grasping state data as the input of the grasping policy model, combines the current desired grasping posture, and iteratively trains the grasping policy model based on the deep reinforcement learning method until the termination condition is met.

[0137] It should be noted that the grasping policy model described in the embodiments of the present application can adopt various existing or future-developed network models, such as: Feed-Forward Networks, RNN, LSTM, Transformer, GNN, GAN, AE, Convolutional Neural Network (CNN). Common CNN models may include, but are not limited to: LeNet, AlexNet, ZFNet, VGG, GoogLeNet, Residual Net, DenseNet, R-CNN, SPP-NET, Fast-RCNN, Faster-RCNN, FCN, Mask-RCNN, YOLO, SSD, GCN.

[0138] In one embodiment, step 330 may include the following method steps:

[0139] Step 331 uses the operation text prompt information and the current grasping state data as the input of the grasping policy model, and outputs the next moment's grasping posture of the robot.

[0140] In one embodiment, the next moment's grasping posture includes, but is not limited to: the end pose information of the robot at the next moment and the pose data of the grasping actuator at the next moment.

[0141] Step 332 calculates the pose difference between the next moment's grasping posture and the current desired grasping posture; calculates the force difference between the current force data and the desired force data; and calculates the corresponding reward function according to the pose difference and the force difference.

[0142] The embodiments of the present application jointly calculate the reward function by combining the pose difference and the force difference, so that the next moment's grasping posture output by the grasping policy model data is close to the current desired grasping posture, and the grasping posture at the next moment meets the requirement that the current force data and the desired force data are as evenly distributed as possible. Thus, the next moment's grasping posture can be limited from multiple dimensions, thereby improving the generality of the grasping policy model.

[0143] Specifically, the reward function can be arbitrarily set as needed. For example, the reward function can be defined such that the smaller the difference, the greater the reward.

[0144] Step 333 optimizes the grasping policy model based on the reward function.

[0145] Step 334 generates a grasping instruction based on the next moment's grasping posture to instruct the robot to perform a grasping motion.

[0146] In one embodiment, the controller generates a grasping instruction based on the grasping pose at the next moment to instruct the robot to perform a grasping motion, so that the grasping pose at the next moment can be applied to a real or virtual physical environment, and then the next frame of state data output by the state sensor is obtained, and the next frame of state data is used as the current state data in step 331 above.

[0147] Step 335 repeats the method steps of steps 310 to 334 above until a termination condition is met.

[0148] Specifically, the preset termination condition can be the maximum number of iterations or the above difference reaches a preset threshold.

[0149] In the embodiment of the present application, when using reinforcement learning to learn an operation task (for example: taking grasping a target object as an example), a reward function needs to be provided. This reward function needs to give corresponding evaluations to different grasping strategies tried by the robot. The quality of the reward function will affect the finally learned grasping effect. This reward function needs to define the best grasping pose of the finger and the object, and the definition of this best pose should be given according to different operation scenarios and objects, and factors such as the contact points, forces, and stability between the finger and the object also need to be considered. However, designing such a reward function is very difficult and complex, so this solution proposes to calculate the corresponding reward in a learning manner, thereby reducing the model training difficulty and improving the generality of the grasping strategy model.

[0150] In the embodiment of the present application, taking multi-dimensional operation text prompt information and partial current grasping state data as the input of a preset expected grasping model, the current expected grasping pose can be output (the preset expected grasping model is a pre-trained expected grasping model, that is, the current expected grasping pose is obtained based on a learning method); then taking the operation text prompt information and the current grasping state data as the input of the grasping strategy model, combining the current expected grasping pose, and iteratively training the grasping strategy model based on the deep reinforcement learning method. The present application obtains the corresponding expected grasping pose in a learning manner, and calculates the corresponding reward function based on this, and can optimize to obtain a more general grasping strategy model.

[0151] In one embodiment, before step 320 takes the operation text prompt information and at least the current image data as the input of the preset expected grasping model, the following steps can also be included to train the expected grasping model.

[0152] Step 340 obtains training samples; the training samples include: operation text prompt information and target image data in the grasping state; using the training grasping state data corresponding to the operation text prompt information as the annotation of the training samples; the training grasping state data includes the training force data in the grasping state and the training robot grasping pose in the grasping state.

[0153] In one embodiment, the controller obtains the pre-generated training samples from the memory according to the preset address. For the generation method of the training samples, refer to the following embodiments.

[0154] Step 350 trains the expected grasping model based on the training samples and the corresponding annotations to obtain a preset expected grasping model.

[0155] In one embodiment, step 350 may include the following method steps:

[0156] Step 351 takes the current training sample as the input of the expected grasping model and outputs the expected grasping pose;

[0157] Step 352 calculates the difference between the expected grasping pose and the annotation corresponding to the current training sample;

[0158] Step 353 iteratively updates the parameters of the expected grasping model based on the difference until the preset training termination condition is met.

[0159] Specifically, the above training termination condition can be set arbitrarily according to needs. For example, it can be that the difference meets the preset requirement range, or meets the preset number of iterations, etc.

[0160] The above expected grasping model can be various existing or future-developed models capable of grasping pose recognition. For example: Feed-Forward Networks, RNN, LSTM, Transformer, GNN, GAN, AE, Convolutional Neural Network (CNN). Common CNN models may include but are not limited to: LeNet, AlexNet, ZFNet, VGG, GoogLeNet, Residual Net, DenseNet, R-CNN, SPP-NET, Fast-RCNN, Faster-RCNN, FCN, Mask-RCNN, YOLO, SSD, GCN.

[0161] In the embodiment of the present application, the expected grasping model is trained by the method of supervised learning, and the expected grasping pose is obtained based on the learning method. Therefore, the expected grasping pose of the present application has universality. In addition, training the expected grasping model based on the method of supervised learning is simpler than reinforcement learning training.

[0162] In one embodiment, before step 340, the following method steps may also be included to generate the above training samples.

[0163] Step 360 generates the corresponding annotation data and the target image data in the grasping state based on the operation text prompt information.

[0164] It should be noted that the "generation of corresponding expected grasping postures based on operation text prompt information" can be generated based on various existing or future-developed method steps. For example: in the way of teleoperation, based on the operation text prompt information, the operator is instructed to grasp an object and then release the object, so as to obtain the grasping posture, which will be further described in detail in the following implementation examples; or the temperature distribution on the surface of the object recorded through a thermal imager is obtained, and the grasping posture is obtained based on the temperature distribution. Continuing with the operation text prompt information "pour water + cup" as an example, for the operator's corresponding task of pouring water + the target cup, the operator will grasp the handle of the cup. At this time, the temperature of the operator's hand can be transmitted to the object. Therefore, by recording the distribution of this temperature on the surface of the object, it can be known how the object is grasped by the operator during the operation, and thus the grasping posture of the operator's hand can be correspondingly obtained, and this grasping posture is used as the grasping posture of the robot.

[0165] In one embodiment, step 360 may include the following method steps. The following method steps are generally executed by the controller 230 of the training sample application system based on teleoperation described in the above embodiments Figure 2 in the system.

[0166] Step 361 sends operation text prompt information to the operator so that the operator performs a grasping operation on the target object based on the operation text prompt information.

[0167] Exemplarily, based on the operation text prompt information "pour water", the operator can subjectively perform the action of grasping the water cup by visually observing the specific position of the water cup at the slave end with the naked eye, or by watching the virtual image of the water cup in the three-dimensional space demonstrated by the demonstrator. Here, in view of the purpose of achieving pouring water, therefore, the operator usually subjectively grabs the handle or the grip of the water cup, rather than the cup mouth; in another example, when the operation text prompt information is "wash the cup", usually the operator will grasp the cup mouth of the water cup.

[0168] Step 362 acquires the action data of the operator performing the target operation collected by the master-end interactors; generates a robot grasping instruction based on the action data to instruct the slave-end robot to follow and grasp the target object.

[0169] Step 363, after the slave-end robot follows the completion of the target operation, acquires the grasping posture in the state of completed grasping and the state data output based on the state sensors at the slave end.

[0170] In one embodiment, the grasping posture of the robot when grasping the object can be recognized based on the target image data in the state of completed grasping output by the slave-end image sensor.

[0171] The above-mentioned status data may include but are not limited to: target image data in the grasping state output by the image sensor (the target image data and the operation text prompt information may be used as training samples of the expected grasping model to train the expected grasping model later); training force data in the grasping state output by the force / tactile sensor; and the training grasping posture of the robot in the grasping state obtained based on the joint encoder. For the training grasping posture data of the robot, please refer to the description in the above embodiment and will not be repeated here.

[0172] Step 364 uses the operation text prompt information, the target image data in the completed grasping state, and the corresponding annotation data (training force data and the training grasping posture of the robot in the grasping state) as training samples for the robot.

[0173] The embodiment of the present application can quickly obtain operation text prompt information, as well as the corresponding robot's grasping posture when grasping an object and image data in the above grasping state through a remote operation-based method, with the operation text prompt information as input and the grasping posture and image data in the above grasping state as annotations.

[0174] like Figure 4 As shown, Figure 4 It is a flowchart of an embodiment of the crawling method of the present application.

[0175] Based on the grasping strategy model trained by the grasping strategy model training method described in the above embodiment, the embodiment of the present application further provides a grasping method, which may include the following method steps:

[0176] It should be noted that the grabbing method provided in the embodiments of the present application is generally Figure 1 The target object grasping system described in the embodiment is executed by the controller 130 , and accordingly, the device for the grasping method is generally arranged in the controller 130 .

[0177] Step 410 obtains the operation text prompt information and the current grasping state data; the current grasping state data includes: the current image data output by the image sensor; the current force data output by the force / tactile sensor; the current grasping posture of the robot obtained based on the joint motion output by the robot's encoder.

[0178] For the description of the current grabbing status data, please refer to the above embodiment, which will not be repeated here.

[0179] Step 420 obtains the crawling strategy model.

[0180] Step 430 uses the operation text prompt information and the grasping state data as inputs of the grasping strategy model, and outputs the grasping posture and image data under the grasping state.

[0181] Among them, the grasping posture includes the force data in the grasping state and the grasping posture data of the robot.

[0182] Step 440 generates a grasping instruction based on the grasping posture to instruct the robot to grasp the target object through the grasping instruction.

[0183] Step 450 repeats the method steps of Step 410 to Step 440 until the grasping of the target object is completed.

[0184] In one embodiment, it is possible to determine whether the target object has been grasped based on a reward function. Specifically, the "operation text prompt information" and the "image data in the grasping state" in Step 410 can be used as the input of the above-mentioned preset expected grasping model, and the expected grasping posture is output; the difference between the expected grasping posture and the grasping posture output in Step 430 is obtained; a reward function is obtained based on the difference, and it is determined whether the grasping is successful according to the value of the reward function. For example, when the reward function is less than a preset threshold, it is regarded as a successful grasping. In addition to completing the training of the grasping strategy model based on the expected grasping model, the embodiment of the present application can also be used to determine whether the target object has been successfully grasped during the actual grasping operation of the target object based on this, thus further improving the versatility of the grasping operation. In addition, it is also possible to determine whether the target object has been successfully grasped based on methods such as the collected image or force data, which all fall within the protection scope of the present application.

[0185] In the embodiment of the present application, based on the operation text prompt information and the image data at the current moment captured and output by the image sensor as the input of the grasping strategy model, the grasping posture at the current moment is output. The method steps of Step 410 to Step 440 are repeatedly executed until the grasping of the target object is completed. Then, the current grasping posture obtained when the grasping of the target object is completed is used as the final grasping posture, and this grasping posture is sent to the robot to instruct the robot to complete the grasping of the target object.

[0186] The embodiment of the present application uses multi-dimensional operation text prompt information and partial current grasping state data as the input of the preset expected grasping model, and can output the current expected grasping posture (the preset expected grasping model is a pre-trained expected grasping model, that is, the current expected grasping posture is obtained based on a learning method); then, using the operation text prompt information and the current grasping state data as the input of the grasping strategy model, combined with the current expected grasping posture, the grasping strategy model is iteratively trained based on the deep reinforcement learning method. The present application obtains the corresponding expected grasping posture in a learning manner, and calculates the corresponding reward function based on this, and can optimize to obtain a more general grasping strategy model, thereby improving the versatility of the target object grasping method.

[0187] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0188] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0189] Further references Figure 6 , as a response to the above Figure 3 The present application provides an embodiment of a training device for a grasping strategy model, and the device embodiment is similar to Figure 3 Corresponding to the training method embodiment of the grasping strategy model shown, the device can be specifically applied to various controllers.

[0190] like Figure 6 As shown, the training device 300 of the grasping strategy model of this embodiment includes:

[0191] The data acquisition module 310 is used to acquire the operation text prompt information and the current grasping state data under the current grasping state; the current grasping state data includes: the current image data output by the image sensor; the current force data output by the force / tactile sensor; the current grasping posture of the robot obtained based on the joint motion output by the robot encoder;

[0192] An expectation generation module 320, configured to use the operation text prompt information and at least the current image data as inputs of a preset expected grasping model, and output a current expected grasping posture;

[0193] The model training module 330 is configured to use the operation text prompt information and the current grasping state data as the input of the grasping policy model, and in combination with the current expected grasping posture, iteratively train the grasping policy model based on the deep reinforcement learning method until the termination condition is met.

[0194] For details, please refer to Figure 8 , to solve the above technical problems, an embodiment of the present application further provides a controller (taking the computer device 6 as an example).

[0195] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 6 with components 61 - 63 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field - programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0196] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server, and other computing devices. The computer device can perform human - machine interaction with the user through a keyboard, a mouse, a remote control, a touchpad, a voice control device, or other means.

[0197] The memory 61 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 6. Of course, the memory 61 may also include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as the program code for training the grasping strategy model and / or the grasping method. In addition, the memory 61 may also be used to temporarily store various types of data that have been output or will be output.

[0198] In some embodiments, the processor 62 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the program code stored in the memory 61 or process data, such as running the program code for training the grasping strategy model and / or the grasping method.

[0199] The network interface 63 may include a wireless network interface or a wired network interface, and this network interface 63 is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0200] This application also provides another implementation manner, that is, to provide a computer-readable storage medium storing a program for training the grasping strategy model and / or a grasping program, and the program for training the grasping strategy model and / or the grasping program can be executed by at least one processor, so that the at least one processor executes the steps of the method for training the grasping strategy model and / or the grasping method as described above.

[0201] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0202] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The preferred embodiments of the present application are given in the drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the scope of the patent protection of the present application.

Claims

1. A training method for a grasping strategy model, applied to a grasping training system, comprising: A robot, a state data acquisition sensor and a controller; the robot comprises: a mechanical arm and a grasping actuator arranged at the grasping execution end of the mechanical arm; the state data acquisition sensor comprises an image sensor, a force / tactile sensor arranged at the grasping actuator and a joint encoder; the force / tactile sensor is arranged at the grasping actuator, characterized in that the method comprises the following steps: Acquire operation text prompt information and current grasping state data under the current grasping state; the current grasping state data includes: current image data output by the image sensor; current force data output by the force / tactile sensor; and the current grasping posture of the robot obtained based on the joint motion output by the robot's encoder; Using the operation text prompt information and at least the current image data as inputs of a preset expected grasping model, and outputting a current expected grasping posture; The operation text prompt information and the current grasping state data are used as inputs of the grasping strategy model, and combined with the current expected grasping posture, the grasping strategy model is iteratively trained based on a deep reinforcement learning method until a termination condition is met.

2. The training method for a grasping strategy model according to claim 1, wherein The current expected grasping posture includes: expected force data in the grasping state and the expected grasping posture of the robot in the grasping state.

3. The training method for a grasping strategy model according to claim 2, wherein The method of taking the operation text prompt information and the current grasping state data as inputs of the grasping strategy model, combining the current expected grasping posture, and iteratively training the grasping strategy model based on a deep reinforcement learning method until a termination condition is met includes the following steps: The operation text prompt information and the current grasping state data are used as inputs of the grasping strategy model, and the grasping posture of the robot at the next moment is obtained as output; Obtaining a posture difference between the grasping posture at the next moment and the current expected grasping posture; obtaining a force difference between the current force data and the expected force data; and calculating a corresponding reward function according to the posture difference and the force difference; Based on the reward function, optimizing the grasping strategy model; Generate a grasping instruction based on the grasping posture at the next moment, so as to instruct the robot to perform a grasping motion through the grasping instruction; Repeat the steps of obtaining the operation text prompt information and the current grasping state data under the current grasping state; using the operation text prompt information and at least the current image data as inputs of a preset desired grasping model; using the operation text prompt information and the current grasping state data as inputs of the grasping strategy model, and outputting the grasping posture of the robot at the next moment; Obtain the posture difference between the grasping posture at the next moment and the current expected grasping posture; obtain the force difference between the current force data and the expected force data; calculate the corresponding reward function according to the posture difference and the force difference; optimize the grasping strategy model based on the reward function; generate a grasping instruction based on the grasping posture at the next moment, and instruct the robot to execute the grasping motion through the grasping instruction until the termination condition is met.

4. The training method for a grasping strategy model according to claim 3, wherein The current grasping posture includes: the end position information of the robot at the current moment and the posture data of the grasping actuator at the current moment; and / or, The next moment grasping pose includes: the end pose information of the robot at the next moment and the pose data of the grasping actuator at the next moment; The desired grasping pose includes: the desired end pose information of the robot at the current moment and the desired pose data of the grasping actuator at the current moment.

5. The training method for a grasping strategy model according to claim 1 or 2, wherein Before taking the operation text prompt information and at least the current image data as the input of the preset desired grasping model and outputting the current desired grasping pose, the following steps are further included: Obtain training samples; the training samples include: the operation text prompt information and the training target image data in the corresponding grasping state; use the training grasping pose data corresponding to the operation text prompt information as the annotation of the training samples; the training grasping pose data includes the training force data in the grasping state and the training grasping pose of the robot in the grasping state; Based on the training samples and the corresponding annotations, train the desired grasping model to obtain the preset desired grasping model.

6. The training method for a grasping strategy model according to claim 5, wherein Before obtaining the training samples, the following steps are further included: Generate the corresponding annotations and the training target image data in the grasping state based on the operation text prompt information.

7. A grasping method of the grasping strategy model obtained by the training method for a grasping strategy model according to any one of claims 1 to 6, wherein The method includes the following steps: Obtain the operation text prompt information and the current grasping state data; the current grasping state data includes: the current image data output by the image sensor; the current force data output by the force / tactile sensor; the current grasping pose of the robot obtained by calculating the joint movement amount output by the robot's encoder; Obtain the grasping strategy model; Take the operation text prompt information and the current grasping pose as the input of the grasping strategy model, and output the next moment grasping pose of the robot; Generate a grasping instruction based on the next moment grasping pose to instruct the robot to grasp the target object through the grasping instruction; Repeat the above steps until the grasping of the target object is completed.

8. A training device for a grasping strategy model, wherein The device includes: A data acquisition module, configured to acquire operation text prompt information and current grasping state data in the current grasping state; the current grasping state data includes: the current image data output by the image sensor; the current force data output by the force / tactile sensor; the current grasping pose of the robot obtained by calculating the joint movement amount output by the robot's encoder; A desired generation module, configured to take the operation text prompt information and at least the current image data as the input of the preset desired grasping model, and output the current desired grasping pose; A model training module, configured to take the operation text prompt information and the current grasping state data as the input of the grasping strategy model, and based on the current desired grasping pose, iteratively train the grasping strategy model based on the deep reinforcement learning method until the termination condition is met.

9. A grasping device of the grasping strategy model trained by the training device of the grasping strategy model according to claim 8, characterized in that, The device includes: An information acquisition module, which is used for a text acquisition module to acquire the operation text prompt information and the current grasping state data; the current grasping state data includes: the current image data output by an image sensor; the current force data output by a force / tactile sensor; the current grasping posture of the robot obtained by calculating the joint movement amount based on the output of the robot's encoder. A model acquisition module, which is used to acquire the grasping strategy model. A posture generation module, which is used to use the operation text prompt information and the current grasping state data as the input of the grasping strategy model, and output the next moment's grasping posture of the robot. An instruction generation module, which is used to generate a grasping instruction based on the next moment's grasping posture, so as to instruct the robot to grasp the target object through the grasping instruction. A step repetition module, which is used to repeat the above steps until the grasping of the target object is completed.

10. A grasping system, characterized in that, The system includes: a robot, a state data acquisition sensor, and a controller; the robot includes: a robotic arm and a grasping actuator arranged at the grasping execution end of the robotic arm. The controller is respectively communicatively connected to the robot and the state data acquisition sensor. The controller is used to implement the training method of the grasping strategy model described in any one of claims 1 to 6 and / or the grasping method described in claim 7.