A vision-based dynamic human-machine-object handover method and system

Through a dynamic human-machine object handover method based on vision, a collision-free grabbing configuration is generated using the hand and object point cloud, the problems of low safety and low success rate of human-machine object handover in the prior art are solved, and efficient and safe dynamic handover is achieved.

CN119501954BActive Publication Date: 2025-06-06UNIV OF SCI & TECH OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510080888.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing human-machine object handover process has low safety and low success rate. The robot is prone to collision with the handover's handover's handover and cannot dynamically adapt to the movement changes of the object.

Method used

A dynamic human-computer object handover method based on vision is adopted to collect image data in real time, acquire hand point clouds and object point clouds, generate collision-free grab configurations, and generate new grab configurations according to the dynamic changes of objects.

Benefits of technology

It improves the safety and success rate of human-machine object hand transfer, avoids the collision between the robot and the hand transferor's hand, and can dynamically adapt to the movement changes of the object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119501954B_ABST
    Figure CN119501954B_ABST
Patent Text Reader

Abstract

The present invention discloses a vision-based dynamic human-machine object handover method and system, comprising the following steps: S1, real-time acquisition of image data, obtaining the starting object point cloud and the starting hand point cloud according to the first frame image; S2, obtaining a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud; S3, configuring the robot's grasping action parameters according to the collision-free grasping configuration; S4, obtaining the current frame image and the previous frame image, judging whether the object moves according to the current frame image and the previous frame image; if so, generating a new grasping configuration according to the current frame image and the previous frame image, configuring the robot's grasping action parameters according to the new grasping configuration; if not, maintaining the current grasping action parameters of the robot. The dynamic human-machine object handover method has high safety and high grasping success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent interaction technology for manipulators, and in particular to a vision-based dynamic human-machine-object handover method and system. Background Art

[0002] As the social problem of population aging becomes increasingly serious, there will be a large number of work scenarios in the future society that require human-machine collaboration and interaction. Among them, human-machine object handover is one of the most common operations. Robots use multi-finger dexterous hands to achieve fast and smooth operations in dynamic scenes, which is of great significance for promoting robots to truly integrate into human daily production and life scenes and provide services in multiple tasks.

[0003] However, in the traditional human-machine object handover process, only the handover person's hand or only the object is identified, which makes it easy for the robot to collide with the handover person's hand during the human-machine object handover process, and the robot may hurt the handover person's hand, resulting in low safety of the existing human-machine object handover. In addition, the existing human-machine object handover usually adopts open-loop control, requiring the handover person to remain stationary when the robot performs the handover task. When the handover person adjusts his posture and causes the object to move, the robot cannot dynamically adapt according to the movement of the object, resulting in a low success rate of the existing human-machine object handover. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a vision-based dynamic human-machine-object handover method and system to solve the problems of low safety and low success rate in human-machine-object handover.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: a vision-based dynamic human-machine object handover method, comprising the following steps:

[0006] S1, collecting image data in real time, obtaining the starting object point cloud and the starting hand point cloud according to the first frame image;

[0007] S2. Acquire a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud;

[0008] S3, configuring the robot's grasping action parameters according to the collision-free grasping configuration;

[0009] S4. Obtain the current frame image and the previous frame image, and determine whether the object has moved based on the current frame image and the previous frame image; if so, generate a new grasping configuration based on the current frame image and the previous frame image, and configure the grasping action parameters of the robot based on the new grasping configuration; if not, maintain the current grasping action parameters of the robot.

[0010] Furthermore, in the step S1, obtaining the starting object point cloud and the starting hand point cloud according to the first frame image specifically includes:

[0011] The first frame image includes an RGB image and a depth image; a hand bounding box and an object bounding box are identified according to the RGB image; a segmentation model obtains a hand segmentation mask and an object segmentation mask according to the hand bounding box, the object bounding box and the RGB image, projects the hand segmentation mask onto the depth image to obtain the starting hand point cloud, and projects the object segmentation mask onto the depth image to obtain the starting object point cloud.

[0012] Furthermore, in the step S2, obtaining a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud includes:

[0013] A plurality of candidate grasping configurations are generated according to the starting object point cloud, grasping success rates of the plurality of candidate grasping configurations are evaluated respectively, the plurality of candidate grasping configurations are sorted from large to small according to the grasping success rates, and collision detection is performed in sequence until a collision-free grasping configuration is screened out.

[0014] Furthermore, the method of sorting the plurality of to-be-selected grabbing configurations from large to small according to the grabbing success rate and performing collision detection in sequence includes:

[0015] A robot model is constructed according to the selected grasping configuration, and the distances between all points in the starting hand point cloud and the robot model are calculated respectively to determine whether the distances between all points in the starting hand point cloud and the robot model are less than or equal to a preset distance value; if so, the selected grasping configuration is a collision-free grasping configuration; if not, the selected grasping configuration is a collision-prone grasping configuration.

[0016] Furthermore, the collision-free grasping configuration includes orientation parameters in a camera coordinate system, position parameters in a camera coordinate system, and joint angle parameters.

[0017] Furthermore, in the step S2, the grasping action parameters of the robot are configured according to the collision-free grasping configuration, including: transforming the orientation parameters in the camera coordinate system into orientation parameters in the robot coordinate system, transforming the position parameters in the camera coordinate system into position parameters in the robot coordinate system, and closing the robot's manipulator according to the joint angle parameters.

[0018] Furthermore, in step S4, judging whether the object moves according to the current frame image and the previous frame image includes:

[0019] A point cloud of the object of the previous frame is obtained according to the previous frame image, and a point cloud of the object of the current frame is obtained according to the current frame image, and it is determined whether the distance between the coordinate center of the point cloud of the object of the previous frame and the coordinate center of the point cloud of the object of the current frame is less than a preset distance threshold; if so, the object has not moved; if not, the object has moved.

[0020] Furthermore, in step S4, a new grasping configuration is generated according to the current frame image and the previous frame image, including: obtaining a transformation matrix between the current frame object point cloud and the previous frame object point cloud; and transforming the grasping configuration corresponding to the current grasping action parameters of the robot according to the transformation matrix to obtain the new grasping configuration.

[0021] Furthermore, the method further includes step S5: determining whether the number of times that the continuous object has not moved is equal to a preset number; if so, terminating the process; if not, returning to execute step S4.

[0022] Another object of the present invention is to provide a dynamic human-machine-object handover system based on visual recognition, which is used to implement the above-mentioned dynamic human-machine-object handover method. The dynamic human-machine-object handover system includes an input unit, a control unit and a robot. The input unit and the robot are both connected to the control unit. The input unit is used to collect image data in real time. The control unit includes a scene understanding module, a grasping planning module and a dynamic grasping module. The grasping planning module and the dynamic grasping module are both connected to the scene understanding module, and the grasping planning module is connected to the dynamic grasping module; the scene understanding module is used to obtain a hand point cloud and an object point cloud according to the image data; the grasping planning module is used to obtain a collision-free grasping configuration according to the hand point cloud of the first frame image of the image data and the object point cloud of the first frame image of the image data; the dynamic grasping module is used to analyze whether the object moves according to two consecutive frames of images in the image data, and generate a new grasping configuration when the object moves.

[0023] The beneficial effects of the present invention are as follows: the dynamic human-machine-object handover method provided by the present invention can segment image data into hand point cloud and object point cloud, and a collision-free grasping configuration can be obtained by combining the hand point cloud and the object point cloud, so as to avoid collision between the robot and the manipulator and the hand of the handover person during the process of grasping the object, thereby improving the grasping effect and safety of the human-machine-object handover; in addition, the dynamic human-machine-object handover method does not require the handover person to remain still during the human-machine-object handover process, and the dynamic human-machine-object handover method can adapt to the dynamic changes of the object during the handover process, dynamically adapt according to the dynamic changes of the object and generate a new grasping configuration accordingly, thereby improving the grasping success rate of the dynamic human-machine-object handover method. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A flowchart of the steps of the dynamic human-machine-object handover method of the present invention;

[0025] Figure 2 is a block diagram of the dynamic human-machine-object handover system of the present invention;

[0026] Figure 3 The figure is a flow chart of the dynamic human-machine-object handover system. DETAILED DESCRIPTION

[0027] In order to explain the technical content, achieved objectives and effects of the present invention in detail, the following is an explanation in combination with the implementation modes and the accompanying drawings.

[0028] Please refer to Figures 1 to 3 The present invention provides a vision-based dynamic human-machine object handover method, comprising the following steps:

[0029] S1, collecting image data in real time, obtaining the starting object point cloud and the starting hand point cloud according to the first frame image;

[0030] S2. Acquire a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud;

[0031] S3, configuring the robot's grasping action parameters according to the collision-free grasping configuration;

[0032] S4. Obtain the current frame image and the previous frame image, and determine whether the object has moved based on the current frame image and the previous frame image; if so, generate a new grasping configuration based on the current frame image and the previous frame image, and configure the grasping action parameters of the robot based on the new grasping configuration; if not, maintain the current grasping action parameters of the robot.

[0033] From the above description, it can be seen that the beneficial effects of the present invention are: the dynamic human-machine-object handover method provided by the present invention can segment the image data into hand point cloud and object point cloud, and the collision-free grasping configuration can be obtained by combining the hand point cloud and the object point cloud, so as to avoid the robot and the manipulator colliding with the hand of the handover person in the process of grasping the object, thereby improving the grasping effect of the human-machine-object handover; in addition, the dynamic human-machine-object handover method does not require the handover person to remain still during the human-machine-object handover process. The dynamic human-machine-object handover method can adapt to the dynamic changes of the object during the handover process, dynamically adapt according to the dynamic changes of the object and generate a new grasping configuration accordingly, thereby improving the grasping success rate of the dynamic human-machine-object handover method.

[0034] Furthermore, in the step S1, obtaining the starting object point cloud and the starting hand point cloud according to the first frame image specifically includes:

[0035] The first frame image includes an RGB image and a depth image; a hand bounding box and an object bounding box are identified according to the RGB image; a segmentation model obtains a hand segmentation mask and an object segmentation mask according to the hand bounding box, the object bounding box and the RGB image, projects the hand segmentation mask onto the depth image to obtain the starting hand point cloud, and projects the object segmentation mask onto the depth image to obtain the starting object point cloud.

[0036] From the above description, it can be seen that the dynamic human-machine object handover method provided by the present invention can segment the image data into hand point cloud and object point cloud, and combine the hand point cloud and the object point cloud to obtain a collision-free grasping configuration to avoid the robot and the manipulator colliding with the hand of the handover person during the process of grasping the object.

[0037] Furthermore, in the step S2, obtaining a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud includes:

[0038] A plurality of candidate grasping configurations are generated according to the starting object point cloud, grasping success rates of the plurality of candidate grasping configurations are evaluated respectively, the plurality of candidate grasping configurations are sorted from large to small according to the grasping success rates, and collision detection is performed in sequence until a collision-free grasping configuration is screened out.

[0039] From the above description, it can be seen that the grasping generation network in the present invention can generate a series of grasping configurations according to the starting object point cloud in the first frame image, and evaluate the success rate of each grasping configuration through the grasping evaluation network, arrange the grasping configurations according to the success rate, and then perform collision detection on the sorted grasping configurations in turn until a collision-free grasping configuration is screened out. Therefore, the dynamic human-machine object handover method provided by the present invention can generate a collision-free grasping configuration with a high grasping success rate and no collision between the manipulator and the handover person's hand.

[0040] Furthermore, the method of sorting the plurality of to-be-selected grabbing configurations from large to small according to the grabbing success rate and performing collision detection in sequence includes:

[0041] A robot model is constructed according to the selected grasping configuration, and the distances between all points in the starting hand point cloud and the robot model are calculated respectively to determine whether the distances between all points in the starting hand point cloud and the robot model are less than or equal to a preset distance value; if so, the selected grasping configuration is a collision-free grasping configuration; if not, the selected grasping configuration is a collision-prone grasping configuration.

[0042] From the above description, it can be seen that the dynamic human-machine object handover method provided by the present invention can perform a collision test on the selected grasping configuration to detect whether the manipulator will collide with the hand of the handover person under the selected grasping configuration, and can screen out a collision-free grasping configuration with a high grasping success rate and in which the manipulator will not collide with the hand from multiple selected grasping configuration modules.

[0043] Furthermore, the collision-free grasping configuration includes orientation parameters in a camera coordinate system, position parameters in a camera coordinate system, and joint angle parameters.

[0044] Furthermore, in the step S2, the grasping action parameters of the robot are configured according to the collision-free grasping configuration, including: transforming the orientation parameters in the camera coordinate system into orientation parameters in the robot coordinate system, transforming the position parameters in the camera coordinate system into position parameters in the robot coordinate system, and closing the robot's manipulator according to the joint angle parameters.

[0045] In some preferred embodiments of the present invention, the new grasping configuration includes new orientation parameters in the camera coordinate system and new position parameters in the camera coordinate system. Since the shape of the object will not change, only the orientation parameters and the position parameters need to be updated when the object moves, so as to improve the efficiency of the dynamic human-machine object handover method in generating a new grasping configuration.

[0046] Furthermore, in step S4, judging whether the object moves according to the current frame image and the previous frame image includes:

[0047] A point cloud of the object of the previous frame is obtained according to the previous frame image, and a point cloud of the object of the current frame is obtained according to the current frame image, and it is determined whether the distance between the coordinate center of the point cloud of the object of the previous frame and the coordinate center of the point cloud of the object of the current frame is less than a preset distance threshold; if so, the object has not moved; if not, the object has moved.

[0048] From the above description, it can be seen that the dynamic human-machine-object handover method provided by the present invention analyzes whether the object moves based on two consecutive frames of images in the image data. The dynamic human-machine-object handover method can dynamically adapt to the moving object and can generate accurate new grasping configurations for the moving problem in real time.

[0049] Furthermore, in step S4, a new grasping configuration is generated according to the current frame image and the previous frame image, including: obtaining a transformation matrix between the current frame object point cloud and the previous frame object point cloud; and transforming the grasping configuration corresponding to the current grasping action parameters of the robot according to the transformation matrix to obtain the new grasping configuration.

[0050] From the above description, it can be seen that this method uses a transformation matrix to obtain a new grasping configuration, which can increase the speed of generating the new grasping configuration, thereby improving the dynamic adaptability reliability of the dynamic human-machine object handover method and the smoothness of the robot's movements, and can respond promptly to changes in the movement of objects.

[0051] Furthermore, the method further includes step S5: determining whether the number of times that the continuous object has not moved is equal to a preset number; if so, terminating the process; if not, returning to execute step S4.

[0052] From the above description, it can be seen that in the process of human-machine object handover, the handover person usually adjusts the hand movement at the beginning of the action to cause the object to move. After adjusting the hand movement, he can remain motionless for a certain period of time until the robot or manipulator completes the object handover action. Therefore, in this embodiment, when it is judged that the object has not moved for a preset number of consecutive times, it can be considered that the object is no longer moving, thereby reducing the frequency of dynamic adaptation according to the movement of the object or stopping dynamic adaptation according to the movement of the object.

[0053] Another object of the present invention is to provide a dynamic human-machine-object handover system based on visual recognition, which is used to implement the above-mentioned dynamic human-machine-object handover method. The dynamic human-machine-object handover system includes an input unit, a control unit and a robot. The input unit and the robot are both connected to the control unit. The input unit is used to collect image data in real time. The control unit includes a scene understanding module, a grasping planning module and a dynamic grasping module. The grasping planning module and the dynamic grasping module are both connected to the scene understanding module, and the grasping planning module is connected to the dynamic grasping module; the scene understanding module is used to obtain a hand point cloud and an object point cloud according to the image data; the grasping planning module is used to obtain a collision-free grasping configuration according to the hand point cloud of the first frame image of the image data and the object point cloud of the first frame image of the image data; the dynamic grasping module is used to analyze whether the object moves according to two consecutive frames of images in the image data, and generate a new grasping configuration when the object moves.

[0054] It can be seen from the above description that the dynamic human-machine-object handover system provided by the present invention has at least all the beneficial effects of the above dynamic human-machine-object handover method.

[0055] Embodiment 1

[0056] Please refer to Figures 1 to 3 The first embodiment of the present invention is to provide a vision-based dynamic human-machine object handover method, which includes the following steps:

[0057] S1, collecting image data in real time, obtaining the starting object point cloud and the starting hand point cloud according to the first frame image;

[0058] S2. Acquire a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud;

[0059] S3, configuring the robot's grasping action parameters according to the collision-free grasping configuration;

[0060] S4. Obtain the current frame image and the previous frame image, and determine whether the object has moved based on the current frame image and the previous frame image; if so, generate a new grasping configuration based on the current frame image and the previous frame image, and configure the grasping action parameters of the robot based on the new grasping configuration; if not, maintain the current grasping action parameters of the robot.

[0061] This dynamic human-machine object handover method can be widely used in various human-machine object handover scenarios, for example: in factory workshops, it can help workers hand over parts and tools to improve work efficiency; in the family life field, it can hand over objects needed for daily activities, allowing the elderly and people with limited mobility to live independently at home; in shopping malls and stores, it can hand over customers' personal belongings to provide customers with a more comfortable environment.

[0062] The dynamic human-machine-object handover method provided in this embodiment can segment the hand point cloud and the object point cloud in the image data, and generate a collision-free grasping configuration based on the hand point cloud and the object point cloud. When the robot moves according to the collision-free grasping configuration, it can avoid collision between the robot and the hand of the handover person, thereby improving the safety of the human-machine-object handover and avoiding collision between the robot and the hand of the handover person during the human-machine-object handover process, thereby avoiding injury to the hand of the handover person caused by injury to the hand of the handover person.

[0063] In addition, the dynamic human-machine object handover method provided in this embodiment can also identify whether the object has moved based on the current frame image and the previous frame image in the real-time acquired image data, and generate a new grasping configuration when the object moves. Therefore, the dynamic human-machine object handover method provided in this embodiment can achieve dynamic adaptation when the object moves, and generate a new grasping configuration.

[0064] Please refer to Figure 2In this embodiment, a dynamic human-machine object handover system based on visual recognition is also provided, which is used to implement the above-mentioned dynamic human-machine object handover method. The dynamic human-machine object handover system includes an input unit, a control unit and a robot. The input unit and the robot are both connected to the control unit. The input unit is used to collect image data in real time. The control unit includes a scene understanding module, a grasping planning module and a dynamic grasping module. The grasping planning module and the dynamic grasping module are both connected to the scene understanding module, and the grasping planning module is connected to the dynamic grasping module; the scene understanding module is used to obtain a hand point cloud and an object point cloud according to the image data; the grasping planning module is used to obtain a collision-free grasping configuration according to the hand point cloud of the first frame image of the image data and the object point cloud of the first frame image of the image data; the dynamic grasping module is used to analyze whether the object moves according to two consecutive frames of images in the image data, and generate a new grasping configuration when the object moves.

[0065] The dynamic human-machine-object handover method provided in this embodiment can be implemented based on the above-mentioned dynamic human-machine-object handover system. According to the dynamic human-machine-object handover system, four major modules, namely, input unit, scene understanding module, grasping planning module and dynamic grasping module, can be set corresponding to the dynamic human-machine-object handover method.

[0066] Please refer to further Figure 3 The input unit can be implemented by a depth camera, and a RealSense D435i camera can be specifically selected to collect image data in real time, wherein the image data includes continuous multi-frame RGBD images. The scene understanding module is used to segment the hand and object in the image data collected by the depth camera, and convert it into a hand-object point cloud (i.e., hand point cloud and object point cloud) that the robot can understand. The grasping planning module is used to generate and filter a safe and reliable collision-free grasping configuration based on the hand point cloud and the object point cloud. The dynamic grasping module includes a dynamic adaptation algorithm, which can analyze whether the object moves based on the continuous current frame image and the previous frame image, and quickly generate a new grasping configuration when the object moves, so as to realize dynamic adaptation and dynamic grasping of moving objects.

[0067] Specifically, the input unit and the scene understanding module are in a continuous running state during the human-machine object handover process, so that the robot can perceive the environment and object changes in real time, and the scene understanding module obtains the hand point cloud and object point cloud in real time based on the image data of the input unit. The grasping planning module is only run once. The grasping planning module uses the initial hand point cloud and object point cloud generated by the first frame of the image data corresponding to the scene understanding module to generate a safe and collision-free grasping configuration with the highest handover success rate. The dynamic grasping module uses a dynamic adaptation algorithm to perform matrix changes on the grasping configuration corresponding to the current grasping action parameters of the robot, so as to quickly generate a new grasping configuration according to the movement state of the object.

[0068] In this embodiment, the input unit, the scene understanding module and the dynamic grasping module form a closed loop, so that the dynamic human-machine object handover system can form a closed-loop control according to the real-time collected image data. The dynamic grasping module dynamically adjusts and generates a new grasping configuration according to the object point cloud obtained by the scene understanding module based on the current frame image, so that the robot's grasping action parameters can respond in time according to the movement changes of the object and adjust the robot's motion state in time.

[0069] In this embodiment, the robot may include a manipulator and a robotic arm; preferably, the manipulator may be implemented by a multi-finger dexterous hand with multiple degrees of freedom, and the Inspire Hand multi-finger dexterous hand may be specifically selected. The shapes of objects in life are all based on the grasping of human hands. The use of a multi-finger dexterous hand for dynamic human-machine object handover can better simulate the grasping action of human hands and ensure the stability and flexibility of grasping, so that the robot applying the dynamic human-machine object handover method can achieve robust, universal and reliable dynamic human-machine object handover.

[0070] In step S1 of this embodiment, obtaining a starting object point cloud and a starting hand point cloud according to the first frame image specifically includes:

[0071] The first frame image includes an RGB image and a depth image; a hand bounding box and an object bounding box are identified based on the RGB image; a segmentation model obtains a hand segmentation mask and an object segmentation mask based on the hand bounding box, the object bounding box and the RGB image, projects the hand segmentation mask onto the depth image to obtain the starting hand point cloud, and projects the object segmentation mask onto the depth image to obtain the starting object point cloud.

[0072] Correspondingly, multiple frames of image data collected in real time can be used to obtain object point clouds and hand point clouds through the above method.

[0073] As an example, the yolov9 target detection model can be used to detect the handing objects and the hands of the handover person in the RGB image in real time to obtain the hand bounding box and the object bounding box; then the hand bounding box, the object bounding box, and the RGB image are passed to the segmentation model SAM, and the hand segmentation mask and the object segmentation mask are obtained by segmentation. The hand segmentation mask and the object segmentation mask are projected onto the depth image to obtain the object point cloud and the hand point cloud respectively.

[0074] In step S2 of this embodiment, obtaining a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud includes:

[0075] A plurality of candidate grasping configurations are generated according to the starting object point cloud, grasping success rates of the plurality of candidate grasping configurations are evaluated respectively, the plurality of candidate grasping configurations are sorted from large to small according to the grasping success rates, and collision detection is performed in sequence until a collision-free grasping configuration is screened out.

[0076] Specifically, before applying the dynamic human-machine object handover method, a grasping evaluation network can be established and trained. The grasping data set of the grasping evaluation network includes grasping data and grasping success rate of grasping data; wherein, 0 can be used to represent grasping data with low grasping success rate, and 1 can be used to represent grasping data with high grasping success rate. The trained grasping evaluation network can be used to evaluate the grasping success rate of the grasping configuration generated by the grasping generation network during the handover process. Accordingly, the closer the grasping success rate is to 1, the better the grasping configuration is.

[0077] In this embodiment, the control unit may further include a simulation module, which may be pybullet, and is used to filter unstable grasping data in the grasping data set. The simulation module can load an object model and a manipulator model according to the grasping data in the data set, and simulate the grasping data to filter out high-quality grasping data.

[0078] The above-mentioned sorting of the plurality of to-be-selected grabbing configurations according to the success rate from large to small and performing collision detection in sequence can be implemented in the following manner:

[0079] A robot model is constructed according to the grasping configuration to be selected, and the distances between all points in the starting hand point cloud and the robot model are calculated respectively to determine whether the distances between all points in the starting hand point cloud and the robot model are less than or equal to a preset distance value; if so, the grasping configuration to be selected is a collision-free grasping configuration; if not, the grasping configuration to be selected is a collision-prone grasping configuration.

[0080] In this embodiment, pytorch3d can be used to build the robot model.

[0081] As an example, the penetration distance between all points in the starting hand point cloud and the manipulator model can be calculated to measure whether the manipulator model will penetrate the hand point cloud, and the preset distance value can be set to 0. When the distance between the manipulator model and at least one point of the hand point cloud is greater than 0, it means that the starting hand point cloud has a point inside the manipulator model, and there is an overlap between the starting hand point cloud and the manipulator model, that is: the robot's grasping action parameters are configured according to the selected grasping configuration, and the robot's manipulator will collide with the handover person's hand during the movement.

[0082] When the distance between the manipulator model and all points of the hand point cloud is less than or equal to zero, it can be ensured that the robot's grasping action parameters are configured according to the grasping configuration. The robot's manipulator will not collide with the hand of the handover person during movement, thereby ensuring the safety of dynamic human-machine object handover and preventing the robot from injuring the hand of the handover person.

[0083] In this embodiment, the collision-free grasping configuration includes orientation parameters in the camera coordinate system, position parameters in the camera coordinate system, and joint angle parameters.

[0084] In step S2 described in this embodiment, the robot's grasping action parameters are configured according to the collision-free grasping configuration, including: transforming the orientation parameters in the camera coordinate system into orientation parameters in the robot coordinate system, transforming the position parameters in the camera coordinate system into position parameters in the robot coordinate system, and closing the robot's manipulator according to the joint angle parameters.

[0085] In step S4 of this embodiment, judging whether the object moves according to the current frame image and the previous frame image includes:

[0086] A point cloud of the object of the previous frame is obtained according to the previous frame image, and a point cloud of the object of the current frame is obtained according to the current frame image, and it is determined whether the distance between the coordinate center of the point cloud of the object of the previous frame and the coordinate center of the point cloud of the object of the current frame is less than a preset distance threshold; if so, the object has not moved; if not, the object has moved.

[0087] While the robot is moving according to the grasping action parameters, the person handling the object may move due to posture adjustment and other reasons. The current frame object point cloud and the previous frame object point cloud are obtained according to the current frame image and the previous frame image respectively, so as to judge whether the object has moved according to whether the distance between the coordinate center of the current frame object point cloud and the coordinate center of the previous frame object point cloud is greater than the preset distance threshold. If the distance between the coordinate center of the current frame object point cloud and the coordinate center of the previous frame object point cloud is less than the preset distance threshold, there is no need to adjust the robot's grasping action parameters; if the distance between the coordinate center of the current frame object point cloud and the coordinate center of the previous frame object point cloud is greater than or equal to the preset distance threshold, the robot's grasping action parameters are adjusted accordingly.

[0088] In step S4 of the present embodiment, a new grasping configuration is generated based on the current frame image and the previous frame image, including: obtaining a transformation matrix between the current frame object point cloud and the previous frame object point cloud; and transforming the grasping configuration corresponding to the current grasping action parameters of the robot according to the transformation matrix to obtain the new grasping configuration.

[0089] Specifically, the new grasping configuration includes new orientation parameters in the camera coordinate system and new position parameters in the camera coordinate system. Since the shape of the object does not change, when the object moves, only the orientation parameters and the position parameters can be updated without updating the joint angle parameters, so as to improve the efficiency of the dynamic human-machine object handover method in generating a new grasping configuration.

[0090] As an example, this embodiment can adopt the ICP algorithm to calculate the transformation matrix of the object point cloud of the current frame and the object point cloud of the previous frame, and transform the current robot's grasping configuration according to the change matrix to generate a new grasping configuration, and update the robot's grasping action parameters according to the new grasping configuration, so that the robot moves according to the target position and direction of the new grasping configuration.

[0091] In this embodiment, step S5 is also included: determining whether the number of times that the continuous object has not moved is equal to a preset number; if so, ending the process; if not, returning to step S4.

[0092] During the human-machine object handover process, the handover person usually adjusts the hand movement at the beginning of the action to cause the object to move. After adjusting the hand movement, they can remain motionless for a certain period of time until the robot completes the object handover action. Therefore, in this embodiment, when the object is judged not to have moved for a preset number of consecutive times, it can be considered that the object is no longer moving, thereby reducing the frequency of dynamic adaptation based on the object movement or stopping dynamic adaptation based on the object movement.

[0093] To sum up, the dynamic human-machine-object handover method provided by the present invention can segment image data into hand point cloud and object point cloud, and can obtain a collision-free grasping configuration by combining the hand point cloud and the object point cloud, so as to avoid the robot and the manipulator colliding with the hand of the handover person in the process of grasping the object, thereby improving the grasping effect of the human-machine-object handover; in addition, the dynamic human-machine-object handover method does not require the handover person to remain still during the human-machine-object handover process, and the dynamic human-machine-object handover method can adapt to the dynamic changes of the object during the handover process, dynamically adapt according to the dynamic changes of the object and generate a new grasping configuration accordingly, thereby improving the grasping success rate of the dynamic human-machine-object handover method.

[0094] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's specification and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A vision-based dynamic human-machine object handover method, characterized in that: The following steps are involved: S1, collecting image data in real time, obtaining the starting object point cloud and the starting hand point cloud according to the first frame image; S2. Acquire a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud; S3, configuring the robot's grasping action parameters according to the collision-free grasping configuration; S4, obtaining a current frame image and a previous frame image, and determining whether the object has moved according to the current frame image and the previous frame image; if so, generating a new grasping configuration according to the current frame image and the previous frame image, and configuring the grasping action parameters of the robot according to the new grasping configuration; if not, maintaining the current grasping action parameters of the robot; In the step S1, obtaining the starting object point cloud and the starting hand point cloud according to the first frame image specifically includes: The first frame image includes an RGB image and a depth image; a hand bounding box and an object bounding box are identified according to the RGB image; a segmentation model obtains a hand segmentation mask and an object segmentation mask according to the hand bounding box, the object bounding box and the RGB image, projects the hand segmentation mask onto the depth image to obtain the starting hand point cloud, and projects the object segmentation mask onto the depth image to obtain the starting object point cloud; In the step S2, obtaining a collision-free grasping configuration according to the starting object point cloud and the starting hand point cloud includes: Generate multiple candidate grasping configurations according to the starting object point cloud, evaluate the grasping success rates of the multiple candidate grasping configurations respectively, sort the multiple candidate grasping configurations from large to small according to the grasping success rates and perform collision detection in sequence until a collision-free grasping configuration is screened out; The method of sorting the plurality of to-be-selected grabbing configurations from large to small according to the grabbing success rates and performing collision detection in sequence includes: A manipulator model is constructed according to the to-be-selected grasping configuration, and the penetration distances between all points in the starting hand point cloud and the manipulator model are calculated respectively to measure whether the manipulator model will penetrate the hand point cloud, and whether the penetration distances between all points in the starting hand point cloud and the manipulator model are less than or equal to a preset penetration distance value; if so, the to-be-selected grasping configuration is a collision-free grasping configuration; if not, the to-be-selected grasping configuration is a collision-prone grasping configuration; The collision-free grasping configuration includes orientation parameters in the camera coordinate system, position parameters in the camera coordinate system and joint angle parameters; the new grasping configuration includes new orientation parameters in the camera coordinate system and new position parameters in the camera coordinate system.

2. The dynamic human-machine-object handover method according to claim 1, characterized in that: In step S2, the robot's grasping action parameters are configured according to the collision-free grasping configuration, including: transforming the orientation parameters in the camera coordinate system into orientation parameters in the robot coordinate system, transforming the position parameters in the camera coordinate system into position parameters in the robot coordinate system, and closing the robot's manipulator according to the joint angle parameters.

3. The dynamic human-machine-object handover method according to claim 1, characterized in that: In the step S4, judging whether the object moves according to the current frame image and the previous frame image includes: A point cloud of the object of the previous frame is obtained according to the previous frame image, and a point cloud of the object of the current frame is obtained according to the current frame image, and it is determined whether the distance between the coordinate center of the point cloud of the object of the previous frame and the coordinate center of the point cloud of the object of the current frame is less than a preset distance threshold; if so, the object has not moved; if not, the object has moved.

4. The dynamic human-machine-object handover method according to claim 3, characterized in that: In step S4, a new grasping configuration is generated according to the current frame image and the previous frame image, including: obtaining a transformation matrix between the current frame object point cloud and the previous frame object point cloud; and transforming the grasping configuration corresponding to the current grasping action parameters of the robot according to the transformation matrix to obtain the new grasping configuration.

5. The dynamic human-machine-object handover method according to claim 1, characterized in that: The process also includes step S5: determining whether the number of times that the object has not moved continuously is equal to a preset number; if so, terminating the process; if not, returning to execute step S4.

6. A dynamic human-machine-object handover system based on visual recognition, characterized in that: The dynamic human-machine-object handover system is used to implement the dynamic human-machine-object handover method described in any one of claims 1 to 5, and includes an input unit, a control unit and a robot, wherein the input unit and the robot are both connected to the control unit, the input unit is used to collect image data in real time, the control unit includes a scene understanding module, a grasping planning module and a dynamic grasping module, the grasping planning module and the dynamic grasping module are both connected to the scene understanding module, and the grasping planning module is connected to the dynamic grasping module; the scene understanding module is used to obtain a hand point cloud and an object point cloud according to the image data; the grasping planning module is used to obtain a collision-free grasping configuration according to the hand point cloud of the first frame of the image data and the object point cloud of the first frame of the image data; the dynamic grasping module is used to analyze whether the object moves according to two consecutive frames of images in the image data, and generate a new grasping configuration when the object moves.

Citation Information

Patent Citations

  • Line obstacle monitoring and alarming system and method based on three-dimensional imaging

    CN110889350A

  • Machine learning control of object handovers

    CN114004329A

  • Composite robot 3D grabbing method and system based on object plane characteristics

    CN114939891A

  • Manipulator flexible grabbing planning method and system based on grabbing gesture detection

    CN115401698A

  • Vision-based robot-to-human object transfer method and device, medium and terminal

    CN115635482A