This invention discloses a method, apparatus, device, and storage medium for determining a grasping posture. It acquires a visual image and a depth image of the current scene, as well as
natural language information input by the user. The
natural language information describes the feature information of the object to be grasped. Based on the visual image and the
natural language information, it determines the availability points of the object to be grasped in the visual image. Based on the availability points and the visual image, it determines a
mask image of the object to be grasped. Based on the visual image, depth image, camera parameters, and the
mask image, it determines the grasping posture of the object to be grasped. By directly locating the availability points of the object through natural language information, and using these points to guide local segmentation to generate an accurate
mask, and then combining the depth image and camera parameters to calculate the 3D grasping posture, it avoids the
blindness and category dependence of traditional methods in complex, multi-object, and similar object scenes, significantly improving the accuracy and efficiency of grasping.