This invention relates to the field of
machine vision technology and discloses a method, apparatus, device, and medium for object
pose recognition and grasping. The method includes: receiving a
natural language instruction containing target
semantics, and acquiring visual information corresponding to the warehousing operation environment based on the
natural language instruction, wherein the target
semantics includes a target object; inputting the visual information and the
natural language instruction into a pre-trained end-to-end visual language
action model for cross-
modal alignment and fusion to generate an instance recognition result of the target object and a candidate grasping
pose corresponding to the instance recognition result; transforming the candidate grasping
pose to a pre-constructed
robot base coordinate
system, and optimizing the candidate grasping pose after the coordinate
system transformation based on kinematic constraints to obtain a
robot joint trajectory sequence; controlling the warehousing
robot to execute the joint trajectory sequence to complete the grasping of the target object. This invention improves the accuracy of object pose recognition and grasping.