The invention discloses a grabbing posture generation method and
system based on a multi-
modal large model, and the method comprises the steps: carrying out the cross-
modal matching of visual features and semantic features in the multi-
modal large model when a voice instruction and an
RGB image are inputted, and obtaining the position information of a
control function code and a target object; when an
RGB image with a hand drawing instruction is input, obtaining position information of a
control function code, a target object and a path point; calculating the
point cloud data of the target object according to the position information of the target object and the depth information, inputting the ideal
point cloud of the target object into a target recognition
network model after preprocessing, carrying out the grabbing region recognition of the
point cloud of the target object region, outputting a region with high grabbing confidence, and mapping a real coordinate
system; constructing a point cloud bounding box, and generating a grabbing posture candidate set; the grabbing posture with the highest quality is selected as the grabbing posture of the
robot by calculating the grabbing posture candidate
score; and executing a target grabbing task in combination with the
control function code and the grabbing path.