A Language-Driven Object Grasping Pose Prediction Method, Terminal, and Storage Medium
Through the language-driven object grabbing pose prediction method, combining scene images and language prompt words to predict the object grabbing pose, the problem of poor prediction effect in unstructured sorting scenarios is solved, and more accurate and flexible pose prediction is achieved.
Patent Information
- Application Number
- CN202510162677.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The prior art has poor grasping posture prediction effect in unstructured sorting scenarios, and lacks interactivity, so it is impossible to flexibly specify sorting target categories.
The object grasping attitude prediction method is adopted to predict the object by obtaining scene images and language prompt words, and use modules such as image encoder, object detection and segmenter, prompt encoder and mask decoder to predict the object's grasping attitude.
It realizes more accurate capture pose prediction, expands the interactivity and flexibility of the model, is suitable for unstructured task scenarios, and improves generalization.
Smart Images

Figure CN119625072B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and computer vision, and particularly relates to a language-driven object grasping pose prediction method, a terminal, and a storage medium. Background Art
[0002] In the manufacturing industry, computer vision focuses on creating artificial systems that can capture and analyze visual inputs from the physical world to assist humans in completing various production-related tasks, and provide position information related to environmental interaction such as navigation and grasping for operators or control mechanisms.
[0003] For example, in the scenario of pipeline sorting, industrial robots need the guidance of visual information to perceive objects in the scenario and predict the grasping pose. The four-degree-of-freedom grasping pose is a planar grasping pose representation, including displacements along three coordinate axes and rotation along the z-axis. Such methods represent the grasping pose of an object as a quadruple, where is the two-dimensional grasping center; is the grasping angle, that is, the angle of rotation of the end of the robotic arm around the z-axis; is the grasping width, representing the opening degree of the gripper of the robotic arm. In this pose representation, the robotic arm pose is always perpendicular to the grasping plane.
[0004] However, traditional grasping pose prediction algorithms based on four-degree-of-freedom robotic arms still have certain limitations when facing unstructured sorting scenarios: there are usually multiple targets in the sorting scenario, and traditional algorithms can only be used for a predefined number of object categories, and the prediction effect for newly emerging object categories or unseen object instances is poor. And traditional algorithms are not interactive during inference and do not support flexible specification of the sorting target category.
[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a language-driven object grasping pose prediction method, a terminal, and a storage medium in view of the above-mentioned defects of the existing technology, aiming to solve the problems of poor generalization and low interactivity of the existing grasping pose prediction algorithms.
[0007] The technical solution adopted by the present invention to solve the problem is as follows:
[0008] In a first aspect, an embodiment of the present invention provides a language-driven object grasping pose prediction method, and the method includes:
[0009] Obtain a scene image and a language prompt word;
[0010] Input the scene image and the language prompt into a language-driven object grasping pose prediction model to obtain a grasping pose;
[0011] Among them, the language-driven object grasping pose prediction model includes:
[0012] An image encoder for outputting an image embedding according to the scene image;
[0013] A target detection and segmenter for outputting a bounding box prompt and a point prompt of the target to be grasped according to the scene image and the language prompt;
[0014] A prompt encoder for outputting a prompt embedding according to the bounding box prompt and the point prompt of the target to be grasped;
[0015] A mask decoder for outputting a predicted grasping position mask, a grasping width mask, and a grasping angle mask according to the image embedding and the prompt embedding;
[0016] A grasping pose extraction module for outputting a grasping pose according to the grasping position mask, the grasping width mask, and the grasping angle mask.
[0017] In one implementation, the image encoder includes a number of attention encoder modules and a convolutional layer with 256 channels in the last layer; the attention encoder module includes a number of normalization layers, a multi-head attention layer, and a multi-layer perceptron.
[0018] In one implementation, the prompt encoder includes a point prompt encoder and a bounding box prompt encoder; the point prompt encoder and the bounding box prompt encoder include a number of convolutional layers, normalization layers, and activation function layers; outputting a prompt embedding according to the bounding box prompt and the point prompt of the target to be grasped includes:
[0019] Input the bounding box prompt of the target to be grasped into the bounding box prompt encoder to obtain a first prompt embedding;
[0020] Input the point prompt into the point prompt encoder to obtain a second prompt embedding;
[0021] Use the concatenated data of the first prompt embedding and the second prompt embedding as the prompt embedding.
[0022] In one implementation, the mask decoder includes a composite attention module, an upsampling convolutional module, and a composite perceptron module; outputting a predicted grasping position mask, a grasping width mask, and a grasping angle mask according to the image embedding and the prompt embedding includes:
[0023] Embed the image and the corresponding position embedding into the composite attention module to obtain segmentation features and mask features;
[0024] Input the segmentation features into the upsampling convolution module to obtain low-level position mask features, low-level angle mask features, and low-level width mask features;
[0025] Input the mask features into the composite perceptron module to obtain position tokens, angle tokens, and width tokens;
[0026] Multiply the low-level position mask features by the position tokens to obtain a grasping position mask;
[0027] Multiply the low-level angle mask features by the angle tokens to obtain a grasping angle mask;
[0028] Multiply the low-level width mask features by the width tokens to obtain a grasping width mask.
[0029] In one implementation, according to the grasping position mask, the grasping width mask, and the grasping angle mask, output the grasping pose, including:
[0030] Determine the grasping position peak according to the predicted grasping position mask, and use the grasping position peak as the grasping center;
[0031] At each pixel position of the predicted grasping angle mask, determine the angle index with the maximum confidence, and construct a single-channel grasping angle mask according to the angle indexes at each pixel position;
[0032] Obtain the grasping angle index corresponding to the grasping center according to the single-channel grasping angle mask, and convert the grasping angle index into the corresponding grasping angle;
[0033] Obtain the grasping width corresponding to the grasping center according to the predicted grasping width mask;
[0034] Determine the grasping pose according to the position of the grasping center, the grasping angle, and the grasping width.
[0035] In one implementation, the language-driven object grasping pose prediction model is pre-trained, and the method for generating training samples includes:
[0036] Obtain a number of image data with a preset resolution, where each piece of image data contains a preset number range of objects to be grasped, and the number and positions of the objects to be grasped are randomly determined;
[0037] For each of the described image data, generate a tuple based on the position and grasping pose of each object to be grasped in the image data, generate a grasping label in the form of a mask based on the set of tuples corresponding to the image data, and generate a training sample based on the image data and the grasping label.
[0038] In one implementation, generating a grasping label in the form of a mask based on the set of tuples corresponding to the image data includes:
[0039] For each grasping pose in the set of tuples, determine the neighborhood of the grasping pose;
[0040] Divide the full angular range into several intervals, and for each grasping angle in the set of tuples, calculate the angular mask index of the grasping angle according to the grasping angle and the angular range covered by each interval;
[0041] For each pixel coordinate, determine whether the pixel coordinate is located in the neighborhood of the grasping pose, and generate a grasping position mask label according to the judgment results of the pixel coordinates;
[0042] For each pixel coordinate, determine whether the pixel coordinate is located in the neighborhood of the grasping pose and has a corresponding angular mask index, and generate a grasping angle mask label according to the judgment results of the pixel coordinates;
[0043] For each pixel coordinate, find all width sets covering the pixel coordinate, determine whether the width set is an empty set, and generate a grasping width mask label according to the judgment results of the pixel coordinates;
[0044] Generate the grasping label according to the grasping position mask label, the grasping angle mask label, and the grasping width mask label.
[0045] In one implementation, the training method of the language-driven object grasping pose prediction model includes:
[0046] Initialize the network weights of the image encoder, the object detection and segmentation module, and the prompt encoder according to the pre-trained weights of the vision large model and the vision-language model;
[0047] Input the image data in the training sample into the initial language-driven object grasping pose prediction model to obtain the predicted grasping position mask, grasping angle mask, and grasping width mask during training;
[0048] According to the predicted grasping position mask, grasping angle mask, and grasping width mask during training and the grasping label in the training sample, calculate the model error through the binary cross-entropy loss function, and update the parameters of the image encoder and the mask decoder according to the model error.
[0049] In a second aspect, an embodiment of the present invention further provides a terminal, which includes a memory and more than one processor; the memory stores more than one program; the program contains instructions for executing the language-driven object grasping posture prediction method described in any one of the above; the processor is used to execute the program.
[0050] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor to implement the steps of the language-driven object grasping posture prediction method described in any one of the above.
[0051] Advantages of the present invention: The language-driven object grasping posture prediction model provided by the embodiment of the present invention is a model that introduces language interaction capabilities. It can perform interactive prediction in combination with the language prompt words input by the user, enabling the operator to specify the grasping object through the language prompt words, and the model predicts a more accurate grasping posture. The present invention expands the interactivity and model flexibility of the object grasping posture prediction model, and has strong generalization for unstructured task scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 It is a schematic flowchart of the language-driven object grasping posture prediction method provided by the embodiment of the present invention.
[0054] Figure 2 It is a schematic diagram of the algorithm framework of the language-driven object grasping posture prediction model provided by the embodiment of the present invention.
[0055] Figure 3 It is a schematic diagram of the structure of the image encoder provided by the embodiment of the present invention.
[0056] Figure 4 It is a schematic diagram of the structure of the prompt encoder provided by the embodiment of the present invention.
[0057] Figure 5 It is a schematic diagram of the structure of the mask decoder provided by the embodiment of the present invention.
[0058] Figure 6 It is a schematic block diagram of the terminal according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The present invention discloses a language-driven object grasping pose prediction method, a terminal, and a storage medium. To make the objectives, technical solutions, and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0060] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0061] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.
[0062] In view of the above defects of the prior art, the present invention provides a language-driven object grasping pose prediction method. The method includes obtaining a scene image and a language prompt; inputting the scene image and the language prompt into a language-driven object grasping pose prediction model to obtain a grasping pose. The language-driven object grasping pose prediction model includes: an image encoder for outputting an image embedding according to the scene image; an object detection and segmenter for outputting a bounding box prompt and a point prompt of the object to be grasped according to the scene image and the language prompt; a prompt encoder for outputting a prompt embedding according to the bounding box prompt and the point prompt of the object to be grasped; a mask decoder for outputting a predicted grasping position mask, a grasping width mask, and a grasping angle mask according to the image embedding and the prompt embedding; and a grasping pose extraction module for outputting a grasping pose according to the grasping position mask, the grasping width mask, and the grasping angle mask. The language-driven object grasping pose prediction model provided by the present invention is a model that introduces language interaction ability and can perform interactive prediction in combination with the language prompt input by the user, enabling the operator to specify the grasping object through the language prompt and the model to predict a more accurate grasping pose. The present invention expands the interactivity and model flexibility of the object grasping pose prediction model and has strong generalization ability for unstructured task scenarios.
[0063] As Figure 1 shown, the method includes:
[0064] Step S100: Obtain a scene image and a language prompt;
[0065] Step S200: Input the scene image and the language prompt into a language-driven object grasping pose prediction model to obtain a grasping pose;
[0066] Among them, the language-driven object grasping pose prediction model includes:
[0067] An image encoder for outputting an image embedding according to the scene image;
[0068] An object detection and segmenter for outputting a bounding box prompt and a point prompt of the object to be grasped according to the scene image and the language prompt;
[0069] A prompt encoder for outputting a prompt embedding according to the bounding box prompt and the point prompt of the object to be grasped;
[0070] A mask decoder for outputting a predicted grasping position mask, a grasping width mask, and a grasping angle mask according to the image embedding and the prompt embedding;
[0071] The grasping pose extraction module is used to output the grasping pose according to the grasping position mask, the grasping width mask, and the grasping angle mask.
[0072] Specifically, in this embodiment, a language-driven object grasping pose prediction model is pre-constructed. This model takes the scene image and language prompt words as input data and predicts the grasping pose of the object that matches the language prompt words. The network structure of this model mainly includes an image encoder, a prompt encoder, a mask decoder, and an object detection and segmenter. Among them, the object detection and segmenter can adopt a plug-and-play Grounding SAM object detector (Grounding SAM is a model that can identify and segment objects in pictures). In actual application, first, the scene image is obtained. The scene image can be obtained from the RGB video stream captured by an industrial camera. Then, the language prompt words are given. The language prompt words can be words or phrases input by the user based on the scene image. The language prompt words can be descriptions of the appearance and function of the target object in the scene, used to specify the grasping object. The real-time obtained scene image is input into the image encoder to obtain the image embedding. The real-time obtained scene image and language prompt words are input into the object detection and segmenter to obtain the two-dimensional bounding box of the object to be grasped. The two-dimensional bounding box of the object to be grasped is input into the prompt encoder to obtain the prompt embedding. The image embedding and the prompt embedding are input into the mask decoder to obtain the predicted grasping position mask, grasping width mask, and grasping angle mask. The grasping pose extraction algorithm is used to extract the grasping pose from the predicted grasping position mask, grasping width mask, and grasping angle mask.
[0073] For example, as Figure 2 shown, the image encoder processes the scene image into the image embedding. The Grounding SAM object detector processes the scene image and the user's language prompt words into the bounding box prompt and the point prompt. The prompt encoder processes the bounding box prompt and the point prompt into the prompt embedding. The mask decoder processes the image embedding and the prompt embedding into the position mask, width mask, and angle mask.
[0074] In one implementation, the image encoder includes several attention encoder modules and a convolutional layer with 256 channels in the last layer; the attention encoder module includes several normalization layers, multi-head attention layers, and multi-layer perceptrons.
[0075] Specifically, the image encoder consists of multiple attention encoder modules (Transformer encoder modules) and a final convolutional layer with 256 channels. The number of attention encoder modules can be 12. Each attention encoder module includes several normalization layers, a multi-head attention layer, and a multi-layer perceptron. For example, it can be composed of a stack of a normalization layer, a multi-head attention layer, a normalization layer, and a multi-layer perceptron. The data image encoder takes a training sample image or a scene image obtained in real time as input and outputs an image embedding.
[0076] For example, Figure 3 is the structure of the image encoder. It includes an embedding layer, a dropout layer (i.e., Dropout layer), several stacked encoder modules, and a module composed of several normalization layers and convolutional layers. The number of encoder modules can be 12, and the number of the module composed of normalization layers and convolutional layers can be 2. Specifically, the embedding layer processes the image into a feature embedding of 32x32x768, which is then added to the position embedding. The position embedding is a parameter vector of the same size initialized to 0 and updated during the training process. The encoder module is a standard encoder module in the attention network. The structure of each encoder module is a layer normalization operation, a multi-head attention module, a layer normalization operation, and a multi-layer perceptron. The image encoder also includes a module composed of 2 normalization layers and a convolution with 256 channels.
[0077] In one implementation, the prompt encoder includes a point prompt encoder and a bounding box prompt encoder; the point prompt encoder and the bounding box prompt encoder include several convolutional layers, normalization layers, and activation function layers; according to the bounding box prompt and point prompt of the target to be grasped, a prompt embedding is output, including:
[0078] Input the bounding box prompt of the target to be grasped into the bounding box prompt encoder to obtain a first prompt embedding;
[0079] Input the point prompt into the point prompt encoder to obtain a second prompt embedding;
[0080] Use the concatenated data of the first prompt embedding and the second prompt embedding as the prompt embedding.
[0081] Specifically, the prompt encoder includes a two-dimensional point prompt encoder and a two-dimensional bounding box prompt encoder. The two-dimensional point prompt encoder and the two-dimensional bounding box prompt encoder are stacked by several convolutional layers, several normalization layers, and several activation function layers. Among them, the number of convolutional layers is preferably 3, the number of normalization layers is preferably 2, and the number of activation function layers is preferably 2. The input data of the prompt encoder is the bounding box prompt and point prompt of the target to be grasped output by the target detection and segmenter, and the output data of the prompt encoder is the prompt embedding, which is composed of the outputs of the two-dimensional point prompt encoder and the two-dimensional bounding box encoder spliced together.
[0082] For example, Figure 4 is the prompt encoder structure. Among them, the point prompt is represented as , and is processed by the point embedding layer into an embedding vector of 1x1x256. The two-dimensional bounding box is represented as , and is processed by the bounding box embedding layer into an embedding vector of 1x2x256. Finally, through vector concatenation, the obtained prompt embedding size is 1x3x256.
[0083] In one implementation, the mask decoder includes a composite attention module, an upsampling convolution module, and a composite perceptron module; according to the image embedding and the prompt embedding, it outputs the predicted grasping position mask, grasping width mask, and grasping angle mask, including:
[0084] Input the image embedding and the corresponding position embedding into the composite attention module to obtain the segmentation feature and the mask feature;
[0085] Input the segmentation feature into the upsampling convolution module to obtain the low-level position mask feature, low-level angle mask feature, and low-level width mask feature;
[0086] Input the mask feature into the composite perceptron module to obtain the position token, angle token, and width token;
[0087] Multiply the low-level position mask feature by the position token to obtain the grasping position mask;
[0088] Multiply the low-level angle mask feature by the angle token to obtain the grasping angle mask;
[0089] Multiply the low-level width mask feature by the width token to obtain the grasping width mask.
[0090] Specifically, the mask decoder is composed of a composite attention module, an upsampling convolution module, and a composite perceptron module. Among them, the composite attention module takes the image embedding and the corresponding position embedding as input data, and the output data is the segmentation feature and the mask feature. The position embedding is an immediately generated vector with the same size as the feature embedding. The upsampling convolution module takes the segmentation feature as input data, and the output data is the low-level position mask feature, angle position mask feature, and low-level width mask feature. The composite perceptron module takes the mask feature as input data, and the output data is the position token, angle token, and width token. Finally, multiply the low-level position mask feature by the position token to obtain the grasping position mask; multiply the low-level angle mask by the angle token to obtain the grasping angle mask; multiply the low-level width mask by the width token to obtain the grasping width mask.
[0091] Further, the composite attention module consists of a self-attention layer, a cross-attention layer, a multi-layer perceptron, and two parallel cross-attention layers located at the output heads of the composite attention module, which respectively output segmentation features and mask features.
[0092] For example, Figure 5 It is a mask decoder structure. The image embedding comes from the output data of the image encoder, and the position embedding is a learnable parameter of the same size. The prompt embedding comes from the output data of the prompt encoder. This network structure includes three cross-attention structures. In the first cross-attention module after calculating self-attention, the prompt embedding is used as the Q matrix, and the image embedding is used as the K and V matrices to calculate attention. In the cross-attention structure of the upsampling convolution branch on the output side, the image embedding is used as the Q matrix, and the prompt embedding is used as the K and V matrices to calculate attention. In the cross-attention structure of the multi-layer perceptron branch on the output side, the prompt embedding is used as the Q matrix, and the image embedding is used as the K and V matrices to calculate attention. The output end of the mask decoder contains three groups of upsampling convolutions, and each group of upsampling convolutions is composed of two convolutions with a kernel size of 2, a stride of 2, and a channel number of 128 stacked together. The output end of the mask decoder contains three groups of multi-layer perceptrons, where the channel numbers of the two groups of multi-layer perceptrons corresponding to the position mask and the width mask are 1, and the channel number of the multi-layer perceptron corresponding to the angle mask is 120. The output of each group of upsampling convolutions is multiplied pointwise with the output of the corresponding group of multi-layer perceptrons to obtain the predicted grasping position mask, grasping angle mask, and grasping width mask.
[0093] In one implementation, according to the grasping position mask, the grasping width mask, and the grasping angle mask, the output of the grasping pose includes:
[0094] Determine the grasping position peak according to the predicted grasping position mask, and use the grasping position peak as the grasping center;
[0095] At each pixel position of the predicted grasping angle mask, determine the angle index with the maximum confidence, and construct a single-channel grasping angle mask according to the angle index at each pixel position;
[0096] Obtain the grasping angle index corresponding to the grasping center according to the single-channel grasping angle mask, and convert the grasping angle index into the corresponding grasping angle;
[0097] Obtain the grasping width corresponding to the grasping center according to the predicted grasping width mask;
[0098] Determine the grasping pose according to the position of the grasping center, the grasping angle, and the grasping width.
[0099] Specifically, in this embodiment, a predicted grasping position mask is obtained through a language-driven object grasping pose prediction model , a grasping angle mask and a grasping width mask . After that, a grasping position peak is found from the predicted grasping position mask as the grasping center: . At each pixel position of the grasping angle mask , the angle index with the maximum confidence is determined to construct a single-channel grasping angle mask: . At the grasping center , the corresponding grasping angle index is obtained . Then, the grasping angle index is converted into the corresponding grasping angle . The grasping width is obtained . Finally, the grasping pose is obtained as: .
[0100] In one implementation, the language-driven object grasping pose prediction model is pre-trained, and the generation method of training samples includes:
[0101] Obtain a number of image data with a preset resolution, where each of the image data contains objects to be grasped within a preset quantity range, and the quantity and position of the objects to be grasped are randomly determined;
[0102] For each of the image data, a tuple is generated according to the position and grasping pose of each object to be grasped in the image data, a grasping label in the form of a mask is generated according to the tuple set corresponding to the image data, and a training sample is generated according to the image data and the grasping label.
[0103] Specifically, there are multiple training samples for training the language-driven object grasping pose prediction model, and the generation method of each training sample is the same. Therefore, in this embodiment, one training sample is taken as an example to illustrate the generation process: First, obtain multiple image data with a specified resolution, and each image data needs to contain no more than a specified quantity of objects to be grasped with random quantity and position. The processing methods of each image data are the same. Taking one image data as an example, a tuple is constructed through the objects and grasping poses therein, and a grasping label for this image data is generated through all the constructed tuples. The grasping label includes a grasping position mask label, a grasping angle label, and a grasping width label. Finally, the image data and the corresponding grasping label are combined into a training sample.
[0104] For example, 10k RGB images with a resolution of 500x500 containing grasping targets are obtained for model training, and each image contains no more than 10 objects to be grasped with random quantity, position, and pose. For the th object, the A grasping pose is represented as a tuple , where represents the pixel position of the grasping center, represents the grasping angle, represents the grasping width. The grasping poses contained in each image can be represented as a set , where represents the number of objects, represents the -th number of grasping poses of the object. The grasping labels in the training samples are represented in the form of a mask. Finally, each image-label pair is used as a training sample. In practical applications, 10k training images and their corresponding grasping labels can be obtained from the Jacquard publicly available grasping dataset.
[0105] In one implementation, generating a grasping label in the form of a mask according to the set of tuples corresponding to the image data includes:
[0106] For each grasping pose in the set of tuples, determine the neighborhood of the grasping pose;
[0107] Divide the full angular range into several intervals. For each grasping angle in the set of tuples, calculate the angular mask index of the grasping angle according to the grasping angle and the angular range covered by each interval;
[0108] For each pixel coordinate, determine whether the pixel coordinate is located in the neighborhood of the grasping pose, and generate a grasping position mask label according to the judgment results of each pixel coordinate;
[0109] For each pixel coordinate, determine whether the pixel coordinate is located in the neighborhood of the grasping pose and has a corresponding angular mask index, and generate a grasping angle mask label according to the judgment results of each pixel coordinate;
[0110] For each pixel coordinate, find all width sets covering the pixel coordinate, determine whether the width set is an empty set, and generate a grasping width mask label according to the judgment results of each pixel coordinate;
[0111] Generate the grasping label according to the grasping position mask label, the grasping angle mask label, and the grasping width mask label.
[0112] Specifically, the grasping label includes three types of data, and these three types of data need to be generated separately to obtain the final grasping label. First, calculate the neighborhood of each grasping pose in the tuple set, which is used to describe a subset containing the grasping pose. Then divide the full angular range (i.e., 360°) into multiple intervals at a preset angular interval, and the angular range covered by each interval is the same. For each grasping angle in the tuple set, assign an angular mask index to this grasping angle based on this grasping angle and the preset angular interval (i.e., the angular range covered by each interval). Then determine the data of each pixel point in different mask labels pixel by pixel. Different mask labels use different judgment conditions to determine the value of the pixel point in the mask label. Taking a pixel point as an example, according to the coordinates of this pixel point, judge whether it is located in the neighborhood of the grasping pose. Yes and no correspond to different values respectively. After judging each pixel point, the grasping position mask label can be generated according to the values of all pixel points. According to the coordinates of this pixel point, judge whether it is located in the neighborhood of the grasping pose and has a corresponding angular mask index. Yes and no correspond to different values respectively. After judging each pixel point, the grasping angle mask label can be generated according to the values of all pixel points. According to the coordinates of this pixel point, find the set of widths that cover the coordinates of this pixel point, and judge whether this set of widths is an empty set. Yes and no correspond to different values respectively. After judging each pixel point, the grasping width mask label can be generated according to the values of all pixel points.
[0113] For example, according to the grasping pose The method for generating the corresponding grasping position mask, grasping angle mask, and grasping width mask label is as follows: For each grasping pose , define its neighborhood:
[0114] ;
[0115] Divide the angular range from 0° to 360° into 120 intervals equally, and each interval covers 3°. For each grasping angle , calculate its angular mask index:
[0116] ;
[0117] Among them, , represents the radian range covered by each interval, .
[0118] Generate the grasping position label , for all pixel coordinates , let:
[0119] .
[0120] Generate the grasping angle label , for each angular mask and all pixel coordinates , let:
[0121] .
[0122] Generate a grasping width label , for all pixel coordinates , find all sets of widths that cover this pixel:
[0123] .
[0124] The grasping width label is calculated as:
[0125] .
[0126] In one implementation, the training method of the language-driven object grasping pose prediction model includes:
[0127] Initialize the network weights of the image encoder, the object detection and segmenter, and the prompt encoder according to the pre-trained weights of the vision large model and the vision-language model;
[0128] Input the image data in the training sample into the initial language-driven object grasping pose prediction model to obtain the grasping position mask, grasping angle mask, and grasping width mask predicted during training;
[0129] According to the grasping position mask, grasping angle mask, and grasping width mask predicted during training and the grasping labels in the training sample, calculate the model error through the binary cross-entropy loss function, and update the parameters of the image encoder and the mask decoder according to the model error.
[0130] Specifically, the network weights of the language-driven object grasping pose prediction model are initialized with the pre-trained weights of the vision large model and the vision-language model, where the weights of the object detection and segmenter and the prompt encoder are frozen. The purpose of introducing the vision large model during model training in this embodiment is to utilize its pre-trained weights on a large-scale image dataset to enhance the generalization performance of the model. In this embodiment, the network structure and network parameters are fine-tuned based on the weights of large models such as the vision large model and the vision-language model, which can effectively utilize the extensive visual knowledge of the vision large model. After initialization, the language-driven object grasping pose prediction model is iteratively updated with multiple training samples to optimize the model parameters. In this embodiment, the model before completing model training is defined as the initial language-driven object grasping pose prediction model. Taking a training sample as an example, the image data in the training sample is input into the initial language-driven object grasping pose prediction model, and the grasping position mask, grasping angle mask, and grasping width mask predicted by the initial language-driven object grasping pose prediction model are obtained. The three masks predicted by the model are compared with the three mask labels in the grasping label according to the category correspondence relationship, and the model error is calculated through the binary cross-entropy loss function. The model error can reflect the gap between the model prediction and the ground truth. Guided by the model error, the parameters of the image encoder and the mask decoder are updated, which can effectively improve the model performance.
[0131] For example, the pre-trained weights of the vision large model SAM and the vision-language model Grounding SAM are used to initialize the network weights of the image encoder, the prompt encoder, and the object detection and segmenter. The image in the training sample is input into the image encoder to obtain an image embedding. The image data and the language prompt words in the training sample are input into the object detection and segmenter to obtain the two-dimensional bounding box and the semantic mask of the object to be grasped, and a prompt point is randomly sampled from the semantic mask. The two-dimensional bounding box of the object to be grasped is input into the prompt encoder to obtain a prompt embedding. The image embedding and the prompt embedding are input into the mask decoder to obtain the predicted grasping position mask, grasping angle mask, and grasping width mask. Finally, the three masks and the grasping label are used to calculate the error through the binary cross-entropy loss function, and the network parameters of the image encoder and the mask decoder are updated.
[0132] In one implementation, the binary cross-entropy loss function is the sum of three parts: the grasping position mask loss, the grasping angle mask loss, and the grasping width mask loss.
[0133] For example, the loss function of the language-driven object grasping pose prediction model is expressed as a weighted binary cross-entropy loss:
[0134] ;
[0135] where and represent the length and width of the image, and respectively represent the true value and the predicted value of the mask at ; represents the Sigmoid activation function, represents the weight term.
[0136] is calculated as follows:
[0137] ;
[0138] respectively use binary cross - entropy loss with weights to calculate the grasping position mask loss , the grasping angle mask loss and the grasping width mask loss . Finally, the loss of the language - driven object grasping pose prediction model is .
[0139] Based on the above - mentioned embodiments, the present invention also provides a terminal, and its principle block diagram can be as Figure 6 shown. The terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non - volatile storage medium and an internal memory. The non - volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non - volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes the language - driven object grasping pose prediction method. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.
[0140] Those skilled in the art can understand that Figure 6 the principle block diagram shown in
[0141] is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0142] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0143] In summary, the present invention discloses a language-driven object grasping pose prediction method, a terminal, and a storage medium, which relate to the fields of artificial intelligence and computer vision technologies. The method includes obtaining a scene image and a language prompt word; inputting the scene image and the language prompt word into a language-driven object grasping pose prediction model to obtain a grasping pose. The language-driven object grasping pose prediction model includes: an image encoder for outputting an image embedding according to the scene image; a target detection and segmenter for outputting a bounding box hint and a point hint of the target to be grasped according to the scene image and the language prompt word; a hint encoder for outputting a hint embedding according to the bounding box hint and the point hint of the target to be grasped; a mask decoder for outputting a predicted grasping position mask, a grasping width mask, and a grasping angle mask according to the image embedding and the hint embedding; and a grasping pose extraction module for outputting a grasping pose according to the grasping position mask, the grasping width mask, and the grasping angle mask. The language-driven object grasping pose prediction model provided by the present invention is a model that introduces language interaction capabilities and can perform interactive prediction by combining the language prompt words input by the user, enabling the operator to specify the grasping object through the language prompt words and the model to predict a more accurate grasping pose. The present invention expands the interactivity and model flexibility of the object grasping pose prediction model and has strong generalization ability for unstructured task scenarios.
[0144] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A language-driven object grasping posture prediction method, characterized in that: The method comprises: Obtain scene images and language prompts; Inputting the scene image and the language prompt word into a language-driven object grasping posture prediction model to obtain a grasping posture; Wherein, the language-driven object grasping posture prediction model includes: An image encoder, configured to output an image embedding based on the scene image; An object detector and segmenter, used for outputting a bounding box prompt and a point prompt of a target to be captured according to the scene image and the language prompt word; A hint encoder, comprising: a point hint encoder and a bounding box hint encoder; the point hint encoder and the bounding box hint encoder comprise a plurality of convolutional layers, a normalization layer and an activation function layer; the hint encoder is used to input the bounding box hint of the target to be captured into the bounding box hint encoder to obtain a first hint embedding; input the point hint into the point hint encoder to obtain a second hint embedding; and use the concatenated data of the first hint embedding and the second hint embedding as a hint embedding; a mask decoder for outputting a predicted grasp position mask, a grasp width mask, and a grasp angle mask based on the image embedding and the hint embedding; The grasping posture extraction module is used to output the grasping posture according to the grasping position mask, the grasping width mask and the grasping angle mask.
2. The language-driven object grasping posture prediction method according to claim 1, characterized in that: The image encoder includes several attention encoder modules and a convolutional layer with a last channel number of 256; the attention encoder module includes several normalization layers, multi-head attention layers and a multi-layer perceptron.
3. The language-driven object grasping posture prediction method according to claim 1, characterized in that: The mask decoder includes a composite attention module, an upsampling convolution module, and a composite perceptron module; outputs a predicted grasp position mask, a grasp width mask, and a grasp angle mask according to the image embedding and the prompt embedding, including: Inputting the image embedding and the corresponding position embedding into the composite attention module to obtain segmentation features and mask features; Inputting the segmentation features into the upsampling convolution module to obtain low-level position mask features, low-level angle mask features and low-level width mask features; Input the mask feature into the composite perceptron module to obtain a position token, an angle token and a width token; multiplying the low-level position mask feature with the position token to obtain a grasp position mask; multiplying the low-level angle mask feature with the angle token to obtain a grasped angle mask; The low-level width mask feature is multiplied by the width token to obtain a grabbed width mask.
4. The language-driven object grasping posture prediction method according to claim 1, characterized in that: Outputting a grasping posture according to the grasping position mask, the grasping width mask, and the grasping angle mask includes: Determine a grasping position peak value according to the predicted grasping position mask, and use the grasping position peak value as a grasping center; Determine an angle index with a maximum confidence at each pixel position of the predicted grasping angle mask, and construct a single-channel grasping angle mask according to the angle index at each pixel position; Acquire a grasping angle index corresponding to the grasping center according to the single-channel grasping angle mask, and convert the grasping angle index into a corresponding grasping angle; Acquire a grasping width corresponding to the grasping center according to the predicted grasping width mask; The grasping posture is determined according to the position of the grasping center, the grasping angle, and the grasping width.
5. The language-driven object grasping posture prediction method according to claim 1, characterized in that: The language-driven object grasping posture prediction model is pre-trained, and the method for generating training samples includes: Acquire a plurality of image data of a preset resolution, wherein each of the image data contains objects to be grasped within a preset number range, and the number and position of the objects to be grasped are randomly determined; For each of the image data, a tuple is generated according to the position and grasping posture of each object to be grasped in the image data, a grasping label in mask form is generated according to the tuple set corresponding to the image data, and a training sample is generated according to the image data and the grasping label.
6. The language-driven object grasping posture prediction method according to claim 5, characterized in that: Generating a capture label in the form of a mask according to a set of tuples corresponding to the image data, including: For each grasping posture in the set of tuples, determining a neighborhood of the grasping posture; Divide the full angle range into several intervals, and for each grasping angle in the tuple set, calculate the angle mask index of the grasping angle according to the grasping angle and the angle range covered by each interval; For each pixel coordinate, determine whether the pixel coordinate is located in the neighborhood of the grasping posture, and generate a grasping position mask label according to the determination result of each pixel coordinate; For each pixel coordinate, determine whether the pixel coordinate is located in the neighborhood of the grasping posture and has a corresponding angle mask index, and generate a grasping angle mask label according to the determination result of each pixel coordinate; For each pixel coordinate, find all width sets covering the pixel coordinate, determine whether the width set is an empty set, and generate a captured width mask label according to the determination result of each pixel coordinate; The grasping label is generated according to the grasping position mask label, the grasping angle mask label, and the grasping width mask label.
7. The language-driven object grasping posture prediction method according to claim 5, characterized in that: The training method of the language-driven object grasping posture prediction model includes: Initializing network weights of the image encoder, the object detector and segmenter, and the cue encoder according to pre-trained weights of a visual macro model and a visual language model; Inputting the image data in the training sample into the initial language-driven object grasping posture prediction model to obtain the grasping position mask, grasping angle mask and grasping width mask predicted during training; According to the grasping position mask, grasping angle mask and grasping width mask predicted during training and the grasping labels in the training samples, the model error is calculated through a binary cross entropy loss function, and the parameters of the image encoder and the mask decoder are updated according to the model error.
8. A terminal, characterized in that: The terminal includes a memory and one or more processors; the memory stores one or more programs; the program contains instructions for executing the language-driven object grasping posture prediction method as described in any one of claims 1-7; and the processor is used to execute the program.
9. A computer-readable storage medium having a plurality of instructions stored thereon, characterized in that: The instructions are suitable for being loaded and executed by a processor to implement the steps of the language-driven object grasping posture prediction method as described in any one of claims 1-7.
Citation Information
Patent Citations
Target object grabbing method and system based on multi-source knowledge driving
CN118038221A