Multi-specification target object grabbing method based on large model

By constructing a grab gesture dataset for multi-spec objects and semantic guidance for large-scale objects, the adaptability problem of traditional robot grasping methods to complex and variable objects is solved, and more efficient and adaptable object grasping is achieved, improving the accuracy and adaptability of robot grasping technology.

CN120244978AActive Publication Date: 2025-07-04SHENYANG UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510587079.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-04
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Traditional robot grasping methods are difficult to deal with complex and changeable object specification differences, especially when no objects are seen. They require a lot of manpower and material resources to adjust parameters and optimize models, and it is difficult to formulate an effective grasping plan.

Method used

Build a grab gesture data set for multi-spec objects, use a large model to guide object specifications semantics, integrate prompt word engineering technology and language big model to obtain object shape, components, size, materials and text information of grab gesture descriptions, build a grab generation model for object specification perception, and use the noise prediction network to gradually correct the noise-containing state to get closer to the real grab gesture parameters.

Benefits of technology

It improves the stability and generalization ability of robot grasping methods, can better deal with diversity and no objects, improves the efficiency and adaptability of grasping, and promotes the accuracy and adaptability of robot grasping technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244978A_ABST
    Figure CN120244978A_ABST
Patent Text Reader

Abstract

The invention provides a large model-based multi-specification target object grabbing method, which comprises the following steps of: constructing a grabbing gesture data set for multi-specification objects, and carrying out object specification semantic guidance based on the large model (i.e., fusing a cue word engineering technology and a language large model, obtaining text information of object shapes, parts, sizes, materials and grabbing gesture description, and carrying out object specification semantic guidance on the basis of the large model; the method comprises the steps of object specification representation learning and object specification perception grabbing generation model construction. Depending on the advanced large model technology and the unique generation capability of a diffusion model, the invention aims to propose a grabbing mechanism with object specification information as guidance and construct a generation model with generalization and object specification perceptibility, and focuses on solving the key work of grabbing gesture generation of objects with different specifications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent robot target grasping, and specifically to a multi-specification target object grasping method based on a large model. Background Technique

[0002] Intelligent robots are the high point of competition in the technology industry and play a crucial role in the intelligent transformation of China's industries. Target object grasping is an important link in multiple scenarios, but the multi-specification characteristics of objects pose challenges. Traditional grasping methods have obvious limitations. Facing complex and variable objects, a large amount of manpower and material resources are required to adjust parameters and optimize models, and it is difficult to formulate grasping plans for unseen objects.

[0003] Large models have demonstrated powerful capabilities in multiple fields. Their massive data learning and generalization capabilities bring the possibility of breaking through the limitations of traditional grasping methods. Therefore, how to obtain a grasping dataset containing object specification information, utilize the massive data learning and generalization capabilities of large models, and build a generative model with both generalization and object specification perception is the key to solving the above problems. Summary of the Invention

[0004] The present invention proposes a multi-specification target object grasping method based on a large model, aiming to effectively address the challenges brought by specification differences when grasping objects, solve the problem that traditional grasping methods are difficult to handle complex and variable objects and unseen objects, open up a new path for realizing intelligent, efficient, and adaptable object grasping, and provide an innovative solution with important value.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A multi-specification target object grasping method based on a large model, comprising the following steps:

[0007] Step 1: Construct a grasping gesture dataset for multi-specification objects, that is, construct a grasping dataset composed of object shape, components, dimensions, materials, object specification information, and robotic hand grasping gestures;

[0008] Step 2:

[0009] Perform object specification semantic guidance based on a large model, that is, fuse prompt engineering technology and a language large model to obtain text information describing object shape, components, dimensions, materials, and grasping gestures for object specification representation learning;

[0010] Step 3: Construct a grasping generation model for object specification perception. The model consists of a semantic guidance module and a grasping gesture generation module. Machine hand parameters are generated through the grasping generation model. The grasping gesture generation module gradually perturbs the initial parameters into Gaussian noise through forward diffusion and denoises them inversely. The noise prediction network is gradually corrected to make the noisy state approach the true grasping gesture parameters, and finally, the noise prediction network is supervised and trained by the noise true value term obtained through forward diffusion.

[0011] Preferably, in Step 1, the process of constructing a grasping gesture dataset for multi-specification objects includes:

[0012] Step 1-1: Label the shape, size, material, and components of the object to construct the specification text information of each object;

[0013] Step 1-2: Adopt an optimized sampling method and combine it with interactive simulation software and human experience to collect grasping gestures that are adapted to the object specification semantics and have different degrees of freedom.

[0014] Preferably, in Step 1-1, the shape annotation of the object includes: obtaining the point cloud data of the object using a high-precision 3D scanner, denoising and smoothing it through MeshLab and CloudCompare software, and then constructing the 3D mesh model of the object; manually marking the key geometric feature points of the object, including edge vertices, curvature mutation points, and symmetry axis positions, and structurally annotating the object shape type and geometric feature parameters through an annotation tool;

[0015] The size annotation of the object includes: using a ranging device to measure the size parameters of the object in multiple dimensions. For deformable objects, record the size range in different states and indicate the measurement conditions during annotation.

[0016] Preferably, in Step 1-2, collecting grasping gestures that are adapted to the object specification semantics and have different degrees of freedom includes simulation environment construction, human experience-driven acquisition, and automated assisted acquisition simulation environment construction;

[0017] The process of simulation environment construction is as follows: In the interactive simulation software, import the annotated 3D model of the object and various types of manipulator models, configure the simulation parameters that conform to the real physical laws, including gravitational acceleration, friction, and collision detection threshold. At the same time, set up a virtual camera to record the images and video data of the grasping process.

[0018] The process of automated assisted acquisition is as follows: Develop a rule-based automated acquisition script to automatically generate a series of grasping action sequences according to the object specification features. At the same time, use the reinforcement learning algorithm to conduct autonomous exploration in the simulation environment, optimize the grasping strategy through the reward mechanism, and supplement the edge case samples that are difficult to cover by manual acquisition.

[0019] Preferably, in step 2, semantic guidance for object specifications is performed based on a large model, that is, by integrating prompt engineering techniques and a large language model, the process of obtaining text information on object shape, components, dimensions, materials, and grasping gestures for object specification representation learning includes:

[0020] Step 2-1: Prompt engineering design, including basic prompt construction, task-oriented prompt optimization, and multi-round interaction prompt strategies. The basic prompt construction is to design a basic prompt template for different dimensions of object specification information; the task-oriented prompt optimization is to design task-oriented prompts in combination with the requirements of the grasping task, guiding the model to generate semantic information directly related to the grasping operation, and adding constraint condition prompts to improve the practicality of the generated content; the multi-round interaction prompt strategy is to adopt a multi-round prompt interaction mechanism, inputting basic object specification information in the first round to obtain a preliminary description, and asking questions about ambiguous or key information in the output of the first round in the second round, and guiding the model to output semantic content through iterative interaction.

[0021] Step 2-2: Interaction with the GPT-4v large model, including interface call and parameter configuration, multi-modal information input, and result screening and quality evaluation. The interface call and parameter configuration is to connect to the GPT-4v model through the API interface provided by OpenAI and set parameters; the multi-modal information input is to utilize the multi-modal processing ability of GPT-4v. In addition to inputting object specification information in text form, the rendered image of the object's three-dimensional model, local detail close-up pictures, and material texture images are used as supplementary inputs, and metadata is added to the images through an image annotation tool; the result screening and quality evaluation is to perform double screening of the text information output by the model, automatically and manually.

[0022] Step 2-3: Text information processing and feature extraction, including text cleaning and normalization, text embedding and feature fusion. The text cleaning and normalization is to clean the screened text, remove redundant characters, special symbols, and duplicate content, unify the text format, and perform word segmentation, part-of-speech tagging, and named entity recognition using natural language processing tools to extract key semantic information; the text embedding and feature fusion is to use a pre-trained text embedding model to convert the cleaned text into a vector representation, obtain the text features of the object specifications, and fuse the text features with the object geometric features, material features, etc. extracted in step 1.

[0023] Preferably, in step 2-2, the process of interacting with the GPT-4v large model includes:

[0024] Interface Call and Parameter Configuration: Connect to the GPT-4v model through the API interface provided by OpenAI, and set parameters including temperature, maximum number of tokens, frequency penalty, and presence penalty. Adjust the parameters according to the task requirements to balance the generation efficiency and diversity while ensuring the quality of the generated content;

[0025] Multi-modal Information Input: Utilize the multi-modal processing ability of GPT-4v. In addition to inputting the object specification information in text form, render images of the object's three-dimensional model, close-up pictures of local details, and texture images of materials as supplementary inputs. Add metadata to the images through image annotation tools to help the model better understand the object features and generate more comprehensive semantic descriptions;

[0026] Result Screening and Quality Evaluation: Conduct double screening of the text information output by the model, automatically and manually. For automatic screening, set keyword matching and grammar checking rules to filter out content that does not meet the requirements.

[0027] Preferably, in step three, the semantic guidance module obtains the object semantic geometric features by using a three-dimensional shape feature encoder, interprets the specification semantic instructions with the help of the large language model LLM, obtains the object specification features through text embedding, and designs a modality adapter to fuse the object and specification text features to obtain a general feature representation.

[0028] Preferably, in step three, the process of the semantic leading module for extracting each object feature is as follows:

[0029] Step 3-1a: Obtain the object semantic geometric features by using a three-dimensional shape feature encoder;

[0030] Step 3-2a: Interpret the specification semantic instructions with the help of the large language model LLM, and obtain the object specification features through text embedding;

[0031] Step 3-3a; Design a modality adapter to fuse the object and specification text features to obtain a general feature representation applicable to objects of different specifications.

[0032] Preferably, in step one, the process of generating the robotic arm parameters by the grasping generation model includes:

[0033] Step 3-1b: With the help of a pre-set noise schedule Starting from the initial parameter J0 = J′ t and following the Markov chain rule: Gradually add noise. At each step t, the current state J t is obtained by adding noise that follows a specific Gaussian distribution to the previous state J t-1 As the steps progress, J t gradually approaches the Gaussian noise where Represents a Gaussian distribution with a mean of 0 and a covariance matrix of the identity matrix I;

[0034] Step 3-2b: The reverse denoising process is described by where p θ represents the probability distribution of the previous state J t given the parameters θ, the current noisy state J g and the grasping target-related information F t-1 . It is also a Gaussian distribution, and its mean is determined by μ θ ((J t , t, F g ), and the covariance is determined by ∑ t ;

[0035] The mean function is expressed as:

[0036] where α t = 1 - β t is the noise attenuation coefficient, which is used to measure the degree of noise reduction in each step; is the cumulative attenuation factor, which accumulates the noise attenuation from the first step to the t-th step; ∈ θ is the noise prediction network obtained through training. It predicts the noise based on the current noisy state J t , step t, and the grasping target information F g ;

[0037] Through the mean function Formula , using the prediction result of the noise prediction network, the noisy state J t is gradually corrected to approach the true grasping gesture parameters;

[0038] Step 3-3b: The training of the noise prediction network ∈ θ relies on the true noise term ∈ t calculated in the forward diffusion process for supervision, and is specifically achieved by minimizing the loss function , where represents taking the expectation, and ||·|| 2 represents the square of the Euclidean norm. This loss function measures the gap between the noise ∈ θ (J t , t, F g ) predicted by the noise prediction network and the true noise ∈ t . During the training process, the network parameters θ are continuously adjusted to make the loss function value as small as possible.

[0039] Beneficial effects:

[0040] 1. The present invention utilizes the advantages of the massive data learning and generalization capabilities of large models to break through the limitations of traditional grasping methods in dealing with complex and changeable objects and unseen objects.

[0041] 2. Compared with the prior art, the present invention enriches the dexterous hand grasping dataset, collects and annotates data and models containing information such as shape, size, and material, provides data support for grasping generation, improves the stability, generalization, and diversity handling capabilities of the grasping method, contributes to the breakthrough of grasping technology, and lays a data foundation. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flowchart of the present invention;

[0043] Figure 2 Research plan for generating grasping gestures for multi-specification objects;

[0044] Figure 3 Technical route for generating grasping gestures for multi-specification objects DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] As Figure 1 shown, the method of the present invention mainly includes three steps: constructing a grasping gesture dataset for multi-specification objects, performing object specification semantic guidance based on a large model, and constructing a grasping generation model with object specification perception;

[0047] Specifically, as Figure 2 、 Figure 3 shown, the detailed process is as follows:

[0048] Step 1: Construction of a grasping gesture dataset for multi-specification objects. First, annotate the object shape, size, material, and components to construct specification text information. According to the annotated information, use the optimized sampling method, with the help of interactive simulation software and combined with manual experience, to collect grasping gestures that adapt to object specification semantics and different degrees of freedom. Finally, construct a grasping dataset consisting of objects, object specification information, and robotic hands.

[0049] 1.1 Object information collection and annotation

[0050] Data collection: Collect objects of different specifications from multiple channels such as industrial production scenarios, smart home fields, and logistics and warehousing environments, covering various materials such as metals, plastics, glass, and ceramics, as well as various forms such as regular shapes (such as cuboids, cylinders), irregular shapes (such as special-shaped parts, sculptures), and combined bodies (such as containers with handles, assembled toys), and establish a basic object database.

[0051] Shape annotation: Use a high-precision 3D scanner to obtain the point cloud data of the object. After denoising and smoothing it through software such as MeshLab and CloudCompare, construct a 3D mesh model. Manually mark the key geometric feature points of the object, such as edge vertices, curvature mutation points, and symmetry axis positions, and use annotation tools (such as CVAT) to structurally annotate the object shape type (such as spherical, conical, T-shaped structure, etc.) and feature parameters (such as radius, angle, symmetry distance).

[0052] Dimension measurement: Use equipment such as vernier calipers, laser rangefinders, and coordinate measuring machines to perform multi-dimensional measurements on the size parameters of the object, such as length, width, height, aperture diameter, and wall thickness. For deformable objects (such as hoses, fabrics), record their size ranges in different states (stretching, compression, bending), and indicate the measurement conditions during annotation.

[0053] Material annotation: Detect the chemical composition of the object through spectral analysis equipment (such as Fourier transform infrared spectrometers), and combine tactile perception, visual observation, etc. to annotate the detailed properties of the object material, including surface roughness, hardness grade, friction coefficient, etc. At the same time, establish a material label system to distinguish different characteristics such as "smooth hard plastic" and "matte metal".

[0054] Component annotation: For assembled objects, disassemble their component parts, annotate the name, function, connection method (such as threaded connection, snap connection) and assembly sequence of each part. Use the assembly module in 3D modeling software to generate a component hierarchy diagram, and associate and bind the annotation information with the components of the 3D model.

[0055] 1.2. Optimization of sampling strategy design

[0056] Feature weight calculation: Based on factors such as the shape complexity of the object (such as the number of vertices, curvature change frequency), size (such as the ratio of volume to the manipulator's gripping range), material properties (such as the correlation between friction coefficient and grasping stability), and component structure (such as the position of key stress points), establish a multi-dimensional weight evaluation model. For example, for precision parts with complex shapes and small sizes, assign a higher sampling weight to ensure that enough targeted grasping samples are collected.

[0057] Sampling range division: Divide the object specification parameter space into multiple sub-regions, and group the objects through a clustering algorithm (such as K-Means). For the typical features of each sub-region, formulate a differentiated sampling plan. For example, in the large-size object region, focus on collecting gestures at different grasping positions, and in the high-friction coefficient material region, increase the experimental samples of grasping force changes.

[0058] 1.3. Grasp gesture data collection

[0059] Simulation environment construction: In interactive simulation software such as Gazebo and V-REP, import the annotated 3D models of objects and various types of manipulator models (such as two-finger grippers, multi-finger dexterous hands, suction cup grippers), and configure simulation parameters that conform to real physical laws, including gravitational acceleration, friction, collision detection thresholds, etc. At the same time, set up a virtual camera to record the image and video data of the grasping process.

[0060] Artificial experience-driven collection: Invite experts in the field of robot operation and experienced engineers to manually control the manipulator to perform grasping tasks in the simulation environment according to the object specification annotation information and optimized sampling strategy. During the operation, by adjusting parameters such as the pose (position, rotation angle), joint motion trajectory, and grasping force of the manipulator, try different grasping methods under different degrees of freedom (such as parallel grasping, envelope grasping, adsorption grasping), and record data such as the joint angles of the manipulator, the relative pose between the object and the manipulator, and the grasping success / failure status for each grasping.

[0061] Automated assisted collection: Develop a rule-based automated collection script to automatically generate a series of grasping action sequences according to the object specification features. For example, for a cuboid object, the script automatically generates actions to grasp from different sides and different heights. At the same time, use reinforcement learning algorithms to conduct autonomous exploration in the simulation environment, and optimize the grasping strategy through reward mechanisms (such as grasping stability score, grasping time consumption) to supplement the edge case samples that are difficult to cover by manual collection.

[0062] 1.4. Dataset integration and verification

[0063] Data cleaning: Perform deduplication processing on the collected raw data, and eliminate duplicate grasping samples and invalid data (such as data with incomplete grasping actions). Check the integrity and accuracy of the data through visualization tools, and repair the annotation errors or missing information.

[0064] Data annotation enhancement: Add additional annotations to each grasping sample, including grasping type labels (such as precise grasping, stable grasping), applicable scenario labels (such as industrial assembly, home service), risk level labels (such as easy to slip, easy to damage), etc., to enhance the semantic richness of the dataset.

[0065] Dataset Division: Divide the dataset into a training set, a validation set, and a test set according to the ratio of 7:1:2, ensuring a balanced distribution of object specifications and grasping gestures in different sets. Evaluate the quality of the dataset through methods such as cross-validation. If data deviation or insufficiency is found, supplement the acquisition samples in a timely manner.

[0066] Step 2: Based on the semantic guidance of object specifications from large models, by integrating prompting engineering techniques with large language models such as GPT-4v, obtain rich text information covering object shape, components, dimensions, materials, and grasping gesture descriptions for object specification representation learning.

[0067] 2.1 Prompting Engineering Design

[0068] Basic Prompt Construction: Design basic prompt templates for different dimensions of object specification information (shape, dimensions, material, components). For example, for shape description, use "Please describe the geometric shape of the object in detail, including contour features and special structures: Object shape annotation"; for material description, use "Analyze the material characteristics of the object and its impact on grasping based on the following information: Material annotation. Adjust the expression of the prompt, keyword weights, and instruction complexity through experiments to obtain more accurate model outputs.

[0069] Task-Oriented Prompt Optimization: Combine the requirements of the grasping task to design task-oriented prompts. Guide the model to generate semantic information directly related to the grasping operation. At the same time, add constraint condition prompts to improve the practicality of the generated content.

[0070] Multi-Round Interaction Prompt Strategy: Adopt a multi-round prompt interaction mechanism. In the first round, input basic object specification information to obtain a preliminary description. In the second round, ask questions about the ambiguous or key information in the output of the first round. Through iterative interaction, guide the model to output more detailed and accurate semantic content.

[0071] 2.2 Interaction with Large Models such as GPT-4v

[0072] Interface Call and Parameter Configuration: Connect to the GPT-4v model through the API interface provided by OpenAI, and set appropriate parameters, such as temperature (controlling the randomness of the output, value range 0-1, a lower temperature generates more definite content), maximum number of tokens (limiting the length of the output text), frequency penalty, and presence penalty (reducing duplicate content), etc. Adjust the parameters according to the task requirements to balance the generation efficiency and diversity while ensuring the quality of the generated content.

[0073] Multi-modal Information Input: Leveraging the multi-modal processing capabilities of GPT-4v, in addition to inputting object specification information in text form, rendered images of the object's 3D model, close-up pictures of local details, material texture images, etc. are used as supplementary inputs. Metadata (such as key part descriptions, dimension annotation boxes) is added to the images through image annotation tools to help the model better understand object features and generate more comprehensive semantic descriptions.

[0074] Result Screening and Quality Assessment: Automatically and manually screen the text information output by the model. Automatic screening filters out content that does not meet requirements by setting rules such as keyword matching and grammar checking; manual evaluation is performed by domain experts scoring the screened results, and evaluation metrics include semantic accuracy, completeness, relevance to the grasping task, etc., and high-score high-quality results are retained.

[0075] 2.3. Text Information Processing and Feature Extraction

[0076] Text Cleaning and Normalization: Clean the screened text, remove redundant characters, special symbols, and duplicate content, and unify the text format (such as unifying units, standardizing term expressions). Use natural language processing tools (such as NLTK, spaCy) for word segmentation, part-of-speech tagging, and named entity recognition to extract key semantic information.

[0077] Text Embedding and Feature Fusion: Use a pre-trained text embedding model (such as BERT, RoBERTa) to convert the cleaned text into a vector representation to obtain the text features of the object specifications. At the same time, fuse the text features with the object geometric features, material features, etc. extracted in the first step, and a more comprehensive object specification representation vector can be constructed through concatenation, weighted summation, etc., for subsequent object specification representation learning and grasping model training.

[0078] Step 3: Construction of a grasping generation model for object specification perception, which consists of a semantic guidance module and a grasping gesture generation module.

[0079] The semantic guidance module in Step 3 extracts the features of each object as follows:

[0080] 3.1a. Obtain the semantic geometric features of the object using a 3D shape feature encoder;

[0081] Point cloud data preprocessing: Obtain point clouds from the three-dimensional scan data of an object. These point cloud data may contain noise, outliers, and uneven density. First, use statistical filtering methods to calculate the distances between each point and its neighboring points, set appropriate distance thresholds, and remove outliers that deviate from the normal range; adopt voxel downsampling technology to reduce the density of point cloud data while maintaining the geometric features of the object, reducing the subsequent computational workload; use bilateral filtering algorithms to smooth the data, removing noise while retaining the detailed features of the object surface.

[0082] Selection and configuration of three-dimensional shape feature encoder: Select Point Transformer as the three-dimensional shape feature encoder. It can effectively capture local and global features in point cloud data, and by grouping and feature transformation of the input point cloud, it can explore the geometric relationships between point clouds. When configuring the model, according to the scale and complexity of the object point cloud data, reasonably set parameters such as the number of groups, the number of points in each group, and the feature dimension. For example, for objects with complex shapes, appropriately increase the number of groups and feature dimensions to more detailedly describe the geometric features of the object; while for objects with simple shapes, the corresponding parameters can be reduced to improve computational efficiency.

[0083] Feature extraction and encoding: Input the preprocessed point cloud data into the configured Point Transformer. The model first encodes the position information of each point, and maps the three-dimensional coordinates of the point to a higher-dimensional feature space through a multi-layer perceptron (MLP). Then, use the attention mechanism to calculate the relationship weights between each point and its surrounding neighboring points, enabling the model to focus on key areas and enhancing the ability to capture the geometric features of the object. After multiple layers of feature extraction and transformation, finally output a vector that can represent the semantic geometric features of the object. This vector contains key geometric information such as the shape, size, and surface curvature of the object, providing important geometric basis for subsequent grasp generation.

[0084] 3.2a. With the help of the large language model LLM, interpret the specification semantic instructions, and obtain the object specification features through text embedding;

[0085] Prompt design and optimization: Design targeted prompts according to the type of object and the requirements of the grasping task. Through experiments and analysis, continuously adjust the expression of the prompts, the selection of keywords, and the detail level of the instructions to guide the large language model to output more accurate and rich information. For example, add a guiding language such as "Please describe from the perspective of easy grasping" to the prompt to make the model's answer more suitable for the grasping task.

[0086] Interact with the large language model: Use large language models such as GPT-4v and input the designed prompts into the model. During the interaction, reasonably set the parameters of the model. For example, the temperature parameter controls the randomness of the output. A lower temperature value makes the output more certain and more in line with common patterns; the maximum number of tokens limits the length of the output text to avoid information redundancy caused by overly long output. Obtain the descriptive text generated by the model regarding the shape, components, dimensions, materials, etc. of the object.

[0087] Text embedding and feature extraction: Use the pre-trained BERT-large-uncased model to process the text output by the large language model. First, tokenize the text, splitting the continuous text sequence into individual words or tokens; then perform part-of-speech tagging and syntactic analysis on each token to understand the syntactic structure of the text. Next, extract the [CLS] token in each sentence. This token in the BERT model can comprehensively represent the semantic information of the entire sentence. Apply max pooling operation to the extracted [CLS] tokens to extract the most representative features from the [CLS] tokens of multiple sentences, generating a 1×C semantic feature vector. This vector integrates the large language model's understanding of the semantics of object specifications, contains rich prior knowledge, and serves as an important part of the object specification features.

[0088] 3.3a. Design a modality adapter to fuse the object and specification text features and obtain a general feature representation applicable to objects of different specifications.

[0089] Specific implementation of the 3D shape feature encoder in step 3.1a:

[0090] 3.1.1. Input data: The object is represented in the form of a point cloud, containing the 3D coordinates (X, Y, Z) of 2048 points, which can be collected through a depth camera (such as Intel RealSense) or a simulation environment (such as NVIDIA IsaacGym).

[0091] The point cloud needs to be normalized to the range of the unit cube (-1, 1) to eliminate scale differences.

[0092] 3.1.2. Encoder selection and processing: Use Point Transformer as the backbone network. Its core is the multi-head self-attention mechanism, which can capture the local and global geometric relationships of the point cloud. The network structure consists of 4 layers of Transformer blocks, and each layer outputs 256-dimensional features. Perform farthest point sampling (FPS) grouping on the point cloud, dividing it into 64 local regions, and extracting 256-dimensional features from each region. Finally, aggregate all local features through max pooling to generate global geometric features (256-dimensional).

[0093] 3.1.3. Output: A matrix with geometric features of 64×256 (64 local region features) and a global feature vector of 1×256.

[0094] LLM Interaction and Semantic Parsing in Step 3.2:

[0095] 3.2.1. Model Input Format: Convert the structured instruction into key-value pair text, with the delimiter being a semicolon: "Action: Pinch; Target Part: Upper Part of the Cup; Constraint: Three-Finger Contact, Tilt Angle 30°; Material: Ceramic"

[0096] 3.2.2. Large Model Invocation (Taking OpenAI as an Example)

[0097] 1. API Request Construction:

[0098] Model Selection: `text-embedding-3-large` (3072-dimensional output)

[0099] Input Text: The above key-value pair string

[0100] Parameter Setting: Turn off sensitive content filtering (`filter = false`)

[0101] 2. Response Parsing: Extract the returned `embedding` field to obtain a 3072-dimensional floating-point array

[0102] Typical Value Example: [0.024, -0.157,..., 0.382]` (each dimension ranges from [-1, 1])

[0103] Localization Alternative: Use an open-source model (such as `bge-small-en-v1.5`): Input the same text, and output a 384-dimensional vector. An additional training adaptation layer (MLP) is required to map the 384 dimensions to 256 dimensions

[0104] 3. Post-Processing of Text Embedding:

[0105] ① Feature Dimensionality Reduction (Optional): If the original embedding dimension is high (such as 3072 dimensions), reduce it to 256 dimensions through PCA. Calculate the covariance matrix of the dataset and retain the eigenvectors corresponding to the first 256 largest eigenvalues

[0106] ② Normalization: Perform Min-Max normalization on each dimension and scale it to the [0, 1] interval

[0107] Example: If the original range of a certain dimension is [-0.5, 1.2]

[0108] Normalization Formula: normalized_value = (original value + 0.5) / 1.7

[0109] ③ Feature Fusion Preparation: Concatenate the processed text features with geometric features (128 - dimensional).

[0110] Direct Concatenation: Generate a 384 - dimensional vector (128 + 256).

[0111] Weighted Summation: Text features × 0.6+Geometric features × 0.4 (parameters need to be tuned through experiments).

[0112] 4. Verification and Debugging:

[0113] ① Semantic Similarity Test, calculate the cosine similarity of different instruction embeddings:

[0114] "Pinch the upper part of the cup" vs "Grasp the top of the cup with three fingers" → should be ≥ 0.85

[0115] "Press the button" vs "Rotate the bottle cap" → should be ≤ 0.3

[0116] ② Feature Visualization: Use t - SNE to reduce the 256 - dimensional text features to a 2D plane:

[0117] Instructions of the same type (such as different expressions of "rotation action") should be clustered

[0118] Instructions of different types (such as "rotation" and "press") should be clearly separated.

[0119] How to design the modality adapter in Step 3.3:

[0120] 3.3.1. Input Feature Pre - processing

[0121] 1. Geometric Feature Standardization: 128 - dimensional object geometric features from PointNet++:

[0122] Perform Z - score standardization on each dimension (mean to zero, standard deviation is 1)

[0123] Handle outliers: Truncate to within the range of [- 3, + 3] standard deviations

[0124] 2. Text Feature Normalization: 256 - dimensional text embeddings from LLM:

[0125] Perform Min - Max normalization, linearly map each dimension to the interval [0, 1]

[0126] Eliminate sparse dimensions (maximum value < 0.1), retain 200 - dimensional effective features.

[0127] 3.3.2. Cross - Modality Fusion Architecture:

[0128] 1. Feature Projection Layer (Dimension Alignment)

[0129] Geometric Feature Pathway: Through a fully connected layer (128→256 dimensions) + LayerNorm

[0130] Activation Function: GELU (Gaussian Error Linear Unit)

[0131] Text Feature Pathway: Through a fully connected layer (256→256 dimensions) + Dropout (with a probability of 0.2)

[0132] Activation Function: Swish(x * sigmoid(x))

[0133] 2. Cross-Attention Mechanism: 1. Query / Key / Value Generation:

[0134] Query: Projected geometric features (256 dimensions)

[0135] Key / Value: Projected text features (256 dimensions)

[0136] 3. Multi-Head Attention Calculation: Number of heads: 4 heads, 64 dimensions per head

[0137] Attention Weight Calculation: Scaled Dot-Product

[0138] Temperature Coefficient: √64 = 8 (used to scale the dot-product result)

[0139] 4. Output Merging: After concatenating the outputs of each head, pass through a linear layer (256→256 dimensions)

[0140] Residual Connection: Fused pre-geometry features + attention output

[0141] Gated Fusion Module: Design a dynamic weight gate:

[0142] Gating Value = σ(Wg · [Geometric Features; Text Features] + bg)

[0143] Output = Gating Value × Geometric Features + (1 - Gating Value) × Attention Output

[0144] σ: Sigmoid Function

[0145] Wg: Learnable weight matrix (512×1 dimensions)

[0146] 5. Output Layer Design: ① Feature Enhancement: Further refine through two layers of MLP:

[0147] First Layer: 256→512 dimensions, GELU activation

[0148] Second layer: 512→256 dimensions, linear output. Skip connection: Add the input of the fusion module and the output of the MLP and normalize. Normalization and constraint: Execute before output: LayerNorm (prevent feature scale drift)

[0149] L2 normalization (limit the norm of the feature vector to 1)

[0150] 6. Training strategy:

[0151] ① Loss function combination: Main loss: Contrastive Loss

[0152] Distance between positive sample pairs (same object + matching instruction) ≤ 0.1

[0153] Distance between negative sample pairs (different objects / conflicting instructions) ≥ 0.8

[0154] Auxiliary loss: Modal alignment loss (MMD): Reduce the difference in geometric / text feature distributions

[0155] Reconstruction loss: The decoder reconstructs the point cloud (supervise feature interpretability)

[0156] ② Optimization parameters: Optimizer: AdamW Initial learning rate: 3e-5 Weight decay: 1e-6 Batch size: 64 Learning rate scheduling: Cosine annealing (minimum 1e-6)

[0157] 7. Key implementation details:

[0158] ① Gradient clipping, set the global gradient norm threshold: 1.0, scale the gradient proportionally when it exceeds the threshold

[0159] ② Feature visualization monitoring: Use TensorBoard to display in real time: Heat map of attention weights (4-head attention distribution) Histogram of gating values (statistics of the distribution in the range of 0 to 1)

[0160] ③ Typical failure handling: Problem 1: Modal dominance imbalance (e.g., text features cover geometric information)

[0161] Solution:

[0162] 1) Reduce the Dropout rate of the text path to 0.1

[0163] 2) Increase the weight of the geometric feature reconstruction loss (×2.0)

[0164] Problem 2: Attention divergence (no difference in multi-head weights) Solution:

[0165] 1) Add attention diversity loss (encourage differences between heads)

[0166] 2) Add random noise (standard deviation 0.01) to the Key vector

[0167] 8. Output verification criteria: Geometric consistency test:

[0168] When different text instructions are applied to the same object, the generated features should be maintained:

[0169] Same component description (such as "handle") → Cosine similarity ≥ 0.9

[0170] Contradictory instructions (such as "press" vs "rotate") → Similarity ≤ 0.3

[0171] Cross-modal retrieval accuracy: For text → geometric retrieval, the Top-1 accuracy should be ≥ 85%

[0172] For geometric → text retrieval, the Top-3 accuracy should be ≥ 90%

[0173] The steps of the grasping gesture generation module in step 3, which is also the method for generating robotic hand parameters, are as follows:

[0174] 3.1b. With the aid of a pre-set noise schedule

[0175] Here, T represents the total number of steps in the diffusion process, and β t is the correlation coefficient of the noise added at each step.

[0176] Starting from the initial parameter J0 = J′ i and following the Markov chain rule: Add noise step by step. This means that at each step t, the current state J t is obtained by adding noise that follows a specific Gaussian distribution to the previous state J t-1 . As the steps progress, J t gradually approaches Gaussian noise where represents a Gaussian distribution with a mean of 0 and a covariance matrix of the identity matrix I. Simply put, it makes the initial parameter more and more "chaotic" and turns it into a pure noise state.

[0177] 3.2b. The reverse denoising process is described by .

[0178] Here, p θ represents the probability distribution of the previous state J t given the parameter θ, the current noisy state J g and the grasping target-related information F t-1 . It is also a Gaussian distribution, and its mean is determined by μ θ ((J t , t, F g ) and the covariance is determined by ∑ t (usually determined according to specific settings).

[0179] Mean function:

[0180] where α t = 1 - β t is the noise attenuation coefficient, which is used to measure the degree of noise reduction at each step; is the cumulative attenuation factor, which accumulates the noise attenuation from the first step to the t-th step; ∈ θ is the noise prediction network obtained through training. It predicts the noise according to the current noisy state J t , step t, and the grasping target information F g to predict the noise. Through this formula, using the prediction result of the noise prediction network, the noisy state J t is gradually corrected to approach the true grasping gesture parameters.

[0181] 3.3b. Noise prediction network ∈ θ is trained relying on the true noise term ∈ t calculated during the forward diffusion process for supervision. Specifically, it is achieved by minimizing the loss function .

[0182] where denotes taking the expectation, and ||·|| 2 denotes the square of the Euclidean norm. This loss function measures the gap between the noise ∈ θ (J t , t, F g ) predicted by the noise prediction network and the true noise ∈ t . During the training process, the network parameters θ are continuously adjusted to make the loss function value as small as possible, so that the noise prediction network can predict the noise more and more accurately, and then play a better role in the reverse denoising process to help generate high-quality grasping gestures.

[0183] This invention is expected to significantly improve the semantic understanding ability and generation quality of the grasping gesture generation model for object specification perception, enhance the generalization ability for multi-type specification objects, provide more effective support for robots in complex and diverse object grasping scenarios, contribute to the development of fields such as industrial automation, logistics sorting, and intelligent services, and promote the robot grasping technology to reach a new level of higher precision and stronger adaptability

[0184] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for grasping multi-specification target objects based on large models, characterized in that, It includes the following steps: Step 1: Construct a grasping gesture dataset for multi-specification objects, that is, construct a grasping dataset composed of object shape, components, dimensions, materials, object specification information, and robotic hand grasping gestures; Step 2: Perform object specification semantic guidance based on a large model, that is, fuse prompt engineering techniques and language large models to obtain text information describing object shape, components, dimensions, materials, and grasping gestures for object specification representation learning; Step 3: Construct a grasping generation model that perceives object specifications. The model consists of a semantic guidance module and a grasping gesture generation module. Generate robotic hand parameters through the grasping generation model. The grasping gesture generation module gradually perturbs the initial parameters into Gaussian noise through forward diffusion and denoises them reversely. Use the noise prediction network to gradually correct the noisy state to approach the real grasping gesture parameters. Finally, supervise and train the noise prediction network with the noise true value term obtained through forward diffusion.

2. The method for grasping multi-specification target objects based on a large model according to claim 1, wherein, In Step 1, the process of constructing a grasping gesture dataset for multi-specification objects includes: Step 1-1: Label the shape, dimensions, materials, and components of the object to construct the specification text information of each object; Step 1-2: Adopt an optimized sampling method and combine it with interactive simulation software and human experience to collect grasping gestures with different degrees of freedom that are adapted to the object specification semantics.

3. The method for grasping multi-specification target objects based on a large model according to claim 2, wherein, In Step 1-1, the shape labeling of the object includes: obtaining the point cloud data of the object using a high-precision 3D scanner, denoising and smoothing it through software such as MeshLab and CloudCompare, and then constructing a 3D mesh model of the object; manually marking the key geometric feature points of the object, including edge vertices, curvature mutation points, and symmetry axis positions, and structurally labeling the object shape type and geometric feature parameters through a labeling tool; The dimension labeling of the object includes: using a ranging device to perform multi-dimensional measurements on the dimension parameters of the object. For deformable objects, record their dimension ranges in different states and indicate the measurement conditions during labeling.

4. The method for grasping multi-specification target objects based on a large model according to claim 2, wherein, In Step 1-2, collecting grasping gestures with different degrees of freedom that are adapted to the object specification semantics includes simulation environment construction, human experience-driven acquisition, and automated assisted acquisition simulation environment construction; The process of simulation environment construction is: in the interactive simulation software, import the labeled 3D model of the object and various types of robotic hand models, configure simulation parameters that conform to real physical laws, including gravitational acceleration, friction, and collision detection thresholds. At the same time, set up a virtual camera to record the images and video data of the grasping process; The process of automated assisted acquisition is: develop a rule-based automated acquisition script, automatically generate a series of grasping action sequences according to the object specification characteristics. At the same time, use reinforcement learning algorithms to perform autonomous exploration in the simulation environment, optimize the grasping strategy through a reward mechanism, and supplement edge case samples that are difficult to cover by manual acquisition.

5. The method for grasping multi-specification target objects based on a large model according to claim 1, wherein In Step 2, the process of performing object specification semantic guidance based on a large model, that is, fusing prompt engineering techniques and language large models to obtain text information describing object shape, components, dimensions, materials, and grasping gestures for object specification representation learning includes: Step 2-1: Prompt engineering design, including basic prompt construction, task-oriented prompt optimization, and multi-round interaction prompt strategy. The basic prompt construction is to design basic prompt templates for different dimensions of object specification information. The task-oriented prompt optimization is to design task-oriented prompts in combination with the requirements of the grasping task, guiding the model to generate semantic information directly related to the grasping operation, and adding constraint condition prompts to improve the practicality of the generated content. The multi-round interaction prompt strategy is to adopt a multi-round prompt interaction mechanism. In the first round, the basic object specification information is input to obtain a preliminary description. In the second round, questions are asked about the fuzzy or key information in the output of the first round, and the model is guided to output semantic content through iterative interaction. Step 2-2: Interaction with the GPT-4v large model, including interface call and parameter configuration, multi-modal information input, and result screening and quality assessment. The interface call and parameter configuration is to connect with the GPT-4v model through the API interface provided by OpenAI and set parameters. The multi-modal information input is to utilize the multi-modal processing ability of GPT-4v. In addition to inputting the object specification information in text form, the rendered image of the object's three-dimensional model, the close-up picture of local details, and the texture image are used as supplementary inputs. Metadata is added to the images through an image annotation tool. The result screening and quality assessment is to conduct dual screening of the text information output by the model, including automatic and manual screening. Step 2-3: Text information processing and feature extraction, including text cleaning and normalization, text embedding and feature fusion. The text cleaning and normalization is to clean the screened text, remove redundant characters, special symbols, and duplicate content, unify the text format, and perform word segmentation, part-of-speech tagging, and named entity recognition using natural language processing tools to extract key semantic information. The text embedding and feature fusion is to convert the cleaned text into a vector representation using a pre-trained text embedding model to obtain the text features of the object specification, and fuse the text features with the object geometric features, material features, etc. extracted in Step 1.

6. The method for grasping multi-specification target objects based on a large model according to claim 5, wherein In Step 2-2, the process of interacting with the GPT-4v large model includes: Interface call and parameter configuration: Connect with the GPT-4v model through the API interface provided by OpenAI, set parameters, including temperature, maximum tokens, frequency penalty, and presence penalty, and adjust the parameters according to the task requirements to balance the generation efficiency and diversity while ensuring the quality of the generated content. Multi-modal information input: Utilize the multi-modal processing ability of GPT-4v. In addition to inputting the object specification information in text form, the rendered image of the object's three-dimensional model, the close-up picture of local details, and the texture image are used as supplementary inputs. Metadata is added to the images through an image annotation tool to help the model better understand the object features and generate a more comprehensive semantic description. Result screening and quality assessment: Conduct dual screening of the text information output by the model. Automatic screening filters out the content that does not meet the requirements by setting keyword matching and grammar checking rules.

7. The method for grasping multi-specification target objects based on a large model according to claim 1, wherein In step three, the semantic guidance module obtains the object semantic geometric features by using the three-dimensional shape feature encoder, interprets the specification semantic instructions with the help of the large language model LLM, obtains the object specification features through text embedding, and designs a modality adapter to fuse the object and specification text features to obtain a general feature representation.

8. The method for grasping multi-specification target objects based on a large model according to claim 4, wherein In step three, the semantic leading module processes the extraction of each object feature as follows: Step 3-1a: Use the three-dimensional shape feature encoder to obtain the object semantic geometric features; Step 3-2a: Interpret the specification semantic instructions with the help of the large language model LLM, and obtain the object specification features through text embedding; Step 3-3a; Design a modality adapter to fuse the object and specification text features to obtain a general feature representation applicable to objects of different specifications.

9. The method for grasping multi-specification target objects based on a large model according to claim 1, wherein, In step one, the process of generating the robotic arm parameters by the grasping generation model includes: Step 3-1b: With the aid of a preset noise schedule starting from the initial parameter J0 = J i ′, according to the Markov chain rule: gradually add noise. At each step t, the current state J t is obtained from the previous state J t-1 by adding noise that follows a specific Gaussian distribution. As the steps progress, J t gradually approaches Gaussian noise where represents a Gaussian distribution with a mean of 0 and a covariance matrix of the identity matrix I; Step 3-2b: The reverse denoising process is described by where p θ represents the probability distribution of the previous state J t given the parameters θ, the current noisy state J g and the grasping target-related information F t-1 . It is also a Gaussian distribution, whose mean is determined by μ θ ((J t , t, F g ) and the covariance is determined by ∑ t ; The mean function is expressed as: where α t = 1 - β t is the noise attenuation coefficient, which is used to measure the degree of noise reduction at each step; is the cumulative attenuation factor, which accumulates the noise attenuation from the first step to the t-th step; ∈ θ is the noise prediction network obtained through training, which predicts the noise based on the current noisy state J t , step t, and the grasping target information F g to predict the noise; Through the mean function Formula , the prediction result of the noise prediction network is used to gradually correct the noisy state J t , approaching the true grasping gesture parameters; Step 3-3b: The noise prediction network ∈ θ is trained relying on the true noise terms calculated during the forward diffusion process ∈ t for supervision, specifically by minimizing the loss function which is achieved, where denotes taking the expectation, ||·|| 2 denotes the square of the Euclidean norm. This loss function measures the noise ∈ θ (J t , t, F g ) predicted by the noise prediction network and the true noise ∈ t . During the training process, the network parameters θ are continuously adjusted to make the value of the loss function as small as possible.

Citation Information

Patent Citations

  • Unmanned laboratory-oriented transparent object grabbing method

    CN118386230A

  • Man-machine co-fusion mechanical arm self-adaptive grabbing method and system based on multi-modal large model

    CN118789551A

  • Robot grabbing detection method based on visual language action multi-mode alignment strategy

    CN119526405A

  • Grasping of an object by a robot based on grasp strategy determined using machine learning model(s)

    WO2019136124A1

Cited By

  • Robot grabbing control method and system

    CN121083655A

  • Robot lightweight retrieval and grabbing method for human-computer interaction, medium and equipment

    CN121234066A