Large model-based multi-specification target object grasping method

By constructing a multi-size object grasping gesture dataset and large-scale model semantic guidance, the problem that traditional grasping methods are difficult to handle complex and varied objects is solved, achieving efficient and highly adaptable object grasping, and improving the stability and generalization ability of grasping technology.

CN120244978BActive Publication Date: 2026-01-13SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510587079.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2026-01-13
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Traditional object grasping methods struggle to handle complex and varied object sizes, especially when no object is visible. They require significant manpower and resources to adjust parameters and optimize models, and it is difficult to develop effective grasping solutions.

Method used

We construct a dataset of grasping gestures for objects of various sizes. We use a large model to provide semantic guidance on object size, integrate prompt word engineering technology with a large language model, and obtain textual information on object shape, components, size, material, and grasping gesture descriptions. We use a noise prediction network to gradually correct noisy states, generate robot arm parameters, and achieve efficient grasping.

Benefits of technology

It breaks through the limitations of traditional grasping methods, improves the stability and generalization ability of grasping methods, can efficiently deal with complex and ever-changing objects and unseen objects, provides data support for dexterous hand grasping datasets, and enhances the adaptability of grasping technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244978B_ABST
    Figure CN120244978B_ABST
Patent Text Reader

Abstract

The application provides a multi-specification target object grasping method based on a large model, which comprises three steps of constructing a grasping gesture dataset for multi-specification objects, performing object specification semantic guidance based on a large model (that is, fusing a prompt word engineering technology and a language large model to obtain text information of object shape, components, size, material and grasping gesture description for object specification characterization learning), and constructing a grasping generation model with object specification perception. The application relies on advanced large model technology and the unique generation capability of a diffusion model, and is committed to proposing a grasping mechanism guided by object specification information, and constructing a generation model with generalization and object specification perception, and focuses on solving the key work of grasping gesture generation of different specification objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent robot target grasping technology, specifically a method for grasping multi-size target objects based on a large model. Background Technology

[0002] Intelligent robots represent a commanding height in the competition of the technology industry and play a crucial role in my country's industrial intelligent transformation. Object grasping is an important step in many scenarios, but the diverse characteristics of objects present challenges. Traditional grasping methods have obvious limitations. When dealing with complex and variable objects, a lot of manpower and resources are needed to adjust parameters and optimize models, and it is difficult to formulate grasping plans for objects that have never been seen.

[0003] Large models demonstrate powerful capabilities across multiple domains, and their massive data learning and generalization capabilities make it possible to break through the limitations of traditional crawling methods. Therefore, the key to solving the above problems lies in how to crawl datasets containing object specification information, utilize the massive data learning and generalization capabilities of large models, and construct a generative model that combines generalization and object specification perception. Summary of the Invention

[0004] This invention proposes a multi-size target object grasping method based on a large model. Its purpose is to effectively address the challenges brought about by the size differences of objects when grasping them. It solves the problem that traditional grasping methods are unable to handle complex and variable objects and unseen objects. It opens up a new path for achieving intelligent, efficient and highly adaptable object grasping and provides an innovative solution of great value.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] A method for grasping multi-size target objects based on large models includes the following steps:

[0007] Step 1: Construct a grasping gesture dataset for objects of various sizes, that is, construct a grasping dataset consisting of object shape, parts, size, material, object specification information and robotic hand grasping gestures;

[0008] Step Two:

[0009] Object specification semantic guidance is based on a large model. That is, the prompt word engineering technology and language large model are integrated to obtain text information describing the object's shape, components, size, material and grasping gesture, which is used for object specification representation learning.

[0010] Step 3: Construct an object size-aware grasping generation model. The model consists of a semantic guidance module and a grasping gesture generation module. The grasping generation model generates robot hand parameters. The grasping gesture generation module gradually perturbs the initial parameters into Gaussian noise through forward diffusion and performs inverse denoising. The noise prediction network is used to gradually correct the noisy state to approach the real grasping gesture parameters. Finally, the noise prediction network is trained under the supervision of the noise ground truth terms obtained by forward diffusion.

[0011] Preferably, in step one, the process of constructing a grasping gesture dataset for objects of various sizes includes:

[0012] Step 1-1: Label the shape, size, material, and components of the objects to create specification text information for each object;

[0013] Steps 1-2: Using optimized sampling methods and combining interactive simulation software with human experience, collect grasping gestures that are adapted to the semantics of object specifications and have different degrees of freedom.

[0014] Preferably, in step 1-1, the shape annotation of the object includes: acquiring the point cloud data of the object using a high-precision 3D scanner, denoising and smoothing it using MeshLab and CloudCompare software, and then constructing a 3D mesh model of the object; manually marking the key geometric feature points of the object, including edge vertices, curvature change points, and symmetry axis positions, and using annotation tools to structurally annotate the object's shape type and geometric feature parameters.

[0015] The dimensioning of an object includes: using a distance measuring device to measure the object's dimensional parameters in multiple dimensions; for deformable objects, recording the dimensional range under different states, and specifying the measurement conditions during annotation.

[0016] Preferably, in steps 1-2, collecting grasping gestures that are adapted to the semantics of the object's specifications and have different degrees of freedom includes simulation environment construction, human experience-driven collection, and automated assisted collection simulation environment construction.

[0017] The process of setting up the simulation environment is as follows: In the interactive simulation software, import the labeled 3D model of the object and various types of robot arm models, configure simulation parameters that conform to the laws of real physics, including gravitational acceleration, friction, and collision detection threshold, and set up a virtual camera to record the image and video data of the grasping process.

[0018] The automated assisted acquisition process is as follows: develop rule-based automated acquisition scripts, automatically generate a series of grasping action sequences based on the object's specifications and characteristics, and simultaneously use reinforcement learning algorithms to autonomously explore in a simulation environment, optimize the grasping strategy through a reward mechanism, and supplement edge case samples that are difficult to cover by manual acquisition.

[0019] Preferably, in step two, the semantic guidance for object specifications is based on a large model. That is, by integrating prompt word engineering technology with a large language model, textual information describing the object's shape, components, size, material, and grasping gesture is obtained. The process of learning object specification representation includes:

[0020] Step 2-1: Prompt word engineering design, including basic prompt word construction, task-oriented prompt word optimization, and multi-round interactive prompt word strategy. Basic prompt word construction involves designing basic prompt word templates for different dimensions of object specification information. Task-oriented prompt word optimization involves designing task-oriented prompt words based on the requirements of the grasping task, guiding the model to generate semantic information directly related to the grasping operation, and adding constraint prompt words to improve the usability of the generated content. The multi-round interactive prompt word strategy employs a multi-round prompt word interaction mechanism. In the first round, basic object specification information is input to obtain a preliminary description; in the second round, follow-up questions are asked regarding ambiguous or key information in the first round's output, guiding the model to output semantic content through iterative interaction.

[0021] Step 2-2: Interacting with the GPT-4v large model, including interface calls and parameter configuration, multimodal information input, and result filtering and quality assessment; the interface calls and parameter configuration refer to connecting with the GPT-4v model and setting parameters through the API provided by OpenAI; the multimodal information input refers to utilizing the multimodal processing capabilities of GPT-4v to input not only the object specification information in text form, but also the object's 3D model rendering image, close-up images of local details, and material texture images as supplementary inputs, and adding metadata to the images through image annotation tools; the result filtering and quality assessment refer to the automatic and manual dual filtering of the text information output by the model;

[0022] Steps 2-3: Text information processing and feature extraction, including text cleaning and normalization, text embedding and feature fusion. Text cleaning and normalization involves cleaning the filtered text, removing redundant characters, special symbols and duplicate content, unifying the text format, and using natural language processing tools for word segmentation, part-of-speech tagging and named entity recognition to extract key semantic information. Text embedding and feature fusion involves using a pre-trained text embedding model to convert the cleaned text into a vector representation, obtaining text features of the object's specifications, and fusing these text features with the object's geometric features, material features, etc., extracted in Step 1.

[0023] Preferably, in step 2-2, the process of interacting with the GPT-4v large model includes:

[0024] Interface calls and parameter configuration: Connect to the GPT-4v model through the API interface provided by OpenAI, set parameters, including temperature, maximum number of tokens, frequency penalty and existence penalty, and adjust parameters according to task requirements to balance generation efficiency and diversity while ensuring the quality of generated content;

[0025] Multimodal Information Input: Leveraging the multimodal processing capabilities of GPT-4v, in addition to textual object specification information, the input also includes rendered images of the object's 3D model, close-up images of local details, and material texture images as supplementary inputs. Metadata is added to the images using image annotation tools to help the model better understand object features and generate more comprehensive semantic descriptions.

[0026] Results screening and quality assessment: The text information output by the model is screened by both automatic and manual methods. Automatic screening filters out content that does not meet the requirements by setting keyword matching and grammar checking rules.

[0027] Preferably, in step three, the semantic guidance module obtains the semantic geometric features of the object by using a three-dimensional shape feature encoder, interprets the specification semantic instructions with the help of a large language model (LLM) and obtains the object specification features through text embedding, and designs a modal adapter to fuse the object and specification text features to obtain a general feature representation.

[0028] Preferably, in step three, the semantic guidance module extracts the features of each object as follows:

[0029] Step 3-1a: Obtain the semantic geometric features of the object using a 3D shape feature encoder;

[0030] Step 3-2a: Use a large language model (LLM) to interpret specification semantic instructions and obtain object specification features through text embedding;

[0031] Step 3-3a: Design a modal adapter to fuse object and specification text features to obtain a universal feature representation applicable to objects of different specifications.

[0032] Preferably, in step one, the process of generating robot arm parameters by capturing and generating the model includes:

[0033] Step 3-1b: Using pre-set noise scheduling From the initial parameter J0 = J′ t Initially, following the Markov chain rules: Noise is added gradually, and at each step t, the current state J... t It is from the previous state J t-1 Obtained by adding noise that follows a specific Gaussian distribution, J... t Gradually approaching Gaussian noise in This represents a Gaussian distribution with a mean of 0 and a covariance matrix of I.

[0034] Step 3-2b: Reverse denoising process using Description, where p θ This indicates that, given the parameter θ and the current noisy state J t And capture target-related information F g At that time, the previous state J t-1 The probability distribution of is also a Gaussian distribution, with a mean given by μ. θ ((J t ,t,F g The covariance is determined by ∑ t Decide;

[0035] The mean function is expressed as:

[0036] Where α t =1-β t It is the noise attenuation coefficient, used to measure the degree of noise reduction at each step; It is the cumulative attenuation factor, which accumulates the noise attenuation from step 1 to step t; ∈ θ It is a trained noise prediction network that predicts noise based on the current noisy state J. t Step t and capturing target information F g To predict noise;

[0037] Through the mean function formula Using the prediction results of the noise prediction network, the noisy state J is gradually corrected. t To approximate the parameters of a real grasping gesture;

[0038] Step 3-3b: Noise prediction network ∈ θ The training relies on the noise truth term ∈ calculated during the forward diffusion process. t To supervise, specifically by minimizing the loss function. Implementation, in which To express the expectation, ||·|| 2 The loss function represents the square of the Euclidean norm, and measures the noise level predicted by the noisy prediction network. θ (J t ,t,F g ) and true noise ∈ t To minimize the difference between the two, the network parameters θ are continuously adjusted during training to make the loss function value as small as possible.

[0039] Beneficial effects:

[0040] 1. This invention leverages the advantages of large-scale model learning and generalization capabilities to overcome the limitations of traditional grasping methods in dealing with complex and variable objects and unseen objects.

[0041] 2. Compared with existing technologies, this invention enriches the dexterous hand grasping dataset, collects and annotates data and models containing information such as shape, size, and material, provides data support for grasping generation, improves the stability, generalization and ability to cope with diversity of grasping methods, helps breakthroughs in grasping technology, and lays the foundation for data. Attached Figure Description

[0042] Figure 1 This is a flowchart of the present invention;

[0043] Figure 2 Research scheme for generating grasping gestures for objects of various sizes;

[0044] Figure 3 Technical route for generating grasping gestures for objects of various sizes Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] like Figure 1 As shown, the method of the present invention mainly includes three steps: constructing a grasping gesture dataset for objects of multiple sizes, providing semantic guidance on object size based on a large model, and constructing a grasping generation model that is aware of object size.

[0047] Specifically, such as Figure 2 , Figure 3 As shown, the detailed process is as follows:

[0048] Step 1: Constructing a grasping gesture dataset for objects of various sizes. First, label the shape, size, material, and components of the objects to construct specification text information. Based on the labeled information, use an optimized sampling method, interactive simulation software, and human experience to collect grasping gestures with different degrees of freedom that are adapted to the semantics of the object specifications. Finally, construct a grasping dataset consisting of objects, object specification information, and a robotic arm.

[0049] 1.1 Object Information Collection and Labeling

[0050] Data collection: Collect objects of different specifications from multiple channels such as industrial production scenarios, smart home fields, and logistics warehousing environments, covering a variety of materials such as metal, plastic, glass, and ceramics, as well as various forms such as regular shapes (such as cuboids and cylinders), irregular shapes (such as irregular parts and sculptures), and composite objects (such as containers with handles and assembled toys), to establish a basic object database.

[0051] Shape annotation: Point cloud data of the object is acquired using a high-precision 3D scanner. After denoising and smoothing using software such as MeshLab and CloudCompare, a 3D mesh model is constructed. Key geometric feature points of the object, such as edge vertices, curvature abrupt change points, and symmetry axis positions, are manually marked. Then, annotation tools (such as CVAT) are used to structurally annotate the object's shape type (such as sphere, cone, T-shaped structure, etc.) and feature parameters (such as radius, angle, symmetry distance).

[0052] Dimensional Measurement: Using equipment such as vernier calipers, laser rangefinders, and coordinate measuring machines, multi-dimensional measurements are performed on the object's length, width, height, aperture, wall thickness, and other dimensional parameters. For deformable objects (such as hoses and fabrics), the dimensional range under different conditions (tension, compression, bending) is recorded, and the measurement conditions are noted in the annotation.

[0053] Material labeling: The chemical composition of an object is detected using spectral analysis equipment (such as a Fourier transform infrared spectrometer). Combined with tactile perception and visual observation, detailed material properties are labeled, including surface roughness, hardness grade, and coefficient of friction. Simultaneously, a material labeling system is established to differentiate between different characteristics such as "smooth hard plastic" and "frosted metal."

[0054] Component labeling: For assembled objects, disassemble them into their constituent components and label each component with its name, function, connection method (e.g., threaded connection, snap-fit ​​connection), and assembly sequence. Using the assembly module in 3D modeling software, generate a component hierarchy diagram and associate the labeling information with the components in the 3D model.

[0055] 1.2 Optimize sampling strategy design

[0056] Feature weight calculation: A multi-dimensional weight evaluation model is established based on factors such as the object's shape complexity (e.g., number of vertices, frequency of curvature change), size (e.g., the ratio of volume to the gripping range of the robotic arm), material properties (e.g., the correlation between the coefficient of friction and gripping stability), and component structure (e.g., the location of key force points). For example, for precision parts with complex shapes and small dimensions, higher sampling weights are assigned to ensure that a sufficient number of targeted gripping samples are collected.

[0057] Sampling range segmentation: The object's specification parameter space is divided into multiple sub-regions, and the objects are grouped using clustering algorithms (such as K-Means). Differentiated sampling schemes are developed for the typical characteristics of each sub-region. For example, in areas with large objects, the focus is on collecting gestures at different gripping positions, while in areas with high-friction materials, experimental samples of gripping force variations are added.

[0058] 1.3 Grasp Gesture Data Acquisition

[0059] Simulation environment setup: Import annotated 3D object models and various types of robotic hand models (such as two-finger grippers, multi-finger dexterous hands, and suction cup grippers) into interactive simulation software such as Gazebo and V-REP. Configure simulation parameters that conform to real physical laws, including gravitational acceleration, friction, and collision detection thresholds. Simultaneously, set up a virtual camera to record image and video data of the grasping process.

[0060] Human experience-driven data acquisition: Robotics experts and experienced engineers were invited to manually control a robotic arm to perform grasping tasks in a simulation environment, based on object specifications and optimized sampling strategies. During operation, by adjusting parameters such as the robotic arm's pose (position, rotation angle), joint motion trajectory, and grasping force, different grasping methods under different degrees of freedom (such as parallel grasping, envelope grasping, and suction grasping) were attempted. Data such as the robotic arm joint angles, the relative pose of the object and the robotic arm, and the success / failure status of each grasp were recorded.

[0061] Automated Assisted Data Acquisition: Develop rule-based automated acquisition scripts that automatically generate a series of grasping action sequences based on the object's specifications and characteristics. For example, for a cuboid object, the script automatically generates grasping actions from different sides and heights. Simultaneously, reinforcement learning algorithms are used to autonomously explore the simulation environment, optimizing the grasping strategy through reward mechanisms (such as grasping stability scores and grasping time), supplementing edge case samples that are difficult to cover manually.

[0062] 1.4 Dataset Integration and Validation

[0063] Data cleaning: The collected raw data is deduplicated, removing duplicate samples and invalid data (such as data from incomplete scraping processes). Visualization tools are used to check the completeness and accuracy of the data, and any incorrect or missing annotations are corrected.

[0064] Enhanced data annotation: Add extra annotations to each crawled sample, including crawl type labels (such as precise crawling, stable crawling), applicable scenario labels (such as industrial assembly, home service), risk level labels (such as easy to slip, easy to damage), etc., to improve the semantic richness of the dataset.

[0065] Dataset partitioning: The dataset was divided into training, validation, and test sets in a 7:1:2 ratio to ensure a balanced distribution of object sizes and grasping gestures across the different sets. Dataset quality was evaluated using methods such as cross-validation; if data bias or deficiencies were found, additional samples were collected promptly.

[0066] Step 2: Object specification semantic guidance based on large models. By integrating prompt word engineering technology with language large models such as GPT-4v, rich text information covering object shape, components, size, material and grasping gesture descriptions is obtained for object specification representation learning.

[0067] 2.1 Prompt Word Engineering Design

[0068] Basic prompt word construction: Basic prompt word templates are designed for different dimensions of object specification information (shape, size, material, components). For example, for shape description, "Please describe the geometry of this object in detail, including contour features and special structures: Object Shape Annotation" is used; for material description, "Analyze the material properties of the object and their impact on grasping based on the following information: Material Annotation" is used. The expression of the prompt words, keyword weights, and instruction complexity are adjusted through experiments to obtain more accurate model output.

[0069] Task-oriented prompt word optimization: Task-oriented prompt words were designed based on the requirements of the crawling task. These guide the model to generate semantic information directly related to the crawling operation. Simultaneously, constraint prompt words were added to improve the usability of the generated content.

[0070] Multi-round interactive prompt strategy: A multi-round prompt interaction mechanism is adopted. In the first round, basic object specification information is input to obtain an initial description. In the second round, follow-up questions are asked on the vague or key information in the output of the first round. Through iterative interaction, the model is guided to output more detailed and accurate semantic content.

[0071] 2.2 Interaction with large models such as GPT-4v

[0072] Interface calls and parameter configuration: Connect to the GPT-4v model through the API provided by OpenAI, and set appropriate parameters, such as temperature (controlling the randomness of the output, with a value range of 0-1; lower temperatures generate more deterministic content), maximum number of tokens (limiting the length of the output text), frequency penalty, and existence penalty (reducing duplicate content). Adjust the parameters according to the task requirements to balance generation efficiency and diversity while ensuring the quality of the generated content.

[0073] Multimodal Information Input: Leveraging the multimodal processing capabilities of GPT-4v, in addition to textual object specification information, the input also includes rendered images of the object's 3D model, close-up images of local details, and material texture images as supplementary inputs. Metadata (such as descriptions of key parts and dimension bounding boxes) is added to the images using image annotation tools, helping the model better understand object features and generate more comprehensive semantic descriptions.

[0074] Results Screening and Quality Assessment: The text information output by the model undergoes both automatic and manual screening. Automatic screening filters out content that does not meet the requirements by setting rules such as keyword matching and grammar checking; manual assessment involves domain experts scoring the screened results, with evaluation indicators including semantic accuracy, completeness, and relevance to the crawling task, retaining high-scoring and high-quality results.

[0075] 2.3 Text Information Processing and Feature Extraction

[0076] Text cleaning and normalization: The filtered text is cleaned to remove redundant characters, special symbols, and duplicate content, and the text format is standardized (e.g., standardized units and terminology). Natural language processing tools (such as NLTK and spaCy) are used for word segmentation, part-of-speech tagging, and named entity recognition to extract key semantic information.

[0077] Text embedding and feature fusion: A pre-trained text embedding model (such as BERT or RoBERTa) is used to convert the cleaned text into a vector representation to obtain textual features of object specifications. Simultaneously, these textual features are fused with the object's geometric and material features extracted in step one. Through concatenation, weighted summation, and other methods, a more comprehensive object specification representation vector is constructed for subsequent object specification representation learning and grasping model training.

[0078] Step 3: Constructing an object size-aware grasping generation model, which consists of a semantic guidance module and a grasping gesture generation module.

[0079] Step three, the semantic guidance module, involves the following steps for extracting features from each object:

[0080] 3.1a. Obtain the semantic geometric features of the object using a 3D shape feature encoder;

[0081] Point cloud data preprocessing: Point clouds are obtained from the 3D scan data of the object. These point cloud data may contain noise, outliers, and uneven density. First, statistical filtering is used to calculate the distance between each point and its neighbors, and an appropriate distance threshold is set to remove outliers that deviate from the normal range. Voxel downsampling is used to reduce the density of the point cloud data while maintaining the geometric features of the object, thus reducing the amount of subsequent computation. Bilateral filtering is used to smooth the data, removing noise while preserving the detailed features of the object's surface.

[0082] 3D Shape Feature Encoder Selection and Configuration: The Point Transformer was selected as the 3D shape feature encoder. It effectively captures local and global features in point cloud data, and by grouping and transforming the input point cloud, it uncovers the geometric relationships between points. When configuring the model, parameters such as the number of groups, the number of points in each group, and the feature dimension are set appropriately based on the scale and complexity of the object's point cloud data. For example, for objects with complex shapes, the number of groups and the feature dimension are appropriately increased to describe the object's geometric features in more detail; while for objects with simple shapes, the corresponding parameters can be reduced to improve computational efficiency.

[0083] Feature Extraction and Encoding: The preprocessed point cloud data is input into the configured Point Transformer. The model first encodes the positional information of each point, mapping the 3D coordinates of the points to a higher-dimensional feature space using a Multilayer Perceptron (MLP). Next, an attention mechanism is used to calculate the relationship weights between each point and its surrounding neighbors, enabling the model to focus on key regions and enhance its ability to capture the geometric features of objects. After multiple layers of feature extraction and transformation, a vector representing the semantic geometric features of the object is finally output. This vector contains key geometric information such as the object's shape, size, and surface curvature, providing crucial geometric basis for subsequent grasping and generation.

[0084] 3.2a. Using a large language model (LLM) to interpret specification semantic instructions, the object specification features are obtained through text embedding;

[0085] Prompt word design and optimization: Based on the type of object and the requirements of the grasping task, targeted prompt words are designed. Through experimentation and analysis, the wording, keyword selection, and level of detail are continuously adjusted to guide the large language model to output more accurate and richer information. For example, adding prompts such as "Please describe from a grasping perspective" makes the model's response more relevant to the grasping task.

[0086] Interacting with large language models: Using large language models such as GPT-4v, input the designed prompt words into the model. During the interaction, reasonably set the model parameters, such as the temperature parameter to control the randomness of the output; a lower temperature value makes the output more deterministic and more in line with common patterns; the maximum number of tokens limits the length of the output text to avoid excessively long output content leading to information redundancy. Obtain the descriptive text generated by the model regarding the object's shape, parts, size, material, etc.

[0087] Text Embedding and Feature Extraction: The pre-trained BERT-large-uncased model is used to process the text output by the large language model. First, the text is segmented, dividing the continuous text sequence into individual words or tokens. Then, part-of-speech tagging and syntactic analysis are performed on each token to understand the grammatical structure of the text. Next, the [CLS] token is extracted from each sentence; this token, in the BERT model, comprehensively represents the semantic information of the entire sentence. Max pooling is applied to the extracted [CLS] tokens to extract the most representative features from the [CLS] tokens of multiple sentences, generating a 1×C semantic feature vector. This vector integrates the large language model's understanding of object specification semantics, containing rich prior knowledge, and serves as an important component of object specification features.

[0088] 3.3a. Design a modal adapter to integrate object and specification text features to obtain a universal feature representation applicable to objects of different specifications.

[0089] The specific implementation of the 3D shape feature encoder in step 3.1a:

[0090] 3.1.1 Input data: The object is represented in the form of a point cloud, containing the three-dimensional coordinates (X,Y,Z) of 2048 points, which can be acquired by a depth camera (such as Intel RealSense) or a simulation environment (such as NVIDIA IsaacGym).

[0091] Point clouds need to be normalized to the range of a unit cube (-1,1) to eliminate scale differences.

[0092] 3.1.2 Encoder Selection and Processing: The Point Transformer is used as the backbone network, with its core being a multi-head self-attention mechanism capable of capturing both local and global geometric relationships in the point cloud. The network structure consists of four Transformer layers, each outputting 256-dimensional features. The point cloud is grouped into 64 local regions using Farthest Point Sampling (FPS), with 256-dimensional features extracted from each region. Finally, max pooling is used to aggregate all local features to generate global geometric features (256-dimensional).

[0093] 3.1.3 Output: A 64×256 matrix of geometric features (64 local region features) and a 1×256 global feature vector.

[0094] Step 3.2: LLM interaction and semantic parsing

[0095] 3.2.1 Model Input Format: Convert structured instructions into key-value pair text, separated by semicolons: "Action: Pinch; Target Part: Upper part of cup; Constraint: Three-finger contact, tilt angle 30°; Material: Ceramic"

[0096] 3.2.2 Large Model Calling (Taking OpenAI as an Example)

[0097] 1. API request construction:

[0098] Model selection: `text-embedding-3-large` (3072-dimensional output)

[0099] Input text: the above key-value pair string

[0100] Parameter settings: Disable sensitive content filtering (`filter=false`)

[0101] 2. Response parsing: Extract the returned `embedding` field to obtain a 3072-dimensional floating-point array.

[0102] Typical value example: [0.024, -0.157, ..., 0.382]` (range of each dimension [-1, 1])

[0103] Localization alternative: Use an open-source model (such as `bge-small-en-v1.5`): The input is the same text, the output is a 384-dimensional vector, requiring an additional training of an adaptation layer (MLP) to map the 384 dimensions to 256 dimensions.

[0104] 3. Post-processing of embedded text:

[0105] ① Feature dimensionality reduction (optional): If the original embedding dimension is high (e.g., 3072 dimensions), reduce it to 256 dimensions using PCA. Calculate the covariance matrix of the dataset and retain the eigenvectors corresponding to the first 256 largest eigenvalues.

[0106] ② Normalization: Perform Min-Max normalization on each dimension, scaling it to the [0,1] interval.

[0107] Example: If the original range of a certain dimension is [-0.5, 1.2]

[0108] Normalization formula: normalized_value = (original value + 0.5) / 1.7

[0109] ③ Feature fusion preparation: Concatenate the processed text features with geometric features (128 dimensions).

[0110] Direct concatenation: generates a 384-dimensional vector (128+256).

[0111] Weighted summation: Text features × 0.6 + Geometric features × 0.4 (experimental parameter tuning required)

[0112] 4. Verification and Debugging:

[0113] ① Semantic similarity test: Calculate the cosine similarity of different instruction embeddings:

[0114] "Pinch the top of the cup" vs. "Grasp the top of the cup with three fingers" → should be ≥0.85

[0115] "Press button" vs. "Rotate cap" → should be ≤0.3

[0116] ② Feature visualization: Using t-SNE to reduce the dimensionality of 256-dimensional text features to a 2D plane:

[0117] Similar instructions (such as "rotation" in different ways) should be grouped together.

[0118] Different types of instructions (such as "rotate" and "press") should be clearly separated.

[0119] How to design the modal adapter in step 3.3:

[0120] 3.3.1 Input Feature Preprocessing

[0121] 1. Geometric feature standardization: 128-dimensional object geometric features from PointNet++:

[0122] Each dimension is Z-score standardized (mean zeroed, standard deviation 1).

[0123] Handling outliers: Truncate to the range of [-3, +3] standard deviations.

[0124] 2. Text Feature Normalization: 256-dimensional text embedding from LLM:

[0125] Performing Min-Max normalization linearly maps each dimension to the [0,1] interval.

[0126] Sparse dimensions (maximum value < 0.1) are removed, retaining 200 effective features.

[0127] 3.3.2 Cross-modal fusion architecture:

[0128] 1. Feature projection layer (dimensional alignment)

[0129] Geometric feature pathway: via fully connected layer (128→256 dimensions) + LayerNorm

[0130] Activation function: GELU (Gaussian Error Linear Unit)

[0131] Text feature pathway: through a fully connected layer (256→256 dimensions) + Dropout (0.2 probability)

[0132] Activation function: Swish(x*sigmoid(x))

[0133] 2. Cross-attention mechanism: 1. Query / Key / Value generation:

[0134] Query: Projected geometric features (256 dimensions)

[0135] Key / Value: Projected text features (256 dimensions)

[0136] 3. Multi-head attention calculation: Number of heads: 4, 64 dimensions per head

[0137] Attention weight calculation: Scaled Dot-Product

[0138] Temperature coefficient: √64 = 8 (used for scaling dot product results)

[0139] 4. Output merging: The outputs from each head are spliced ​​together and then processed through a linear layer (256→256 dimensions).

[0140] Residual connection: fusing pre-geometric features + attention output

[0141] Gating fusion module: Design of dynamic weighted gating:

[0142] Gating value = σ(Wg·[geometric features; text features] + bg)

[0143] Output = Gated value × Geometric feature + (1 - Gated value) × Attention output

[0144] σ: Sigmoid function

[0145] Wg: Learnable weight matrix (512×1 dimension)

[0146] 5. Output Layer Design: ① Feature Enhancement: Further refinement through two layers of MLP:

[0147] First layer: 256→512 dimensions, GELU activation

[0148] Second layer: 512→256 dimensions, linear output. Skip connections: fusion module input and MLP output are added together. Normalization and constraints: LayerNorm is performed before output (to prevent feature scale drift).

[0149] L2 normalization (restricting the feature vector magnitude to 1)

[0150] 6. Training strategies:

[0151] ① Loss function combination: Main loss: Contrastive Loss

[0152] Positive sample pairs (same object + matching instruction) have a distance ≤ 0.1.

[0153] Negative sample pairs (different objects / conflicting commands) with a distance ≥ 0.8

[0154] Auxiliary Loss: Modality Alignment Loss (MMD): Reduces the difference in geometric / textual feature distributions.

[0155] Reconstruction Loss: Decoder Reconstructs Point Cloud (Supervised Feature Interpretability)

[0156] ② Optimization parameters: Optimizer: AdamW; Initial learning rate: 3e-5; Weight decay: 1e-6; Batch size: 64; Learning rate scheduling: Cosine annealing (minimum 1e-6).

[0157] 7. Key Implementation Details:

[0158] ① Gradient clipping: Set a global gradient norm threshold of 1.0. When the threshold is exceeded, the gradient is scaled proportionally.

[0159] ② Feature visualization and monitoring: Real-time display using TensorBoard: Attention weight heatmap (4-head attention distribution) and gating value histogram (statistical distribution of 0-1 interval).

[0160] ③ Typical failure handling: Problem 1: Modality dominance imbalance (e.g., text features covering geometric information)

[0161] Solution:

[0162] 1) Reduce the text path Dropout rate to 0.1.

[0163] 2) Increase the weight of the geometric feature reconstruction loss (×2.0)

[0164] Problem 2: Divergent attention (no difference in weighting among multiple heads) Solution:

[0165] 1) Add attention diversity loss (encourage inter-head differences)

[0166] 2) Add random noise (standard deviation 0.01) to the Key vector.

[0167] 8. Output verification standard: Geometric consistency test:

[0168] When different text commands are applied to the same object, the generated features should be preserved:

[0169] Descriptions of identical components (e.g., "handle") → Cosine similarity ≥ 0.9

[0170] Conflicting instructions (e.g., "press" vs. "rotate") → similarity ≤ 0.3

[0171] Cross-modal retrieval accuracy: Top-1 accuracy for text → geometric retrieval must be ≥85%.

[0172] Geometry → Text Retrieval Top-3 accuracy must be ≥90%

[0173] The steps of the grasping gesture generation module, i.e., the robot arm parameter generation method, in step three are as follows:

[0174] 3.1b. Using pre-set noise scheduling

[0175] Here, T represents the total number of steps in the diffusion process, and β t It is the correlation coefficient of noise added at each step.

[0176] From the initial parameter J0 = J′ i Initially, following the Markov chain rules: Noise is added gradually. This means that at each step t, the current state J... t It is from the previous state J t-1 This is obtained by adding noise that follows a specific Gaussian distribution. As the steps progress, J... t Gradually approaching Gaussian noise in This represents a Gaussian distribution with a mean of 0 and a covariance matrix of identity matrix I. Simply put, it involves making the initial parameters increasingly "random," turning it into a state of pure noise.

[0177] 3.2b, Reverse denoising process using describe.

[0178] Here p θ This indicates that, given the parameter θ and the current noisy state J t And capture target-related information F g At that time, the previous state J t-1 The probability distribution of is also a Gaussian distribution, with a mean given by μ. θ ((J t ,t,F g The covariance is determined by ∑ t Decision (generally determined based on specific settings).

[0179] Mean function:

[0180] Where α t =1-β t It is the noise attenuation coefficient, used to measure the degree of noise reduction at each step; It is the cumulative attenuation factor, which accumulates the noise attenuation from step 1 to step t; ∈ θ It is a trained noise prediction network that predicts noise based on the current noisy state J. t Step t and capturing target information F g This is used to predict noise. Using this formula, the prediction results from the noise prediction network are used to progressively correct the noisy state J. t It aims to approximate the parameters of a realistic grasping gesture.

[0181] 3.3b, Noise Prediction Network ∈ θ The training relies on the noise truth term ∈ calculated during the forward diffusion process. t To supervise. Specifically, by minimizing the loss function. accomplish.

[0182] in To express the expectation, ||·|| 2 This represents the square of the Euclidean norm. This loss function measures the noise ∈ [the range of noise predicted by the noise prediction network]. θ (J t ,t,F g ) and true noise ∈ t The difference lies in the network parameters. During training, the network parameters θ are continuously adjusted to minimize the loss function value. This allows the noise prediction network to predict noise more and more accurately, thus playing a better role in the inverse denoising process and helping to generate high-quality grasping gestures.

[0183] This invention is expected to significantly improve the semantic understanding and generation quality of object size perception grasping gesture generation models, enhance their generalization ability to multiple object sizes, and provide more effective support for robots in complex and diverse object grasping scenarios. This will contribute to the development of industrial automation, logistics sorting, intelligent services, and other fields, propelling robot grasping technology to a new level of higher precision and greater adaptability.

[0184] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for grasping a multi-specification target object based on a large model, characterized in that, The method comprises the following steps: Step 1: Construct a multi-specification object grasping gesture dataset, that is, construct a grasping dataset composed of object shape, components, size, material, object specification information and robot hand grasping gesture; Step 2: Object specification semantic guidance based on a large model, that is, fuse prompt word engineering technology and a language large model to obtain text information describing object shape, components, size, material and grasping gesture for object specification representation learning; Step 3: Construct an object specification perception grasping generation model, which is composed of a semantic guidance module and a grasping gesture generation module, generate robot hand parameters through the grasping generation model, gradually perturb the initial parameters to Gaussian noise through forward diffusion in the grasping gesture generation module and reverse denoising, gradually correct the noise state to the real grasping gesture parameter by using a noise prediction network, and finally supervise the training of the noise prediction network through the noise true value item obtained by forward diffusion; In step 2, object specification semantic guidance based on a large model, that is, fuse prompt word engineering technology and a language large model to obtain text information describing object shape, components, size, material and grasping gesture for object specification representation learning, which comprises the following steps: Step 2-1: Prompt word engineering design, including basic prompt word construction, task-oriented prompt word optimization and multi-round interaction prompt word strategy, the basic prompt word construction, that is, designing basic prompt word templates for different dimensions of object specification information; the task-oriented prompt word optimization, that is, designing task-oriented prompt words combined with the requirements of the grasping task to guide the model to generate semantic information directly related to the grasping operation, and adding constraint condition prompt words to improve the practicality of the generated content; the multi-round interaction prompt word strategy, that is, adopting a multi-round prompt word interaction mechanism, inputting basic object specification information in the first round to obtain preliminary description, and asking questions in the second round for the fuzzy or key information in the first round output, guiding the model to output semantic content through iterative interaction; Step 2-2: Interact with the GPT-4v large model, including interface calling and parameter configuration, multi-modal information input, and result screening and quality evaluation; the interface calling and parameter configuration, that is, connecting and setting parameters through the API interface provided by OpenAI and the GPT-4v model; the multi-modal information input, that is, using the multi-modal processing capability of GPT-4v, in addition to inputting text form object specification information, also inputting object three-dimensional model rendering image, local detail close-up picture, material texture image as supplementary input, adding metadata to the image through image annotation tool; the result screening and quality evaluation, that is, automatically and manually screening the text information output by the model; Step 2-3: Text information processing and feature extraction, including text cleaning and normalization, text embedding and feature fusion. The text cleaning and normalization is to clean the filtered text, remove redundant characters, special symbols and repeated content, unify the text format, and use natural language processing tools for word segmentation, part-of-speech tagging and named entity recognition to extract key semantic information. The text embedding and feature fusion is to convert the cleaned text into vector representation using a pre-trained text embedding model, obtain the text features of the object specification, and fuse the text features with the object geometric features and material features extracted in step one.

2. The large model-based multi-specification target object grasping method according to claim 1, characterized in that, In step one, the process of constructing a multi-specification object grasping gesture dataset includes: Step 1-1: Label the shape, size, material and components of the object to construct the specification text information of each object; Step 1-2: Collect grasping gestures with different degrees of freedom that are compatible with the semantic specification of the object by using an optimized sampling method, interactive simulation software and human experience.

3. The large model-based multi-specification target object grasping method according to claim 2, characterized in that, In step 1-1, the shape labeling of the object includes: using a high-precision three-dimensional scanner to obtain the point cloud data of the object, denoising and smoothing the data through MeshLab and CloudCompare software, and then constructing a three-dimensional mesh model of the object; manually marking the key geometric feature points of the object, including edge vertices, curvature discontinuity points and symmetry axis positions, and structuring the shape type and geometric feature parameters of the object through labeling tools; The size labeling of the object includes: using a distance measuring device to measure the size parameters of the object in multiple dimensions, and for deformable objects, recording the size range in different states and indicating the measurement conditions during labeling.

4. The large model-based multi-specification target object grasping method according to claim 2, wherein, In step 1-2, collecting grasping gestures with different degrees of freedom that are compatible with the semantic specification of the object includes simulation environment setup, human experience driven collection and automated auxiliary collection simulation environment setup; The process of simulation environment setup is: in the interactive simulation software, import the labeled three-dimensional model of the object and multiple types of robot models, configure simulation parameters that conform to real physical laws, including gravitational acceleration, friction and collision detection threshold, and set up virtual cameras to record image and video data of the grasping process; The process of automated auxiliary collection is: develop a rule-based automated collection script to automatically generate a series of grasping action sequences based on the specification characteristics of the object, and use reinforcement learning algorithms to explore the simulation environment autonomously, optimize the grasping strategy through a reward mechanism, and supplement the edge case samples that are difficult to cover by human collection.

5. The multi-specification target object grasping method based on a large model according to claim 1, characterized in that, In step 2-2, the process of interacting with the GPT-4v large model includes: Interface calling and parameter configuration: connect with the GPT-4v model through the API interface provided by OpenAI, set parameters including temperature, maximum token number, frequency penalty and existence penalty, adjust parameters according to task requirements, balance generation efficiency and diversity while ensuring the quality of generated content; Multi-modal information input: Utilizing the multi-modal processing capability of GPT-4v, in addition to inputting the object specification information in text form, the three-dimensional model rendering image, local detail close-up picture, and material texture image of the object are also input as supplementary input; metadata is added to the image through an image annotation tool to help the model better understand the object features and generate more comprehensive semantic descriptions; Result screening and quality evaluation: The text information output by the model is automatically and manually screened, and the automatic screening is filtered out by setting keyword matching, grammar checking rules, and other requirements.

6. The large model-based multi-specification target object grasping method according to claim 1, wherein, In step three, the semantic guidance module obtains the semantic geometric features of the object by using the three-dimensional shape feature encoder, interprets the specification semantic instructions by using the large language model LLM, and obtains the object specification features by text embedding, and designs a modal adapter to fuse the object and specification text features to obtain a general feature representation.

7. The large model-based multi-specification target object grasping method according to claim 6, characterized in that, In step three, the semantic guidance module is the process of extracting features of each object as follows: Step 3-1a: Obtain the semantic geometric features of the object by using the three-dimensional shape feature encoder; Step 3-2a: Interpret the specification semantic instructions by using the large language model LLM, and obtain the object specification features by text embedding; Step 3-3a; Design a modal adapter to fuse the object and specification text features to obtain a general feature representation suitable for different specification objects.

Citation Information

Patent Citations

  • Unmanned laboratory-oriented transparent object grabbing method

    CN118386230A

  • Man-machine co-fusion mechanical arm self-adaptive grabbing method and system based on multi-modal large model

    CN118789551A