An open-vocabulary grasping method and system based on Gaussian force field

By combining the Gaussian model and Gaussian force field embedded in the complex three-dimensional environment, the gripper posture is optimized, and the problem of slow rendering speed, insufficient reliance on a large amount of labeled data and generalization capabilities in the existing technology is solved, and efficient semantic understanding and stable capture of complex environments is achieved.

CN119810839BActive Publication Date: 2025-06-27ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510304167.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

When handling complex three-dimensional environments and irregularly shaped objects, the rendering speed is slow, it is difficult to extract clear geometric information, rely on a large amount of labeled data, and the generalization ability is weak, making it difficult to meet the real-time operation needs of robots in complex environments.

Method used

The open vocabulary grasping method based on Gaussian force field is adopted to combine the language-embedded three-dimensional Gaussian model with the Gaussian force field. By constructing the target object point set, the holder model point set and the attraction point set, the repulsive force and attraction function are defined, and the iterative gradient descent method is used to optimize the holder posture to achieve semantic understanding and stable grasping of complex three-dimensional environments.

Benefits of technology

It realizes efficient semantic understanding and stable capture of complex three-dimensional environments, reduces dependence on labeled data, improves the processing ability of irregularly shaped objects, and is suitable for real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810839B_ABST
    Figure CN119810839B_ABST
Patent Text Reader

Abstract

The present invention discloses an open-vocabulary grasping method and system based on a Gaussian force field, which relates to the field of artificial intelligence technology. An input grasping instruction is input into a posture grasping model that has been trained in a robot to obtain a robot grasping result. The training process of the posture grasping model is as follows: constructing a target object point set, a gripper model point set, and an attraction point set; obtaining a transformed gripper model point set and a transformed attraction point set through a transformation matrix; calculating the sum of exponential functions of distances between all point pairs in the target object point set, the transformed gripper model point set, and the transformed attraction point set; minimizing the sum of the exponential functions to achieve a balance between repulsive force and attractive force, and using an iterative gradient descent method to optimize the trainable parameters in the posture grasping model. This open-vocabulary grasping method and system combines a language-embedded three-dimensional Gaussian model with a Gaussian force field to achieve open-vocabulary grasping of a robot in a complex three-dimensional environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an open vocabulary grasping method and system based on Gaussian force field. Background Art

[0002] 1) Robot grasping pose detection;

[0003] The goal of robot grasping pose detection is to determine the optimal pose of the robot's hand (such as a manipulator or gripper) to firmly grasp the target object. Early methods usually simplified the grasping task to a two-dimensional problem and defined the grasping pose as a directed rectangle with a fixed height. Due to ignoring the complexity in three-dimensional space, these methods can only achieve three-degree-of-freedom operations and cannot meet the requirements of accuracy and flexibility in practical applications.

[0004] To overcome this limitation, researchers began to explore six-degree-of-freedom grasping technologies. These methods combine depth information or point cloud data, enabling the robot to perceive the three-dimensional geometric shape and spatial position information of the object, thereby generating a more accurate grasping pose. With accurate depth data, these methods have achieved remarkable results. However, in the real world, it is not always easy to obtain high-quality depth information. For example, for transparent or reflective objects, the performance of depth sensors may be severely affected.

[0005] At the same time, with the rapid development of deep learning and computer vision technologies, neural network-based grasping pose detection methods have gradually emerged. These methods use a large amount of data to train convolutional neural networks, enabling them to extract rich features from images. However, these methods usually require a large amount of labeled data for training, and obtaining this data is often costly. In addition, these methods have weak generalization ability when dealing with objects with irregular or complex shapes and are difficult to meet the requirements of practical applications.

[0006] 2) Three-dimensional feature field for operation;

[0007] To improve the robot's understanding ability of the three-dimensional environment, some researchers have proposed the idea of constructing a three-dimensional feature field. The three-dimensional feature field can represent the geometric and semantic information in the scene in a unified way, providing the robot with a more comprehensive environment perception ability. Currently, the methods for constructing a three-dimensional feature field are mainly divided into two categories: implicit representation and explicit representation.

[0008] Implicit representation methods, such as Neural Radiance Field (NeRF), learn a continuous scene representation through neural networks, which can achieve high-quality scene reconstruction and novel view rendering. However, NeRF requires a large amount of computing resources and has a slow rendering speed, making it difficult to meet the real-time requirements. In addition, it is difficult to directly extract explicit geometric information from implicit representation, which limits its application in robot operation.

[0009] Explicit representation methods use a clear three-dimensional geometric structure to represent the scene. For example, the 3D Gaussian Splatting (3DGS) technique uses Gaussian points in space to explicitly represent the scene. This method can achieve high-quality scene rendering at a relatively low computational cost and is also convenient for directly extracting geometric information. The method based on 3DGS further combines a vision-language model, embeds language features into Gaussian points, and constructs a Language Embedded Gaussian model. This model can support open-vocabulary scene segmentation and semantic queries, providing new possibilities for robots to understand and operate in the environment.

[0010] However, existing methods still have some deficiencies. First, they usually rely on a large amount of labeled data, and obtaining and labeling this data requires a lot of manpower and material resources. Second, when dealing with dynamic scenes or when the scene changes, the efficiency of these methods is low, and it is difficult to update the scene information in a timely manner. In addition, for objects with complex or irregular shapes, the generalization ability of existing methods still needs to be improved. In summary, there is an urgent need for an efficient and general method to enable robots to understand complex three-dimensional environments and operate based on natural language instructions.

[0011] Disadvantage 1 of the prior art: Slow rendering speed and difficult to extract clear geometric information. Although methods based on Neural Radiance Fields (NeRF) have made progress in scene modeling, due to their slow rendering speed, they are difficult to be applied to robot operations in real time. At the same time, these methods are difficult to extract clear geometric features from the scene, limiting their application in grasping tasks.

[0012] Disadvantage 2 of the prior art: Dependence on a large amount of labeled data. Existing grasping methods usually require a large amount of labeled grasping pose data, and the process of obtaining and labeling this data is time-consuming and costly, making it difficult to meet the requirements of practical applications.

[0013] Disadvantage 3 of the prior art: Poor handling of objects with irregular or complex shapes. When faced with objects with irregular or complex shapes, the generalization ability of existing methods is weak, and it is difficult to accurately predict stable grasping poses.

[0014] The above disadvantages limit the ability of robots to perform open-vocabulary grasping tasks based on natural language instructions in complex three-dimensional environments, and there is an urgent need for an efficient and general method to solve these problems. Summary of the Invention

[0015] Based on the technical problems existing in the background art, the present invention proposes an open-vocabulary grasping method and system based on a Gaussian force field, which combines a three-dimensional Gaussian model of language embedding with a Gaussian force field to achieve open-vocabulary grasping of a robot in a complex three-dimensional environment.

[0016] An open-vocabulary grasping method and system based on a Gaussian force field proposed by the present invention inputs a grasping instruction into a posture grasping model that has been trained in a robot to obtain a robot grasping result;

[0017] The training process of the posture grasping model is as follows:

[0018] Construct a target object point set, a gripper model point set, and an attraction point set, and define that there is a repulsive force between the target object point set and the gripper model point set, and there is an attractive force between the target object point set and the attraction point set;

[0019] Construct a transformation matrix for each point in the target object point set, and based on the transformation matrix, map the gripper model point set and the attraction point set to the local coordinate system of each point to obtain a transformed gripper model point set and a transformed attraction point set;

[0020] Calculate the sum of the exponential functions of the distances between all point pairs in the target object point set, the transformed gripper model point set, and the transformed attraction point set , and use as the repulsive force function, and use as the attractive force function;

[0021] Minimize the sum of the exponential functions , balance between the repulsive force and the attractive force, and use the iterative gradient descent method to optimize the trainable parameters in the posture grasping model.

[0022] Further, the construction process of the target object point set is as follows:

[0023] Obtain multi-view RGB video sequences and natural language instructions;

[0024] Based on a visual segmentation model, track and segment the objects in the RGB video sequence to obtain a temporally consistent object mask;

[0025] Extract semantic features from the segmented object mask based on the CLIP model;

[0026] Through dimensionality reduction processing, embed the extracted semantic features into a three-dimensional Gaussian model to construct a Gaussian model of language embedding;

[0027] Extract semantic features from the natural language instructions based on the CLIP model to obtain language instruction features;

[0028] Calculate the semantic similarity between the computational language instruction features and the semantic features of each Gaussian point in the scene, and segment the Gaussian points corresponding to the semantic similarity greater than the fixed threshold to obtain the target object point set.

[0029] Further, by sampling the three-dimensional model of the gripper, a gripper model point set is obtained; an attraction point set is obtained by uniformly sampling the spatial region between the gripper jaws.

[0030] Further, construct the transformation matrix for each point in the target object point set , specifically:

[0031] ;

[0032] where is the spatial coordinate of point , is the rotation matrix composed of the normal vectors of point , and are the learnable parameters of the pose grasping model, corresponding to the translation amount along the normal vector direction of point and the rotation angle around the normal vector of point respectively. Point is the th point in the target object point set.

[0033] Further, the transformed gripper model point set , is the gripper model point set, is the transformation matrix;

[0034] The transformed attraction point set , is the attraction point set.

[0035] Further, the absolute values of the repulsive force function and the attractive force function are the same, and the calculation formula of the repulsive force function is as follows:

[0036]

[0037] where is the repulsive force between the transformed gripper model point set and the target object point set , and are the learnable parameters of the pose grasping model, represents the th point in the target object point set and is the th point in the target object point set is the Euclidean distance between two points, is the temperature parameter, is the number of points in the gripper model point set, is the number of points in the target object point set.

[0038] Furthermore, during the training of the pose grasping model and the online execution after training is completed, record the minimum force function value corresponding to each point in the target object point set;

[0039] Select from all points in the target object point set the point that makes the total force function minimum and its corresponding parameters and ;

[0040] Based on the parameters and calculate the optimal transformation matrix , and apply it to the gripper model point set and the attraction point set, then the optimal gripper pose can be obtained.

[0041] An open-vocabulary grasping system based on Gaussian force field inputs the grasping instruction into the pose grasping model that has been trained in the robot to obtain the robot grasping result;

[0042] The training process of the pose grasping model includes a point set construction unit, a point set transformation unit, a calculation unit, and a gradient optimization unit;

[0043] The point set construction unit is used to construct the target object point set, the gripper model point set, and the attraction point set, and define that there is a repulsive force between the target object point set and the gripper model point set, and there is an attractive force between the target object point set and the attraction point set;

[0044] The point set transformation unit is used to construct the transformation matrix for each point in the target object point set, and map the gripper model point set and the attraction point set to the local coordinate system of each point based on the transformation matrix to obtain the transformed gripper model point set and the transformed attraction point set;

[0045] The calculation unit is used to calculate the sum of the exponential functions of the distances between all point pairs in the target object point set, the transformed gripper model point set, and the transformed attraction point set , take as the repulsive force function, and take as the attractive force function;

[0046] The gradient optimization unit is used to minimize the sum of the exponential functions , balance between the repulsive force and the attractive force, and optimize the trainable parameters in the pose grasping model by using the iterative gradient descent method.

[0047] Furthermore, by sampling the 3D model of the gripper, a gripper model point set is obtained; an attraction point set is obtained by uniformly sampling the spatial region between the gripper jaws.

[0048] Furthermore, the construction process of the target object point set in the point set construction module is as follows:

[0049] Obtain RGB video sequences from multiple perspectives and natural language instructions;

[0050] Based on a visual segmentation model, track and segment the objects in the RGB video sequence to obtain temporally consistent object masks;

[0051] Extract semantic features from the segmented object masks based on the CLIP model;

[0052] Through dimensionality reduction processing, embed the extracted semantic features into a three-dimensional Gaussian model to construct a Gaussian model with language embedding;

[0053] Extract semantic features from the natural language instructions based on the CLIP model to obtain language instruction features;

[0054] Calculate the semantic similarity between the language instruction features and the semantic features of each Gaussian point in the scene, and segment the Gaussian points corresponding to the semantic similarity greater than a fixed threshold to obtain the target object point set.

[0055] The advantages of an open-vocabulary grasping method and system based on Gaussian force field provided by the present invention are as follows: By constructing a Gaussian model with language embedding, the system has the ability to understand scenes with open vocabulary, can segment and recognize any object in a complex scene, supports natural language-based queries and operations, and realizes semantic understanding of complex three-dimensional environments. At the same time, the proposed Gaussian force field method does not require a large amount of labeled data, and the six-degree-of-freedom pose of the gripper can be estimated in a self-supervised manner, reducing the cost and difficulty of data acquisition. By balancing the repulsive force and the attractive force in the Gaussian force field and using the gradient descent method to quickly optimize the pose of the gripper, efficient and stable grasping of the target object is achieved, which is suitable for real-time applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic diagram of the training process of the pose grasping model of the present invention;

[0057] Figure 2 It is a schematic diagram of the Gaussian force field; showing the force relationship between the target object point set, the gripper model point set and the attraction point set;

[0058] Figure 3 It is an application schematic diagram of the open-vocabulary grasping method in a simulation environment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] Next, the technical solution of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0060] As Figures 1 to 3 shown, an open-vocabulary grasping method based on Gaussian force field proposed by the present invention inputs a grasping instruction into the pose grasping model that has been trained in the robot to obtain the robot grasping result;

[0061] The training process of the pose grasping model is as follows:

[0062] Construct a target object point set, a gripper model point set, and an attraction point set, and define that there is a repulsive force between the target object point set and the gripper model point set, and there is an attractive force between the target object point set and the attraction point set;

[0063] Construct a transformation matrix for each point in the target object point set, and based on the transformation matrix, map the gripper model point set and the attraction point set to the local coordinate system of each point to obtain the transformed gripper model point set and the transformed attraction point set;

[0064] Calculate the sum of the exponential functions of the distances between all point pairs in the target object point set, the transformed gripper model point set, and the transformed attraction point set , and use as the repulsive force function, and use as the attractive force function;

[0065] Minimize the sum of the exponential functions , balance between the repulsive force and the attractive force, and use the iterative gradient descent method to optimize the trainable parameters in the pose grasping model.

[0066] In this embodiment, by using 3D Gaussian Splatting (3DGS) and a vision-language foundation model, a Gaussian model embedded with language is constructed to achieve 3D scene segmentation under open-language queries. In addition, a Gaussian force field, a self-supervised gripper pose estimation method, is proposed to avoid the dependence on a large amount of labeled grasping pose data.

[0067] The core idea of this embodiment is to combine the three-dimensional Gaussian model of language embedding with the Gaussian force field to achieve open-vocabulary grasping of a robot in a complex three-dimensional environment. First, using pre-trained models such as SAM2 and CLIP, semantic and feature information is extracted from the video sequence of the scene to construct a Gaussian model of language embedding, providing open-vocabulary understanding of the scene. Then, for a specified natural language instruction, the target object is segmented, and the corresponding Gaussian force field is constructed. The Gaussian force field consists of the point cloud of the target object, the point cloud of the gripper model, and the attraction points. By designing an appropriate force model, the repulsive force between the positive point sets prevents the gripper from colliding with the object, and the attractive force between the positive and negative point sets promotes the gripper to move towards a stable grasping posture. Finally, through iterative gradient optimization on this force field, the optimal six-degree-of-freedom posture of the gripper is predicted and obtained, realizing stable grasping of the target object.

[0068] The implementation of the open-vocabulary grasping method in this embodiment mainly relies on the following modules:

[0069] (1) Scene model construction module: Using three-dimensional Gaussian scattering (3DGS) technology, the input multi-view RGB video sequence is processed to construct a three-dimensional Gaussian model of the scene;

[0070] (2) Gaussian model construction module of language embedding: Using SAM2 and CLIP models, the objects in the scene are tracked and segmented, the language features of the objects are extracted, and embedded into the Gaussian model to achieve open-vocabulary scene representation; among them, SAM2 is a new generation of image and video segmentation model released by Meta; the CLIP model is a multi-modal pre-trained neural network proposed by OpenAI in 2021;

[0071] For the input multi-view RGB video sequence, the SAM2 model is used to track and segment the objects in the video sequence to obtain temporally consistent object masks. The CLIP model is used to extract semantic features from the segmented object masks, and through dimensionality reduction processing, the high-dimensional language features are embedded into the three-dimensional Gaussian model to construct a Gaussian model of language embedding.

[0072] In one of the embodiments, the construction process of the target object point set is as follows: Obtain RGB video sequences from multiple perspectives and natural language instructions; Based on a visual segmentation model, track and segment the objects in the RGB video sequences to obtain temporally consistent object masks; Extract semantic features from the segmented object masks based on the CLIP model; Through dimensionality reduction processing, embed the extracted semantic features into a three-dimensional Gaussian model to construct a Gaussian model with language embedding; Extract semantic features from the natural language instructions based on the CLIP model to obtain language instruction features; Calculate the semantic similarity between the language instruction features and the semantic features of each Gaussian point in the scene, and segment the Gaussian points corresponding to the semantic similarity greater than a fixed threshold to obtain the target object point set.

[0073] By constructing a Gaussian model with language embedding, the system is equipped with the ability to understand scenes with open vocabulary, can segment and recognize any object in a complex scene, supports natural language-based queries and operations, and realizes semantic understanding of complex three-dimensional environments.

[0074] (3) Gaussian force field construction module: For a specified natural language instruction, segment the Gaussian point cloud of the target object, and combine it with the gripper model to construct a Gaussian force field.

[0075] The construction of the Gaussian force field is the core of this embodiment. By simulating the interaction between the gripper and the target object in the Gaussian force field, self-supervised optimization of the gripper posture is achieved. Specifically, the Gaussian force field consists of a target object point set, a gripper model point set, and an attraction point set. By setting the force relationship between the positive and negative point sets, the gripper is guided to reach the optimal grasping posture.

[0076] (3-1) Definition of the Gaussian force field;

[0077] The Gaussian force field consists of the following three types of point sets:

[0078] Target object point set : It is composed of the Gaussian point set of the target object obtained from (2), denoted as where , is the number of points in the target object point set, is the th Gaussian point of the target object.

[0079] Gripper model point set : Obtained by sampling the three-dimensional model of the gripper, denoted as where , is the number of points in the gripper model.

[0080] Attraction point set : Uniform sampling is obtained in the spatial region between the gripper jaws, denoted as , where , is the number of points in the attraction point set.

[0081] In the Gaussian force field: The target object point set and the gripper model point set are defined as positive point sets, denoted by the symbol "+". There is a repulsive force between the positive point sets to prevent the gripper from colliding with the target object. The attraction point set is defined as a negative point set, denoted by the symbol "-". There is an attractive force between the target object point set and the attraction point set to guide the gripper towards the optimal grasping position of the target object.

[0082] (3-2) Gripper pose parameterization;

[0083] For each point in the target object point set , set the learnable parameters and for the pose of the gripper. and are respectively defined as the translation amount along the normal vector direction of the point (usually taken as the normal vector of the target object surface) and the rotation angle around the normal vector of the point . Based on the above parameters, define the rigid body transformation of the gripper at the point . First, let the spatial coordinates of the point be , and the rotation matrix formed by the corresponding normal vectors be . Among them, and are important groups describing rigid body motion, representing the Special Euclidean Group and the Special Orthogonal Group in three-dimensional space respectively. is the group composed of all 3×3 rotation matrices, is the group composed of all 4×4 rigid body transformation matrices. Then, the transformation matrix is expressed as:

[0084] ; (1)

[0085] Among them, the first matrix represents placing the gripper at the point and aligning its pose, and the second matrix Represents rotation and translation about the normal vector.

[0086] Definition of the force function (3-3);

[0087] Through the transformation matrix , the set of points of the gripper model and the set of attraction points are mapped to the local coordinate system of the point to obtain the transformed set of points of the gripper model and the transformed set of attraction points . Next, the repulsive force and the attractive force between the gripper and the target object are defined respectively as: the repulsive force function and the attractive force function . Among them, the repulsive force function is the repulsive force between the transformed set of points of the gripper model and the set of points of the target object , and the attractive force function is the attractive force between the transformed set of attraction points and the set of points of the target object . The absolute values of the repulsive force function and the attractive force function are the same, and are summarized as the sum of exponential functions of the distances between all point pairs. The calculation formula of the repulsive force function is as follows:

[0088] ; (2)

[0089] where is the th point in the transformed set of points of the gripper model , is the th point in represents the Euclidean distance between two points, is the temperature parameter, which controls the attenuation degree of the force, and are the learnable parameters of the attitude grasping model, is the number of points in the set of points of the gripper model, is the number of points in the set of points of the target object. The attractive force function is similar to the repulsive force function, and only differs by a negative sign in the expression form.

[0090] (4) Gripper attitude optimization module: By optimizing the Gaussian force field and using the gradient descent method, the optimal six-degree-of-freedom attitude of the gripper is predicted and obtained.

[0091] The goal of this embodiment is to find the optimal parameters and such that the total force function is minimized. By minimizing the above formula (2), a balance can be achieved between the repulsive force and the attractive force, avoiding collisions between the gripper and the target object while guiding the gripper towards the optimal grasping position and posture. The iterative gradient descent method is used to optimize the parameters and , and the parameters are updated according to the gradient descent method. During the optimization process, the minimum force function values corresponding to each point are recorded. Finally, the point that minimizes the total force function is selected from all points and its corresponding parameters and . The optimal transformation matrix is applied to the gripper model point set and the attraction point set, and the optimal gripper posture can be obtained. This posture ensures the minimization of the repulsive force and the maximization of the attractive force between the gripper and the target object, realizing the stable grasping of the target object.

[0092] In this embodiment, by balancing the repulsive force and the attractive force in the Gaussian force field and using the gradient descent method to quickly optimize the posture of the gripper, the efficient and stable grasping of the target object is realized, which is suitable for real-time applications.

[0093] In summary, in this embodiment, by constructing a Gaussian model with language embedding, the system is equipped with the ability to understand scenes with open vocabulary, can segment and recognize any object in a complex scene, support natural language-based queries and operations, and realize the semantic understanding of complex three-dimensional environments. At the same time, the proposed Gaussian force field method does not require a large amount of labeled data, and the six-degree-of-freedom posture of the gripper can be estimated in a self-supervised manner, reducing the cost and difficulty of data acquisition. By balancing the repulsive force and the attractive force in the Gaussian force field and using the gradient descent method to quickly optimize the posture of the gripper, the efficient and stable grasping of the target object is realized, which is suitable for real-time applications. In addition, the method of this embodiment can still perform the grasping task on unseen objects in the zero-shot case, especially when dealing with objects with complex or irregular shapes, showing strong generalization ability and significantly improving the operation ability of the robot in complex environments.

[0094] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. An open vocabulary crawling method based on Gaussian force field, characterized in that: Input the grasping instruction into the posture grasping model that has been trained in the robot to obtain the robot grasping result; The training process of the posture grasping model is as follows: Constructing a target object point set, a gripper model point set and an attraction point set, and defining that there is a repulsive force between the target object point set and the gripper model point set, and there is an attractive force between the target object point set and the attraction point set; Constructing a transformation matrix for each point in the target object point set, mapping the gripper model point set and the attraction point set to the local coordinate system of each point based on the transformation matrix, and obtaining a transformed gripper model point set and a transformed attraction point set; Calculate the sum of the exponential functions of the distances between all pairs of points in the target object point set, the transformed gripper model point set, and the transformed attractor point set ,Will As a repulsive force function, As an attraction function; Minimize the sum of exponential functions , an iterative gradient descent method is used to optimize the trainable parameters in the posture grasping model; The absolute values ​​of the repulsive force function and the attractive force function are the same. The calculation formula of the repulsive force function is as follows: ; in, is the transformed gripper model point set and the target object point set The repulsive force between and are the learnable parameters of the posture grasping model, express The Points, The target object point set The Points, is the Euclidean distance between two points, is the temperature parameter, is the number of points in the gripper model point set, is the number of points in the target object point set.

2. The open vocabulary crawling method based on Gaussian force field according to claim 1 is characterized in that: The process of constructing the target object point set is as follows: Obtain multi-view RGB video sequences and natural language instructions; Track and segment objects in RGB video sequences based on a visual segmentation model to obtain temporally consistent object masks; Extract semantic features from the segmented object mask based on the CLIP model; Through dimensionality reduction processing, the extracted semantic features are embedded into a three-dimensional Gaussian model to construct a language embedding Gaussian model; Extract semantic features from natural language instructions based on the CLIP model to obtain language instruction features; The semantic similarity between the language instruction features and the semantic features of each Gaussian point in the scene is calculated, and the Gaussian points corresponding to the semantic similarity greater than a fixed threshold are segmented to obtain the target object point set.

3. The open vocabulary crawling method based on Gaussian force field according to claim 1, characterized in that: The gripper model point set is obtained by sampling the three-dimensional model of the gripper; and the attraction point set is obtained by uniformly sampling the space area between the gripper jaws.

4. The open vocabulary crawling method based on Gaussian force field according to claim 1, characterized in that: Construct the transformation matrix for each point in the target object point set , specifically: ; in, For point The spatial coordinates of For point The rotation matrix of the normal vector is and are the learnable parameters of the posture grasping model, corresponding to the points along The translation in the direction of the normal vector and around the point The rotation angle of the normal vector of point is the first points.

5. The open vocabulary crawling method based on Gaussian force field according to claim 1, characterized in that: The transformed gripper model point set , is the gripper model point set, is the transformation matrix; Transformed attraction point set , The set of attraction points.

6. The open vocabulary crawling method based on Gaussian force field according to claim 1, characterized in that: During the training of the posture grasping model and the online execution after the training is completed, the minimum force function value corresponding to each point in the target object point set is recorded; Select from all points in the target object point set such that the total force function The smallest point and its corresponding parameters and ; Based on parameters and Calculate the optimal transformation matrix , and applied to the gripper model point set and the attraction point set, the optimal gripper posture can be obtained.

7. An open vocabulary crawling system based on Gaussian force field, characterized in that: Input the grasping instruction into the posture grasping model that has been trained in the robot to obtain the robot grasping result; The training process of the posture grasping model includes a point set construction unit, a point set transformation unit, a calculation unit, and a gradient optimization unit; The point set construction unit is used to construct the target object point set, the gripper model point set and the attraction point set, and defines that there is a repulsive force between the target object point set and the gripper model point set, and there is an attractive force between the target object point set and the attraction point set; The point set transformation unit is used to construct a transformation matrix for each point in the target object point set, and based on the transformation matrix, the gripper model point set and the attraction point set are mapped to the local coordinate system of each point to obtain a transformed gripper model point set and a transformed attraction point set; The calculation unit is used to calculate the sum of the exponential functions of the distances between all pairs of points in the target object point set, the transformed gripper model point set, and the transformed attraction point set. ,Will As a repulsive force function, As an attraction function; The gradient optimization unit is used to minimize the sum of exponential functions , a balance is achieved between repulsion and attraction, and the trainable parameters in the posture grasping model are optimized using an iterative gradient descent method; Among them, the absolute values ​​of the repulsive force function and the attractive force function are the same, and the calculation formula of the repulsive force function is as follows: ; in, is the transformed gripper model point set and the target object point set The repulsive force between and are the learnable parameters of the posture grasping model, express The Points, The target object point set The Points, is the Euclidean distance between two points, is the temperature parameter, is the number of points in the gripper model point set, is the number of points in the target object point set.

8. The open vocabulary crawling system based on Gaussian force field according to claim 7, characterized in that: The gripper model point set is obtained by sampling the three-dimensional model of the gripper; and the attraction point set is obtained by uniformly sampling the space area between the gripper jaws.

9. The open vocabulary crawling system based on Gaussian force field according to claim 7, characterized in that: The construction process of the target object point set in the point set construction module is as follows: Obtain multi-view RGB video sequences and natural language instructions; Track and segment objects in RGB video sequences based on a visual segmentation model to obtain temporally consistent object masks; Extract semantic features from the segmented object mask based on the CLIP model; Through dimensionality reduction processing, the extracted semantic features are embedded into a three-dimensional Gaussian model to construct a language embedding Gaussian model; Extract semantic features from natural language instructions based on the CLIP model to obtain language instruction features; The semantic similarity between the language instruction features and the semantic features of each Gaussian point in the scene is calculated, and the Gaussian points corresponding to the semantic similarity greater than a fixed threshold are segmented to obtain the target object point set.

Citation Information

Patent Citations

  • Navigation control method for roadway inspection robot

    CN117192569A

  • Grabbing method and system based on Gaussian spatter, robot and storage medium

    CN118769237A