System and method for segmenting unknown article instance grabbed by robot
By combining a depth camera with the Transformer architecture, the problems of insufficient data dependency and real-time processing capabilities in robot recognition of unknown objects are solved, efficient and accurate instance segmentation of unknown objects is achieved, and the segmentation effect of complex scenes is improved.
Patent Information
- Application Number
- CN202510870541.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies for robot unknown object recognition have problems such as strong dependence on large amounts of labeled data, disconnection between synthetic data and reality, insufficient real-time processing capabilities and flexibility, and poor object instance segmentation.
A depth camera is used to obtain RGB and depth information, patches are generated through superpixel segmentation, and the Transformer architecture and multi-layer fully connected neural network are combined to extract and classify patch features. The feature vector is constructed using the centroid coordinates and normal vector, and an undirected graph is constructed for instance segmentation.
It achieves efficient and accurate instance segmentation of unknown objects without the need for external semantic labels, improves the segmentation ability of irregular and occluded targets in complex scenes, and can achieve accurate object instance segmentation without category prior conditions.
Smart Images

Figure CN120765665A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot visual perception, and in particular relates to a system and method for segmenting instances of unknown objects grasped by a robot. Background Art
[0002] To autonomously perform tasks in diverse indoor environments, robots require a range of skills, such as perception. Regarding perception, robots must be able to accurately identify new objects in their environment. Because constructing 3D models of new objects in real time for identification is impractical, identifying unknown objects is crucial. Imagine a robot's task is to clean a tabletop. In this complex scenario, it might encounter new objects. Therefore, accurately identifying and grasping the target object remains a significant challenge.
[0003] Thanks to the rapid development of deep learning, robotics has achieved accurate object recognition and classification in diverse scenarios, significantly enhancing its application potential across numerous industrial sectors. However, these algorithms rely heavily on large, highly annotated training datasets, requiring complex point-to-point labeling operations, which significantly increases time and human resource costs. Consequently, deep learning models face challenges in generalizing to robotic tasks, particularly the iteration required to build customized datasets for new tasks, which limits their scalability to a wider range of application scenarios. In recent years, researchers have developed novel techniques for unknown object instance segmentation (UOIS) to better address robotic manipulation tasks. These techniques aim to segment and identify regions of unknown objects. They integrate RGB and depth information to learn properties associated with target objects. Furthermore, by generating simulated data and training the network based on this data, these methods reduce the need for data annotation and model training in new scenarios. UOIS techniques have demonstrated promising results and high process efficiency for the task of instance segmentation of unseen objects.
[0004] In summary, existing technical solutions face several major bottlenecks. Traditional supervised learning is limited by its strong reliance on large amounts of labeled data, which is inherently difficult to obtain. Methods based on synthetic data face the problem of a disconnect between simulation and reality. Furthermore, technologies that require interaction lack real-time processing capabilities and flexibility. Furthermore, existing basic models still have room for improvement when it comes to completing the task of segmenting object instances. Therefore, there is an urgent need to explore new technical approaches to address these challenges. Summary of the Invention
[0005] To address the limitations of traditional supervised learning due to its strong reliance on large amounts of labeled data, which is inherently difficult to obtain, and the disconnection between simulation and reality faced by methods based on synthetic data; to address the poor real-time processing capabilities and flexibility of interactive technologies; and to address the fact that even basic models still have room for improvement when completing the task of complete segmentation of object instances. The present invention provides a system and method for segmenting unknown object instances grasped by a robot, aiming to achieve efficient and accurate segmentation of unknown object instances without the need for external semantic labels or auxiliary information.
[0006] The technical solution adopted by the system and method for segmenting unknown objects grasped by a robot in the present invention is:
[0007] A system for segmenting unknown object instances for robotic grasping includes a patch-level feature extraction module. A depth camera is used to acquire RGB and depth information of the current scene. Superpixel segmentation is performed on the RGB and depth images to generate irregular patches. These segmentation results are then merged using a fusion strategy to create a unified segmentation map. For each integrated patch, a comprehensive feature vector containing color, spatial, and geometric information is extracted.
[0008] Patch Pair Interaction Enhancement Module: This module uses a Transformer-based architecture to process the extracted patch pair features. It captures the global contextual relationships and local features of the scene through patch embedding and multi-layer Transformer encoding.
[0009] Patch pair classification module: A four-layer fully connected neural network is used to perform binary classification on the Transformer output features to determine whether the patch pairs belong to the same instance.
[0010] A further improvement of the technical solution of the present invention is that: in the comprehensive feature vector, the centroid coordinates are calculated based on the mean of the patch point cloud, and the normal vector is obtained from the pre-calculated depth normal map.
[0011] A method for segmenting unknown objects grasped by a robot comprises the following steps:
[0012] S1. The SLIC algorithm is used to perform superpixel segmentation on the input RGB image and depth image respectively;
[0013] S2, fuse the RGB and depth superpixel segmentation results into a joint segmentation map;
[0014] S3, extract the RGB mean, 3D centroid coordinate position and normal direction of each patch and construct a feature vector;
[0015] S4. By determining whether the bounding boxes of two patches intersect, we can filter out truly adjacent patch pairs, ensuring that relationship modeling is only carried out between patches that have actual spatial connection significance, thus avoiding redundant calculations.
[0016] S5. The feature vector of the patch pair is input into the multi-layer Transformer encoding module for relationship modeling;
[0017] S6. Use a four-layer fully connected neural network to perform binary classification on the patch features encoded by Transformer and output the adjacency probability;
[0018] S7. Based on the binary classification results between patch pairs, an undirected graph is constructed and a complete object instance segmentation map is obtained using connected component analysis.
[0019] A further improvement of the technical solution of the present invention is that: in step S1, the RGB image is divided into r The depth image is segmented into superpixels, represented by P d Representation, and assign labels based on depth similarity.
[0020] A further improvement of the technical solution of the present invention is that: the method of fusing the RGB and depth superpixel segmentation results into a joint segmentation map in step S2 is specifically to combine P r Offset to thousands and add P d To ensure that each patch corresponds to a unique identifier of the intersection of RGB and depth superpixels, for each superpixel index i, we compute the combined patch identifier The calculation formula is:
[0021]
[0022] Among them, N p Represents the total number of patches.
[0023] A further improvement of the technical solution of the present invention is that: the step S3 is specifically,
[0024] After generating a unified superpixel segmentation map, for each integrated patch Analyze and extract local features for subsequent processing; for each The following features are calculated:
[0025] RGB average value: R i ,G i ,B i ;
[0026] Center of mass coordinates:
[0027] Surface normal to the center of mass:
[0028] The above features constitute each patch The underlying descriptor of ;patch The eigenvector F i Defined as:
[0029]
[0030] Among them, the coordinates of the center of mass ( ) is based on the point cloud in the patch The surface normal vector at the center of mass is calculated from the average position in It is obtained from the pre-computed depth normal map using the center of mass coordinates as the index.
[0031] A further improvement of the technical solution of the present invention is that: specifically, step S5 is that the Transformer encoder is composed of identical layers, each layer contains two basic components: multi-head self-attention and feedforward network, layer normalization and residual connection; the output of the multi-head self-attention is then normalized by layer, and in each encoder layer, the multi-head self-attention mechanism first calculates the self-attention of all patch embeddings, so that the model can capture the complex dependencies in the input sequence; the multi-head self-attention output is layer normalized and combined with the input through residual connection; then, the normalized representation is processed by the feedforward network, which consists of two linear transformations separated by a ReLU activation function, and the output of the feedforward network also undergoes layer normalization and residual connection.
[0032] A further improvement of the technical solution of the present invention is that: step S6 specifically implements patch pair classification using a four-layer FCNN architecture. The four-layer FCNN architecture consists of three hidden layers with ReLU activation functions and a final output layer. Except for the last layer, each FCNN layer uses learnable weights and biases for linear transformation and then performs ReLU activation; then, the output of the network passes through a sigmoid function to generate a continuous probability score between 0 and 1. The higher the score, the greater the possibility that the patches are truly adjacent.
[0033] The further improvement of the technical solution of the present invention is that: the step S7 specifically uses a threshold of 0.5 to perform binary processing on the predicted probability. When the predicted adjacency probability exceeds the threshold, a definite connection is established between the patch pairs; an undirected graph G = (V, E) is constructed, where each vertex V in V i Represents a patch area; the edge e in E=(V i ,V j) is built on vertex pairs with binary prediction value 1; finally, a connected component analysis method based on depth-first search is used to identify different instance regions.
[0034] Due to the adoption of the above technical solution, the technical advancements achieved by the present invention include:
[0035] This paper proposes a system and method for segmenting unknown object instances for robot grasping. By fusing the superpixel segmentation maps of the RGB image and the depth map, the present invention effectively integrates multimodal information and strengthens the expression of boundaries and geometric features. At the same time, superpixel patch pairs are constructed, and the self-attention mechanism of the multi-layer Transformer is used to capture the complex dependencies between superpixels, which greatly improves the model's ability to segment irregular and occluded targets in complex scenes. In addition, a binary classifier based on patch pair relationships and a graph structure reasoning strategy are designed, which enables the model to achieve accurate unknown object instance segmentation through local relationship aggregation without category prior conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a diagram of an example segmentation system for unknown objects grasped by a robot according to the present invention;
[0037] Figure 2 It is a processing flow chart of a method for segmenting instances of unknown objects grasped by a robot according to the present invention;
[0038] Figure 3 This is a diagram of the overall network structure of a method for segmenting unknown objects grasped by a robot according to the present invention;
[0039] Figure 4 This is a visualization example of the segmentation results of the unknown object instance segmentation method grasped by a robot on the TOD-Z dataset;
[0040] Figure 5 and Figure 6 The invention discloses an instance segmentation method for unknown objects grasped by a robot, which is applied to a test process of the robot grasping physical objects. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. In the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0042] The present invention provides an unknown object instance segmentation system for robot grasping, comprising a patch-level feature extraction module, a patch pair interaction enhancement module, and a patch pair classification module.
[0043] The above patch-level feature extraction module, the robot uses a depth camera to obtain the current scene RGB and depth information. Then the RGB and depth images are segmented by superpixels respectively, generating irregular patches. Then these segmentation results are merged by fusion strategy to create a unified segmentation map. For each integrated patch, we will extract a comprehensive feature vector containing color, spatial and geometric information; patch pair interaction enhancement module, a Transformer-based architecture is used to process the extracted patch pair features. Through patch embedding and multi-layer Transformer encoding, the global context relationship and local features in the scene can be effectively captured; patch pair classification module, we process the features output by the patch pair interaction enhancement module with a four-layer fully connected neural network (FCNN) to classify the relationship between patch pairs. Then through graph-based processing to perfect these adjacency predictions, where connected component analysis is used to generate the final instance segmentation map.
[0044] The application also provides a robot grasping unknown object instance segmentation method, as shown in Figure 2 The above segmentation method first uses superpixel segmentation to independently divide the RGB image and the depth image, then fuses them to form superpixel regions with rich spatial and texture information. Then, multi-modal local features are extracted, including color average, spatial position and surface normal information, combined with the patch pair building method to form a patch pair feature matrix. Subsequently, a Transformer encoder is introduced to globally model the patch pairs, dynamically adjusting the dependency relationship between different regions using the self-attention mechanism, enhancing the understanding ability of irregular boundaries and occlusions in complex scenes. In order to realize the judgment of instances, a relationship binary classification module based on multi-layer fully connected network (FCNN) is designed to judge the coded region pairs and identify the region relationship belonging to the same instance. Finally, by constructing a relationship graph, the positively connected superpixel regions are aggregated, and the complete instance mask is obtained using the graph structure post-processing strategy, realizing the accurate segmentation of unknown objects. The whole process from local feature extraction to global relationship modeling, to graph structure reasoning forms a complete and systematic zero-shot segmentation scheme.
[0045] As shown in Figure 3 The specific steps of the robot grasping unknown object instance segmentation method provided by the application are as follows:
[0046] S1. Superpixel Generation. The SLIC algorithm is used to perform superpixel segmentation on the input RGB image and depth image. Direct pixel-level processing is computationally intensive and noisy. Superpixels can divide the image into regions with consistent boundaries, reducing the complexity of subsequent processing and improving feature stability. By generating superpixels for RGB and depth separately, color and geometric information can be fully utilized, allowing subsequent stages to combine these two types of information and enhance understanding of the target boundaries and structure.
[0047] The above RGB image is segmented into r The depth image is segmented into superpixels, represented by P d Representation, and assign labels based on depth similarity.
[0048] S2. RGB-D superpixel fusion. The RGB and depth superpixel segmentation results are fused into a joint segmentation map. The fusion process uses a weighted encoding strategy to map each pixel's label in both superpixel maps to a unique patch ID, thereby constructing a fused, comprehensive segmentation map that combines color and depth information. This approach aims to integrate the advantages of both information sources, addressing potential missegmentation and information loss issues with a single modality, and laying a more accurate foundation for subsequent patch feature extraction and relationship learning.
[0049] To combine the above superpixels into a unified patch representation, we compute a combined label map. This arithmetic coding is done by combining P r Offset to thousands and add P d To ensure that each patch corresponds to a unique identifier of the intersection of RGB and depth superpixels, for each superpixel index i, we compute the combined patch identifier The calculation formula is:
[0050]
[0051] Among them, N p Represents the total number of patches.
[0052] S3. Patch-level feature extraction. Features are extracted from each patch, including geometric and color information such as the RGB mean, 3D centroid coordinates, and normal direction. Computing these patch features provides comprehensive metrics for each region that reflect both appearance and spatial shape. This helps determine the relationships between different patches and achieve a deeper understanding of the object structure in the scene.
[0053] Specifically, after generating a unified superpixel segmentation map, each integrated patch Analyze and extract local features for subsequent processing; for each The following features are calculated:
[0054] RGB average value: R i ,G i ,B i ;
[0055] Center of mass coordinates:
[0056] Surface normal to the center of mass:
[0057] The above features constitute each patch The underlying descriptor of ;patch The eigenvector F i Defined as:
[0058]
[0059] Among them, the coordinates of the center of mass ( ) is based on the point cloud in the patch The surface normal vector at the center of mass is calculated from the average position in It is obtained from the pre-computed depth normal map using the center of mass coordinates as the index.
[0060] S4. Generate patch pairs. By determining whether the bounding boxes of two patches intersect, we filter out truly adjacent patch pairs. This ensures that relationship modeling is performed only between patches that are meaningfully spatially connected, avoiding redundant computation. This more effectively captures potential structural information between patches that belong to the same object, facilitating subsequent relationship inference.
[0061] S5, Transformer encoder. The feature vectors of the patch pairs are input into a multi-layer Transformer encoder module for relationship modeling. The Transformer's self-attention mechanism captures the long-range and complex dependencies between patch pairs, fusing features layer by layer to achieve a richer and deeper representation. This design leverages the Transformer's powerful contextual modeling capabilities, enabling the model to adaptively strengthen its focus on key relationships between patch features, thereby improving relationship classification accuracy.
[0062] Specifically, the Transformer encoder consists of identical layers, each of which contains two basic components: multi-head self-attention (MSA) and a feed-forward network (FFN), layer normalization (LN), and residual connections. The output of the MSA is then normalized using LN. In each encoder layer, the MSA mechanism first computes the self-attention of all patch embeddings, enabling the model to capture complex dependencies in the input sequence. The MSA output passes through the LN and is combined with the input through a residual connection. The normalized representation is then processed by the FFN, which consists of two linear transformations separated by a ReLU activation function. The FFN output also undergoes layer normalization and residual connections.
[0063] S6. Patch pair relationship classification. After Transformer encoding, the patch pair features are fed into a multi-layer, fully connected neural network for binary classification to determine whether the pair of patches belong to the same instance. This decision is made using probabilistic outputs. The classification results directly determine the presence of edges in the next step of the graph structure, ensuring that only patches that truly belong to the same object are connected.
[0064] Specifically, patch pair classification is implemented using a four-layer FCNN architecture. The four-layer FCNN architecture consists of three hidden layers with ReLU activation function and a final output layer. Except for the last layer, each FCNN layer uses learnable weights and biases for linear transformation, and then performs ReLU activation; then, the output of the network passes through a sigmoid function to produce a continuous probability score between 0 and 1. The higher the score, the greater the possibility that the patches are truly adjacent.
[0065] S7. Graph Construction and Instance Segmentation. Based on the binary classification results between patch pairs, an undirected graph is constructed, with all pairs of patches classified as the same instance as edges and the patches themselves as nodes. This graph aggregates patches by searching for connected subgraphs, resulting in a complete object instance segmentation graph.
[0066] Specifically, we use a threshold of 0.5 to binarize the predicted probability. When the predicted adjacency probability exceeds the threshold, a definite connection is established between the patch pairs. Based on this, we construct an undirected graph G = (V, E), where each vertex V in V i Represents a patch area. The edge e in E=(V i ,V j ) is built on vertex pairs with a binary prediction value of 1 (indicating high predicted adjacency). This graph structure effectively encodes the spatial relationships and connectivity patterns between patches. Finally, we employ connected components analysis based on depth-first search (DFS) to identify distinct instance regions.
[0067] Example 1
[0068] Parameter settings
[0069] The proposed method was validated on the TOD-Z dataset, a widely used instance segmentation method for unknown objects in indoor environments. In this embodiment, when using SLIC to perform superpixel segmentation on a depth image, the parameter nsegments was set to 256, meaning that the image was segmented into 256 superpixel regions; the parameter compactness was set to 0.08, which controls the compactness of the superpixel shape; the lower the value, the closer the superpixel shape is to a square; and the parameter sigma was set to 0.01, which is the standard deviation of the Gaussian kernel used to calculate pixel similarity and affects the smoothness of the superpixel edges. When segmenting an RGB image, nsegments was set to 256, compactness was set to 10, and sigma was set to 1.
[0070] In our transformer encoder, the main architecture parameters include: D = 18, the number of attention heads (nh = 6), and the number of transformer-encoder layers (L = 2). The FCNN has four fully connected layers (256, 1024, 256, 1). We used the Adam optimizer with a decay rate of 0.5 and an initial learning rate of 0.001. The network was trained for 10 epochs.
[0071] Algorithm Model
[0072] The computational steps for testing a set of RGB-D images are as follows:
[0073] Input: preprocessed RGB image and depth image;
[0074] Step 1: The RGB image is segmented into r The depth image is segmented into superpixels, represented by P d Representation, and assign labels based on depth similarity. To merge these patterns into a unified patch representation, we compute a combined label map. This arithmetic coding is done by r Offset to thousands and add P d To ensure that each patch corresponds to a unique identifier of the intersection of RGB and depth superpixels, thus avoiding label conflicts. For each superpixel index i, we compute the combined patch identifier The calculation formula is:
[0075]
[0076] Among them, N p Represents the total number of patches.
[0077] Step 2: After generating the unified superpixel segmentation map, we perform Analyze and extract local features for subsequent processing. We will calculate the following features:
[0078] RGB average value: R i ,G i ,B i
[0079] Center of mass coordinates:
[0080] Surface normal to the center of mass:
[0081] These features make up each patch The underlying descriptor of the patch. The eigenvector F i Defined as:
[0082]
[0083] Center of mass coordinates ( ) is based on the point cloud in the patch The surface normal vector at the center of mass is calculated from the average position in It is obtained from the pre-computed depth normal map using the center of mass coordinates as the index.
[0084] Step 3: Generate patch pairs. By determining whether the bounding boxes of two patches intersect, we filter out truly adjacent patch pairs. This ensures that relationship modeling is performed only between patches that are meaningfully connected in space, avoiding redundant computation.
[0085] Step 4: The feature vectors of the patch pairs are input to a multi-layer Transformer encoder module for relationship modeling. The Transformer encoder consists of identical layers, each of which contains two basic components: multi-head self-attention (MSA) and a feed-forward network (FFN), layer normalization (LN), and residual connections. The output of the MSA is then normalized using LN. In each encoder layer, the MSA mechanism first calculates the self-attention of all patch embeddings, enabling the model to capture complex dependencies in the input sequence. The MSA output passes through the LN and is combined with the input through a residual connection. The normalized representation is then processed by the FFN, which consists of two linear transformations separated by a ReLU activation function. The FFN output also undergoes layer normalization and residual connections.
[0086] Step 5: Patch pair classification is implemented using a four-layer FCNN architecture. This architecture consists of three hidden layers with ReLU activation and a final output layer. Each FCNN layer is linearly transformed using learnable weights and biases, followed by ReLU activation (except the last layer). The network output is then passed through a sigmoid function to produce a continuous probability score between 0 and 1, with higher scores indicating a greater likelihood that the patches are truly adjacent.
[0087] Step 6: Use a threshold of 0.5 to binarize the predicted probability. When the predicted adjacency probability exceeds the threshold, a definite connection is established between the patch pairs. Based on this, we construct an undirected graph G = (V, E), where each vertex V in V i Represents a patch area. The edge e in E=(V i ,V j ) is built on vertex pairs with a binary prediction value of 1 (indicating high predicted adjacency). This graph structure effectively encodes the spatial relationships and connectivity patterns between patches. Finally, we employ connected components analysis based on depth-first search (DFS) to identify distinct instance regions.
[0088] Output: Instance segmentation results of objects in the image.
[0089] Experimental evaluation and analysis
[0090] This method has good performance in the field of indoor scene object instance segmentation. We first conducted a verification test on the widely used TOD-Z dataset, such as Figure 4 The segmentation results shown demonstrate that this method can achieve efficient and accurate instance segmentation of unknown objects without the need for external semantic labels or auxiliary information, and effectively segment target objects in RGB-D images.
[0091] In order to further verify the feasibility of the method in actual application scenarios, we conducted field tests using the Fetch mobile robot. The Fetch robot is equipped with an RGB camera and a Primesense Carmine 1.09 depth camera. The resolution of the RGB and depth cameras is 640x 480 pixels, and the viewing angle is 60 degrees. A 7-degree-of-freedom robotic arm with a load of 6 kg was used, and the end of the arm was equipped with two finger grippers for grasping objects. During the test, the scene image was first collected by the RGB-D camera on the head of the robot, and then the objects in the image were segmented using the method proposed in this article. After completing the object segmentation, the system calls Contact-GraspNet for grasping posture planning, and MoveIt performs specific motion planning. As Figure 5 and Figure 6 As shown in the figure, the experiment fully records the entire process of the robot from target recognition, object segmentation to final successful grasping, which fully demonstrates the effectiveness of this method in practical applications.
[0092] In the above embodiment, the present invention provides a system and method for segmenting unknown object instances grasped by a robot. The present invention effectively integrates multimodal information and strengthens the expression of boundaries and geometric features by fusing the superpixel segmentation map of the RGB image and the depth map. At the same time, superpixel patch pairs are constructed, and the self-attention mechanism of the multi-layer Transformer is used to capture the complex dependencies between superpixels, which greatly improves the model's ability to segment irregular and occluded targets in complex scenes. In addition, a binary classifier based on patch pair relationships and a graph structure reasoning strategy are designed, which enables the model to achieve accurate unknown object instance segmentation through local relationship aggregation without category prior conditions.
[0093] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the concept and scope of the present invention. Any modifications and improvements made to the technical solution of the present invention by a person of ordinary skill in the art without departing from the design concept of the present invention shall fall within the scope of protection of the present invention. The technical content for which protection is sought in the present invention is fully set forth in the claims.
Claims
1. A system for segmenting unknown objects grasped by a robot, characterized by: It includes a patch-level feature extraction module: a depth camera is used to obtain the RGB and depth information of the current scene, and superpixel segmentation is performed on the RGB and depth images respectively to generate irregular patches. These segmentation results are merged through a fusion strategy to create a unified segmentation map; for each integrated patch, a comprehensive feature vector containing color, spatial and geometric information is extracted; Patch Pair Interaction Enhancement Module: This module uses a Transformer-based architecture to process the extracted patch pair features. It captures the global contextual relationships and local features of the scene through patch embedding and multi-layer Transformer encoding. Patch pair classification module: A four-layer fully connected neural network is used to perform binary classification on the Transformer output features to determine whether the patch pairs belong to the same instance.
2. The unknown object instance segmentation system for robot grasping according to claim 1, characterized in that: In the comprehensive feature vector, the centroid coordinates are calculated based on the mean of the patch point cloud, and the normal vector is obtained from the pre-calculated depth normal map.
3. A method for segmenting unknown objects grasped by a robot, characterized by: The following steps are included: S1. The SLIC algorithm is used to perform superpixel segmentation on the input RGB image and depth image respectively; S2, fuse the RGB and depth superpixel segmentation results into a joint segmentation map; S3, extract the RGB mean, 3D centroid coordinate position and normal direction of each patch and construct a feature vector; S4. By determining whether the bounding boxes of two patches intersect, we can filter out truly adjacent patch pairs, ensuring that relationship modeling is only carried out between patches that have actual spatial connection significance, thus avoiding redundant calculations. S5. The feature vector of the patch pair is input into the multi-layer Transformer encoding module for relationship modeling; S6. Use a four-layer fully connected neural network to perform binary classification on the Transformer-encoded patch features and output the adjacency probability. S7. Based on the binary classification results between patch pairs, an undirected graph is constructed and a complete object instance segmentation map is obtained using connected component analysis.
4. The method for segmenting unknown objects grasped by a robot according to claim 3, wherein: In step S1, the RGB image is segmented into r The depth image is segmented into superpixels, represented by P d Representation, and assign labels based on depth similarity.
5. The method for segmenting unknown objects grasped by a robot according to claim 4, characterized in that: The method of fusing the RGB and depth superpixel segmentation results into a joint segmentation map in step S2 is specifically as follows: r Offset to thousands and add P d To ensure that each patch corresponds to a unique identifier of the intersection of RGB and depth superpixels, for each superpixel index i, we compute the combined patch identifier The calculation formula is: Among them, N p Represents the total number of patches.
6. The method for segmenting unknown objects grasped by a robot according to claim 5, characterized in that: The step S3 is specifically as follows: After generating a unified superpixel segmentation map, for each integrated patch Analyze and extract local features for subsequent processing; for each The following features are calculated: RGB Average value: R i ,G i ,B i ; Center of mass coordinates: Surface normal to the center of mass: The above features constitute each patch The underlying descriptor of ;patch The eigenvector F i Defined as: Among them, the centroid coordinates Is based on the point cloud in the patch The surface normal vector at the center of mass is calculated from the average position in It is obtained from the pre-computed depth normal map using the center of mass coordinates as the index.
7. The method for segmenting unknown objects grasped by a robot according to claim 6, characterized in that: Specifically, step S5 comprises the following steps: the Transformer encoder is composed of identical layers, each of which contains two basic components: multi-head self-attention and feedforward networks, layer normalization, and residual connections; the output of the multi-head self-attention is then normalized using layers, and in each encoder layer, the multi-head self-attention mechanism first calculates the self-attention of all patch embeddings, enabling the model to capture complex dependencies in the input sequence; The multi-head self-attention output is layer-normalized and combined with the input through a residual connection; the normalized representation is then processed by a feed-forward network, which consists of two linear transformations separated by a ReLU activation function. The output of the feed-forward network also undergoes layer normalization and residual connections.
8. The method for segmenting unknown objects grasped by a robot according to claim 7, wherein: Specifically, step S6 implements patch pair classification using a four-layer FCNN architecture. The four-layer FCNN architecture consists of three hidden layers with ReLU activation functions and a final output layer. Except for the last layer, each FCNN layer uses learnable weights and biases for linear transformation and then performs ReLU activation. Then, the output of the network passes through a sigmoid function to generate a continuous probability score between 0 and 1. The higher the score, the greater the possibility that the patches are truly adjacent.
9. The method for segmenting unknown objects grasped by a robot according to claim 8, characterized in that: The step S7 specifically comprises: using a threshold of 0.5 to perform binarization processing on the predicted probability; when the predicted adjacency probability exceeds the threshold, a definite connection is established between the patch pairs; constructing an undirected graph G = (V, E), where each vertex V in V i Represents a patch area; the edge e in E=(V i ,V j ) is built on vertex pairs with binary prediction value 1; finally, a connected component analysis method based on depth-first search is used to identify different instance regions.