Weak supervision point cloud segmentation method and system for robot assembly

By applying a weakly supervised point cloud segmentation method based on Transformer algorithm in smart factories, using the central attention mechanism and position coding module, the problem of high cost of manual labeling is solved, and efficient and low-cost point cloud segmentation and robot assembly pose understanding is achieved.

CN120126142AActive Publication Date: 2025-06-10NINGDE SKEQI INTELLIGENT EQUIP CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510596972.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-10
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

In the smart factory scenario, the cost of manually labeling the robot assembly posture is high and time-consuming, resulting in limited application of point cloud segmentation in robot three-dimensional posture understanding.

Method used

A weakly supervised point cloud segmentation method based on Transformer algorithm is designed, and the global features of neighboring points are extracted and shared through the central attention mechanism and position coding module to reduce the dependence on manual annotation.

Benefits of technology

Point cloud segmentation with low manual labeling costs is realized, training efficiency and cost-effectiveness are improved, and the performance of robot assembly posture understanding is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126142A_ABST
    Figure CN120126142A_ABST
Patent Text Reader

Abstract

The invention discloses a weak supervision point cloud segmentation method and system for robot assembly, and the method comprises the steps: inputting a point cloud, converting an original picture into standardized input, building a model for the picture, and providing coordinates and RGB image input for the model; encoding based on the characteristics of the standardized input of the input; a central attention mechanism is designed on the basis of a Transform algorithm, and point cloud features are learned by extracting global features of adjacent points and sharing the global features among neighborhoods; and classifying the point clouds based on a central attention mechanism learning result, and outputting a final result. According to the invention, low-manual-annotation-cost location cloud segmentation is realized, so that the training is more efficient and the cost is lower.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent manufacturing, and in particular to a weakly supervised point cloud segmentation method and system for robot assembly. Background Art

[0002] In industrial manufacturing, the ability to successfully grasp objects is an essential skill for a robot. With the development of Industry 5.0, robots already have multiple degrees of freedom, precise perception and control capabilities, and can perform tasks autonomously or in collaboration with workers on the production line. At present, vision-based deep learning and reinforcement learning can achieve certain results in robot object grasping.

[0003] However, the inventors of this application found that the above technology has at least the following technical problems in the process of implementing the technical solution of the invention in the embodiment of this application: In modeling the robot's assembly posture, researchers use point cloud-based visual methods to learn the robot's posture. Point cloud segmentation plays an important role in understanding the robot's 3D posture. However, in the scenario of a smart factory, manually labeling the robot's assembly posture is costly and time-consuming, and how to solve this problem is a challenge. Summary of the invention

[0004] The embodiments of the present application provide a weakly supervised point cloud segmentation method and system for robot assembly, thereby solving the technical problem in the prior art that manual labeling of robot assembly postures is costly and time-consuming in the scenario of smart factories, and achieving point cloud segmentation with low manual labeling cost, thereby making training more efficient and cost-effective.

[0005] The embodiment of the present application provides a weakly supervised point cloud segmentation method for robot assembly, comprising: S1, input point cloud, convert the original image into standardized input, build a model for the image, and provide coordinates and RGB image input for the model; S2, encoding based on the input features of the normalized input; S3, based on the Transformer algorithm, designs a central attention mechanism to learn point cloud features by extracting global features of neighboring points and sharing them among various neighborhoods, thus obtaining two embedded global features and position encoding modules; S4, classifies the point cloud based on the results of the central attention mechanism learning and outputs the final result.

[0006] Furthermore, in step S3, it includes: S31, based on the original image, is converted into a standardized image to obtain coordinates and features, and the standardized coordinates and features are input. After a layer of MLP in the multi-layer perceptron, the features are obtained. and Features , , The following formula: ; ; Among them, F d, P d is the initial feature of the input; S32, the central attention mechanism and To study.

[0007] Furthermore, the central attention mechanism further includes: For each center point , its characteristics Through a linear layer , The dimension is 1; at the same time, the KNN algorithm is used to obtain the center point The coordinates of the k neighbor points , and the features corresponding to the k neighbor points , and then the global features are extracted by integrating the center weights and neighboring point features through the first embedding, as shown in the following formula: ; in, is the feature at coordinate (i, j), is the first embedded linear layer, is the feature at point i, P ij P i The coordinates of the neighbor points of K, N and C are the two neighbor points of K.

[0008] Furthermore, the central attention mechanism further includes: Based on the global features after the first embedding , the second embedding is obtained as shown in the following formula: ; in, is the global feature after the second embedding, is the global feature after the first embedding.

[0009] Furthermore, in step S4, it also includes: The input for classifying the point cloud is the output coordinates of the last central learning module and Features , using a fully connected layer and ReLU activation function, output the classification information Y of the point cloud image d , as shown in the following formula: ; in, is a fully connected layer, is the activation function, To learn vectors.

[0010] A weakly supervised point cloud segmentation system for robot assembly, comprising: The data preprocessing module is used to input the point cloud, convert the original image into a standardized input, build a model for the image, and provide coordinates and RGB image input for the model; A downsampling module for encoding features based on the normalized input; The center learning module is based on the Transformer algorithm and is used to design a center attention mechanism. It learns point cloud features by extracting global features of neighboring points and sharing them among various neighborhoods to obtain two embedded global features and position encoding modules. The classification module classifies the point cloud based on the results of the central attention mechanism learning and outputs the final result.

[0011] Further, in the center learning module, including: Based on the original image, it is converted into a standardized image to obtain coordinates and features. The standardized coordinates and features are input into the multi-layer perceptron. After a layer of MLP, the features are obtained. and ,in , The following formula: ; ; Among them, F d , P d is the initial feature of the input; Then, the central attention mechanism is used to and To study.

[0012] Further, in the center learning module, including: For each center point , its characteristics Through a linear layer , The dimension is 1; at the same time, the KNN algorithm is used to obtain the center point The coordinates of the k neighbor points , and the features corresponding to the k neighbor points , and then the global features are extracted by integrating the center weights and neighboring point features through the first embedding, as shown in the following formula: ; in, is the feature at coordinate (i, j), is the first embedded linear layer, is the feature at point i, P ij P i The coordinates of the neighbor points of K, N and C are the two neighbor points of K.

[0013] Further, in the center learning module, including: Based on the global features after the first embedding , the second embedding is obtained as shown in the following formula: ; in, is the global feature after the second embedding, is the global feature after the first embedding.

[0014] Furthermore, in the classification module, it includes: The input for classifying the point cloud is the output coordinates of the last central learning module and Features , using a fully connected layer and ReLU activation function, output the classification information Y of the point cloud image d , as shown in the following formula: ; in, is a fully connected layer, is the activation function, To learn vectors.

[0015] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: The present invention aims to solve the above problems by designing a center-based attention mechanism and Transformer architecture, enhancing the feature representation of unlabeled points through two embedding processes and a position encoding module, thereby improving the performance of point cloud segmentation under weak supervision settings, and achieving point cloud segmentation with low manual annotation cost, making training more efficient and cost-effective. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of a weakly supervised point cloud segmentation method for robot assembly; Figure 2 It is a schematic diagram of the sampling module; Figure 3 A diagram of the center learning module; Figure 4 It is the central attention mechanism; Figure 5 Figure 2. A weakly supervised point cloud segmentation system for robotic assembly. DETAILED DESCRIPTION

[0017] The invention is a weakly supervised point cloud segmentation method and system for robot assembly. The core technology is to design a center learning module based on Transformer, extract the global features of the center point, and then share them among various neighborhoods. The local features are enhanced by the method of sharing weights of the center point, while retaining the global features, so as to better complete the point cloud feature learning under sparse annotation. At the same time, position encoding is also introduced into the attention mechanism to supplement the geometric features according to the positions of the center point and its neighboring points. The present invention designs a downsampling module, introduces the farthest point sampling method and pooling operation to better learn the features of different center points.

[0018] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0019] See also Figure 1 ,A weakly supervised point cloud segmentation method for robot assembly, including, S1, input point cloud, convert the original image into standardized input, build a model for the image, and provide coordinates and RGB image input for the model; Specifically, the original image is converted into a standardized input. The three-dimensional coordinates of the object are formatted as ,in It is The coordinates of the points, is the number of points. Format the RGB features of the object as ,in It is The RGB features of each point, is the number of points.

[0020] S2, encoding based on the input features of the normalized input; See also Figure 2 , specifically, the input is the normalized coordinates Features First, this module uses the farthest point sampling algorithm (known method) to sample N points. Then, the KNN algorithm is used to cluster the sampled points and select K nearest neighbors for each point. Then, it passes through a layer of MLP (multi-layer perceptron) and a maximum pooling layer to aggregate the features of these k neighbors as the output of this module. and , as shown in Formula 1: (1); in, are the K nearest neighbor features corresponding to each point, is a multi-layer perceptron, is the maximum pooling layer.

[0021] S3, based on the Transformer algorithm, designs a central attention mechanism to learn point cloud features by extracting global features of neighboring points and sharing them among various neighborhoods, thus obtaining two embedded global features and position encoding modules; Specifically, based on the Transformer algorithm, a central attention mechanism is designed to better learn point cloud features by extracting global features of neighboring points and then sharing them among various neighborhoods. The process of this module is as follows: Figure 3 As shown, the process of the central attention mechanism is as follows Figure 4 shown.

[0022] First, the input of this module is and After one layer of MLP, we get the feature and , , The following formula: ; ; Among them, F d , P d is the initial feature of the input; Secondly, the central attention mechanism and Study as follows.

[0023] The central attention mechanism is: for each center point The present invention is characterized by Through a linear layer , The dimension is 1; at the same time, the KNN algorithm is used to obtain the center point The coordinates of k neighbor points of , and the features corresponding to the k neighbor points , and then the global features are extracted by integrating the center weights and neighboring point features through the first embedding, as shown in Formula 3: (3); P ij P i The coordinates of the neighbor points of K, N and C are the two neighbor points of K.

[0024] in, is the feature at coordinate (i, j), is the first embedded linear layer, is the feature at point i.

[0025] The global features improve the representation of unlabeled neighboring points in the weakly supervised point cloud, but they often lack key local features of the neighborhood. , the second embedding is obtained, as shown in Formula 4: (4); in, is the global feature after the second embedding, is the global feature after the first embedding.

[0026] Through the above global characteristics The center point features are shared with the corresponding neighborhoods. This method regards the global features of the center point as the weight of its neighborhood and ensures that the center point features are effectively propagated to each point in its corresponding neighborhood through matrix multiplication.

[0027] In order to build coverage global features And the attention weight of the center point area, this paper combines the above two embeddings with position encoding, as shown in Formula 5: (5); in, is the matrix dot product, the global feature pass (Linear layer) to transform and use learnable parameters , and This integration enhances local features through the center point sharing method while retaining global features. In addition, position encoding is introduced into the attention weight to supplement the geometric features according to the point position, The calculation method is shown in formulas 6, 7, and 8: (6); (7); (8); in, is the altitude angle, is the azimuth, is the coordinate difference between the neighbor point and the center point, is the Euclidean distance, is MLP (Multi-layer Perceptron), is the bias parameter, is a learnable parameter.

[0028] After introducing position encoding, each center point is finally obtained Output characteristics , as shown in Formula 9: (9); in, is the Softmax function, yes The transpose of is a linear layer, Represents matrix dot product.

[0029] Learning the features of the center point and its neighbors through the center attention mechanism After that, we use MLP (Multi-layer Perceptron) to learn the vector , get the output of this module , as shown in Formula 10: (10); S4, classifies the point cloud based on the results of the central attention mechanism learning and outputs the final result.

[0030] Specifically, the input of the classification is the output of the last central learning module , using a fully connected layer and ReLU activation function, output the classification information Y of the point cloud image d , as shown in Formula 11: (11); in, is a fully connected layer, is the activation function, To learn vectors.

[0031] The local features are enhanced by sharing weights among the center points while retaining the global features, which can better complete the point cloud feature learning in the case of sparse annotation. At the same time, position encoding is also introduced into the attention mechanism to supplement the geometric features according to the positions of the center point and its neighboring points.

[0032] A weakly supervised point cloud segmentation system for robot assembly, comprising: Data preprocessing module 01, used to input point cloud, convert the original image into standardized input, build a model for the image, and provide coordinates and RGB image input for the model; Specifically, this module converts the original image into a standardized input. The three-dimensional coordinates of the object are formatted as ,in It is The coordinates of the points, is the number of points. Format the RGB features of the object as ,in It is The RGB features of each point, is the number of points.

[0033] A downsampling module 02, for encoding features of the input based on the normalized input; See also Figure 2 Specifically, the input of the downsampling module is the normalized coordinates Features First, this module uses the farthest point sampling algorithm (known method) to sample N points. Then, the KNN algorithm is used to cluster the sampled points and select K nearest neighbors for each point. Then, it passes through a layer of MLP (multi-layer perceptron) and a maximum pooling layer to aggregate the features of these k neighbors as the output of this module. and , as shown in Formula 1: (1); in, are the K nearest neighbor features corresponding to each point, is a multi-layer perceptron, is the maximum pooling layer.

[0034] Central learning module 03, based on the Transformer algorithm, is used to design a central attention mechanism to learn point cloud features by extracting global features of neighboring points and sharing them among various neighborhoods, thus obtaining two embedded global features and position encoding modules; Specifically, based on the Transformer algorithm, a central attention mechanism is designed to better learn point cloud features by extracting global features of neighboring points and then sharing them among various neighborhoods. The process of this module is as follows: Figure 3 As shown, the process of the central attention mechanism is as follows Figure 4 shown.

[0035] First, the input of this module is and After one layer of MLP, we get the feature and , , The following formula: ; ; Among them, F d, P dis the initial feature of the input; Secondly, the central attention mechanism and Study as follows.

[0036] The central attention mechanism is: for each center point The present invention is characterized by Through a linear layer , The dimension is 1; at the same time, the KNN algorithm is used to obtain the center point The coordinates of the k neighbor points , and the features corresponding to the k neighbor points , and then the global features are extracted by integrating the center weights and neighboring point features through the first embedding, as shown in Formula 3: (3); The global features improve the representation of unlabeled neighboring points in the weakly supervised point cloud, but they often lack the key local features of the neighborhood. , the second embedding is obtained, as shown in Formula 4: (4); Through the above global characteristics The center point features are shared with the corresponding neighborhoods. This method regards the global features of the center point as the weight of its neighborhood and ensures that the center point features are effectively propagated to each point in its corresponding neighborhood through matrix multiplication.

[0037] In order to build coverage global features And the attention weight of the center point area, this paper combines the above two embeddings with position encoding, as shown in Formula 5: (5); in, is the matrix dot product, the global feature pass (Linear layer) to transform and use learnable parameters , and This integration enhances local features through the center point sharing method while retaining global features. In addition, position encoding is introduced into the attention weight to supplement the geometric features according to the point position, The calculation method is shown in formulas 6, 7, and 8: (6); (7); (8); in, is the altitude angle, is the azimuth, is the coordinate difference between the neighbor point and the center point, is the Euclidean distance, is MLP (Multi-layer Perceptron), is the bias parameter, is a learnable parameter.

[0038] After introducing position encoding, each center point is finally obtained Output characteristics , as shown in Formula 9: (9); in, is the Softmax function, yes The transpose of is a linear layer, Represents matrix dot product.

[0039] Learning the features of the center point and its neighbors through the center attention mechanism After that, we use MLP (Multi-layer Perceptron) to learn the vector , get the output of this module , as shown in Formula 10: (10); Classification module 04 classifies the point cloud based on the results of central attention mechanism learning and outputs the final result.

[0040] Specifically, the input of classification module 04 is the output of the last central learning module , using a fully connected layer and ReLU activation function, output the classification information Y of the point cloud image d , as shown in Formula 11: (11); in, is a fully connected layer, is the activation function, To learn vectors.

[0041] The local features are enhanced by sharing weights among the center points while retaining the global features, which can better complete the point cloud feature learning in the case of sparse annotation. At the same time, position encoding is also introduced into the attention mechanism to supplement the geometric features according to the positions of the center point and its neighboring points.

[0042] We design a center-based attention mechanism and Transformer architecture to address the above problems, and enhance the feature representation of unlabeled points through two embedding processes and a position encoding module, thereby improving the performance of point cloud segmentation under weak supervision settings.

[0043] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0044] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0045] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0046] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0047] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0048] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A weakly supervised point cloud segmentation method for robot assembly, characterized in that: include, S1, input point cloud, convert the original image into standardized input, build a model for the image, and provide coordinates and RGB image input for the model; S2, encoding based on the input features of the normalized input; S3, based on the Transformer algorithm, designs a central attention mechanism to learn point cloud features by extracting global features of neighboring points and sharing them among various neighborhoods, thus obtaining two embedded global features and position encoding modules; S4, classifies the point cloud based on the results of the central attention mechanism learning and outputs the final result.

2. A weakly supervised point cloud segmentation method for robot assembly as claimed in claim 1, characterized in that: In step S3, it includes: S31, based on the original image, is converted into a standardized image to obtain coordinates and features, and the standardized coordinates and features are input. After a layer of MLP in the multi-layer perceptron, the features are obtained. and Features ,in , The following formula: ; ; Among them, F d , P d is the initial feature of the input; S32, the central attention mechanism and To study.

3. A weakly supervised point cloud segmentation method for robot assembly as claimed in claim 2, characterized in that: The central attention mechanism also includes: For each center point , its characteristics Through a linear layer , The dimension is 1; at the same time, the KNN algorithm is used to obtain the center point The coordinates of the k neighbor points , and the features corresponding to the k neighbor points , and then the global features are extracted by integrating the center weights and neighboring point features through the first embedding, as shown in the following formula: ; in, is the feature at coordinate (i, j), is the first embedded linear layer, is the feature at point i, P ij P i The coordinates of the neighbor points of K, N and C are the two neighbor points of K.

4. A weakly supervised point cloud segmentation method for robot assembly as claimed in claim 3, characterized in that: The central attention mechanism also includes: Based on the global features after the first embedding , the second embedding is obtained as shown in the following formula: ; in, is the global feature after the second embedding, is the global feature after the first embedding.

5. A weakly supervised point cloud segmentation method for robot assembly as claimed in claim 1, characterized in that: In step S4, it also includes: The input for classifying the point cloud is the output coordinates of the last central learning module and Features , using a fully connected layer and ReLU activation function, output the classification information Y of the point cloud image d , as shown in the following formula: ; in, is a fully connected layer, is the activation function, To learn vectors.

6. A weakly supervised point cloud segmentation system for robot assembly, characterized in that: include, The data preprocessing module is used to input the point cloud, convert the original image into a standardized input, build a model for the image, and provide coordinates and RGB image input for the model; A downsampling module for encoding features based on the normalized input; The center learning module is based on the Transformer algorithm and is used to design a center attention mechanism. It learns point cloud features by extracting global features of neighboring points and sharing them among various neighborhoods to obtain two embedded global features and position encoding modules. The classification module classifies the point cloud based on the results of the central attention mechanism learning and outputs the final result.

7. A weakly supervised point cloud segmentation system for robot assembly as claimed in claim 6, characterized in that: The central learning modules include: Based on the original image, it is converted into a standardized image to obtain coordinates and features. The standardized coordinates and features are input into the multi-layer perceptron. After a layer of MLP, the features are obtained. and Features ,in , The following formula: ; ; Among them, F d , P d is the initial feature of the input; The central attention mechanism and To study.

8. A weakly supervised point cloud segmentation system for robot assembly as claimed in claim 6, characterized in that: The central learning modules include: For each center point , its characteristics Through a linear layer , The dimension is 1; at the same time, the KNN algorithm is used to obtain the center point The coordinates of the k neighbor points , and the features corresponding to the k neighbor points , and then the global features are extracted by integrating the center weights and neighboring point features through the first embedding, as shown in the following formula: =σ( ×g1( )); in, is the feature at coordinate (i, j), is the first embedded linear layer, is the feature at point i, P ij P i The coordinates of the neighbor points of K, N and C are the two neighbor points of K.

9. A weakly supervised point cloud segmentation system for robot assembly as claimed in claim 8, characterized in that: The central learning modules include: Based on the global features after the first embedding , the second embedding is obtained as shown in the following formula: ; in, is the global feature after the second embedding, is the global feature after the first embedding.

10. A weakly supervised point cloud segmentation system for robot assembly as claimed in claim 6, characterized in that: In the classification module, including: The input for classifying the point cloud is the output coordinates of the last central learning module and Features , using a fully connected layer and ReLU activation function, output the classification information Y of the point cloud image d , as shown in the following formula: ; in, is a fully connected layer, is the activation function, To learn vectors.

Citation Information

Patent Citations

  • Point cloud segmentation method based on global feature learning and local feature discriminant aggregation

    CN115131560A

  • Repeatable single object extraction method based on central point detection and clustering

    CN117078988A

  • Block attention point cloud segmentation method based on curve characteristics

    CN119273709A

  • METHOD FOR UPSAMPLING MEASUREMENT POINTS OF A POINT CLOUD

    DE102023106975A1

  • Deep learning-based high-precision point cloud completion method and apparatus

    WO2024060395A1