Method and system for guiding three-dimensional point cloud robot based on natural language

Through a robot system based on the Transformer architecture, visual images are converted into three-dimensional point clouds and integrated with natural language instructions, which solves the problem of robots understanding and performing complex three-dimensional operations and achieves more efficient operation capabilities and action prediction accuracy.

CN119927932BActive Publication Date: 2025-09-12NINGDE SKEQI INTELLIGENT EQUIP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510433791.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-09-12
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Robots need to accurately understand the three-dimensional structure and spatial relationships in their working environment to perform complex three-dimensional spatial operation tasks such as assembly, handling, and quality inspection, but existing technologies make it difficult to effectively integrate natural language commands and three-dimensional point cloud data.

Method used

A robot system based on the Transformer architecture is used to convert visual image data into three-dimensional point clouds, extract spatial features, and combine them with natural language instructions for vector embedding. The attention mechanism is used for information fusion to predict the robot's future action position.

Benefits of technology

It improves the robot's ability to understand and execute complex instructions, enhances the accuracy of predicting future actions, and improves production efficiency and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119927932B_ABST
    Figure CN119927932B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for guiding a three-dimensional point cloud robot based on natural language , A robot based on the Transformer architecture sets an action position t, converts the visual image data of the action position t into a three-dimensional point cloud and a standardized input, and performs downsampling to complete data preprocessing; based on data preprocessing, the point cloud of the generated preprocessed data is encoded, the spatial features of the point cloud are extracted, and visual information is generated; and by performing vector embedding on natural language instructions, the natural language instructions are represented as vectors that the model can understand and process, thereby generating text information; based on the visual information and text information, the generated visual information and contextual information are fused through the attention mechanism; based on the fusion of contextual information, the three-dimensional position of the action position #imgabs0# step is predicted by predicting the heat map and the offset, thereby improving the robot's ability to understand and execute complex instructions and the accuracy of the robot's future action prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart factories, and in particular to a method and system for guiding a three-dimensional point cloud robot based on natural language. Background Art

[0002] Smart factories, at the forefront of industrial automation and intelligent manufacturing, are constantly integrating the latest artificial intelligence technologies to improve production efficiency, reduce human intervention, and enhance the flexibility and adaptability of production lines. Developing robots that can perform various tasks based on natural language instructions is crucial for smart factories. Therefore, a robot guidance system based on natural language instructions is needed to improve production efficiency.

[0003] However, in the process of implementing the technical solutions of the embodiments of the present application, the inventors of the present application discovered that the above technology has at least the following technical problems:

[0004] Robots need to accurately understand the three-dimensional structures and spatial relationships in their working environment, as they often need to perform precise operations in complex three-dimensional spaces, such as assembly, handling, and quality inspection. Therefore, enabling robots to understand and execute complex tasks based on natural language instructions while also being capable of precise three-dimensional manipulation poses a challenge for researchers. Summary of the Invention

[0005] The embodiments of the present application provide a method and system for guiding a three-dimensional point cloud robot based on natural language, thereby solving the problem in the prior art of improving the performance and efficiency of robot operation tasks by using three-dimensional point cloud representation, encoders and multimodal converters, and effective integration with natural language instructions.

[0006] The embodiment of the present application provides a method for guiding a three-dimensional point cloud robot based on natural language, comprising:

[0007] S1, a robot based on the Transformer architecture, sets an action position t, converts the visual image data of the action position t into a 3D point cloud and standardized input, and performs downsampling to complete data preprocessing;

[0008] S2, based on data preprocessing, encodes the point cloud of the generated preprocessed data, extracts the spatial features of the point cloud, and generates visual information; and by embedding the natural language instructions into vectors that the model can understand and process, it generates text information;

[0009] S3, based on visual information and text information, fuses the generated visual information and contextual information through the attention mechanism;

[0010] S4, based on the fusion of context information, predicts the 3D position of the action position t+1 step by predicting the heat map and offset.

[0011] Furthermore, in step S1, the data preprocessing includes:

[0012] Extract useful features from the point cloud, including XYZ coordinates, RGB colors, and normals, and merge the point clouds to generate a merged point cloud;

[0013] The merged point cloud is uniformly downsampled and the normal of each point is estimated through the Open3D toolkit to generate a new point cloud;

[0014] Crop the new point cloud to keep only the points of the object and the robotic arm;

[0015] Only the points of the object and the robotic arm will be retained and 2048 points will be sampled through random sampling.

[0016] Furthermore, in step S2, the visual information is used to learn point cloud information based on the PointNext model, including:

[0017] By using the farthest point sampling method, from the point cloud Collect N points ;

[0018] Set Radius , collect the corresponding radius for each point All neighbor nodes within;

[0019] Learn each point through MLP Corresponding point cloud information ; Aggregate all neighborhood information of each point through maximum pooling to obtain all point cloud information ; The i is a point in the point cloud.

[0020] Furthermore, in step S2, the generation of the text information includes:

[0021] By freezing the encoder and adding a linear layer to the frozen encoder to obtain the embedding vector ,

[0022] The following formula 3 is shown:

[0023] (3);

[0024] in, is the weight, is a linear layer, is the input natural language instruction, and CLIP is the CLIP model.

[0025] Furthermore, in step S3, it includes:

[0026] S31, the input is point cloud information and natural language instruction information; the visual information is learned through the attention mechanism, as shown in the following formulas 4 and 5:

[0027] (4);

[0028] (5);

[0029] in, It is the attention formula;

[0030] is the weight;

[0031] is the hidden layer size;

[0032] S32, the visual information and natural language instruction information are integrated through the attention mechanism, as shown in the following formula 6:

[0033] (6);

[0034] S33, by stacking 1 layer, finally get the output , as shown in Formula 7:

[0035] (7);

[0036] in, is the weight, is the activation function, is the normalization layer.

[0037] A system for guiding 3D point cloud robots based on natural language, including:

[0038] The data preprocessing module is based on the Transformer architecture of the robot and is used to set the action position t, convert the visual image data of the action position t into a 3D point cloud and standardized input, and perform downsampling to complete the data preprocessing;

[0039] The visual module and the text module are based on data preprocessing and are used to encode the point cloud of the generated preprocessed data, extract the spatial features of the point cloud, and generate visual information. They also embed natural language instructions into vectors that the model can understand and process, thereby generating text information.

[0040] The fusion module is based on visual information and text information and is used to fuse the generated visual information and contextual information through the attention mechanism;

[0041] The action module, based on the fusion of contextual information, is used to predict the 3D position of the action position t+1 steps by predicting the heatmap and offset.

[0042] Furthermore, the data preprocessing module includes:

[0043] The electric cloud generation unit is used to extract useful features from the point cloud, including XYZ coordinates, RGB colors, and normals, and merge the point clouds to generate a merged point cloud;

[0044] A new point cloud generation unit is used to generate a new point cloud by uniformly downsampling the merged point cloud through the Open3D toolkit and estimating the normal of each point;

[0045] The electric cloud clipping unit is used to clip the new point cloud to keep only the points of the object and the robotic arm;

[0046] The electric cloud sampling unit is used to sample 2048 points by random sampling method, retaining only the points of the object and the robotic arm.

[0047] Furthermore, in the visual module, it includes:

[0048] By using the farthest point sampling method, from the point cloud Collect N points ;

[0049] Set Radius , collect the corresponding radius for each point All neighbor nodes within;

[0050] Learn each point through MLP Corresponding point cloud information ; Aggregate all neighborhood information of each point through maximum pooling to obtain all point cloud information , where i is a point in the point cloud.

[0051] Furthermore, in the text module, include:

[0052] By freezing the encoder and adding a linear layer to the frozen encoder to obtain the embedding vector ,

[0053] The following formula 3 is shown;

[0054] (3);

[0055] in, is the weight, is a linear layer, is the input natural language instruction, and CLIP is the CLIP model.

[0056] Furthermore, the fusion module includes:

[0057] The attention mechanism learning unit takes point cloud information and natural language instruction information as input; it learns visual information through the attention mechanism, as shown in the following formulas 4 and 5:

[0058] (4);

[0059] (5);

[0060] in, It is the attention formula;

[0061] is the weight;

[0062] is the hidden layer size;

[0063] The attention mechanism fusion unit is used to fuse visual information and natural language instruction information through the attention mechanism, as shown in the following formula 6:

[0064] (6);

[0065] Attention mechanism output unit, by stacking 1 layer, finally get the output , as shown in Formula 7:

[0066] (7);

[0067] in, is the weight, is the activation function, is the normalization layer.

[0068] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0069] 1. By converting the RGB images of the multi-view camera into 3D point clouds for modeling, a Transformer-based encoder is introduced to represent the features of the point cloud data. At the same time, the point cloud features and language embedding are fused based on the attention mechanism to achieve the fusion of visual and language information. This enables the model to comprehensively consider visual and language information, improving the robot's ability to understand and execute complex instructions.

[0070] 2. Thanks to the design of an action decoding module, the robot’s position is estimated by predicting heat maps and offsets, which improves the accuracy of predicting the robot’s future actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1A flow chart of the method for guiding a 3D point cloud robot based on natural language;

[0072] Figure 2 A module diagram for guiding a 3D point cloud robot based on natural language;

[0073] Figure 3 This is the flow chart of the vision module, text module, fusion module, and action module. DETAILED DESCRIPTION

[0074] The present invention provides a method and system for guiding a three-dimensional point cloud robot based on natural language. A robot operating system based on the Transformer architecture is designed to improve the robot's production efficiency by integrating visual and language information.

[0075] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0076] The goal of this invention is to learn a visual strategy through the model , allowing the robot to perform manipulation tasks according to natural language instructions, , , is the visual and action at step t, and are visual and action spaces respectively. In the present invention, the visual space Contains 1) natural language instructions ,in Represents the vectorized words. 2) RGB image 3) Depth Image In the present invention, According to existing research, the size is set to 128. The number of cameras is set to 3 in the present invention, which respectively capture the robot's left shoulder, right shoulder and wrist. Contains rotation posture and open and closed state, where the rotation posture is represented by Cartesian coordinates and quaternions Composition, of which is a three-dimensional vector space, The action coordinates are in a four-dimensional vector space, specifically, the action coordinates are in a three-dimensional vector space, represented by (x, y, z), where x, y, z are real numbers, representing the components of the vector on the three coordinate axes, respectively. The action coordinates are in a four-dimensional vector space, w, x, y, z are real numbers, representing the components of the vector on the four coordinate axes, respectively. The final output of the present invention is to predict the robot's state at step t+1 based on visual information and natural language instruction information, thereby realizing a robot guidance system based on natural language and three-dimensional point clouds, which guides the robot to operate according to natural language instructions.

[0077] See also Figure 1 , an embodiment of the present application provides a method for guiding a three-dimensional point cloud robot based on natural language, comprising:

[0078] S1, a robot based on the Transformer architecture, sets an action position t, converts the visual image data of the action position t into a three-dimensional point cloud and standardized input, and performs downsampling to complete data preprocessing; the Transformer is a neural network;

[0079] Specifically, images from multi-view cameras are converted into 3D point clouds and downsampled to reduce redundant data. After conversion to a 3D point cloud, this module extracts useful features from the point cloud, including XYZ coordinates, RGB colors, and normals. The point cloud is also normalized to provide data support for subsequent model learning.

[0080] Specifically, given the visual image at step t , the present invention converts each pixel into a three-dimensional coordinate, because the RGB image corresponds to the depth visual image, the RGB pixel value can be superimposed on the depth image, so a point cloud is constructed. , including XYZ coordinates and RGB images. For images with different perspectives, the present invention merges their point clouds (three depth images with different perspectives are merged into one depth image). In order to reduce redundancy, the merged point cloud is uniformly downsampled using the Open3D toolkit (a toolkit publicly available on the Internet), and the normal of each point is estimated by Open3D. Afterwards, the present invention crops the point cloud and only retains the points of the object and the robotic arm, which can improve the effectiveness of the point cloud representation. Finally, using the random sampling method, 2048 points are randomly sampled from each point cloud as the output of the data preprocessing module .

[0081] S2, based on data preprocessing, encodes the point cloud of the generated preprocessed data, extracts the spatial features of the point cloud, and generates visual information; and by embedding the natural language instructions into vectors that the model can understand and process, it generates text information;

[0082] Specifically, the visual information is mainly based on learning point cloud information based on the PointNext model, where PointNext is a neural network model in the field of point cloud understanding. The specific process is as follows;

[0083] Using the farthest point sampling method, from the point cloud Collect N points Used for subsequent learning of point cloud information.

[0084] Set the radius r and collect all neighbor nodes within the corresponding radius r for each point.

[0085] Use MLP (Multi-layer Perceptron) to learn each point Corresponding point cloud information , as shown in Formula 1:

[0086] , (1);

[0087] in, Yes According to the radius Get all neighbors, are the coordinates of point i, are the coordinates of point j, It is vector concatenation and MLP is multi-layer perceptron.

[0088] Max pooling is used to aggregate all neighborhood information of each point, as shown in Formula 2:

[0089] (2);

[0090] Get all point cloud information ;

[0091] The i is a point in the point cloud.

[0092] Specifically, the text information uses the encoder in the CLIP model to mark and encode natural language instructions, thereby converting the natural language instructions into an embedding vector. The CLIP model is pre-trained on large-scale image-text pairs and can effectively understand visual-related instructions. The CLIP model is a multimodal pre-training model. The present invention freezes the encoder and adds a linear layer to it to obtain the embedding vector. , as shown in Formula 3:

[0093] (3);

[0094] in, is the weight, is a linear layer, is the input natural language instruction, and CLIP is the CLIP model.

[0095] S3, based on visual information and text information, fuses the generated visual information and contextual information through the attention mechanism;

[0096] Specifically, the attention mechanism is used to fuse the three-dimensional point cloud information of visual information with the natural language instruction information of text information, so that the model can better understand the correlation between text information and visual information.

[0097] The input of this module is point cloud information and natural language instruction information , the specific steps are as follows;

[0098] S31, the input is point cloud information and natural language instruction information; the visual information is learned through the attention mechanism, as shown in the following formulas 4 and 5:

[0099] (4);

[0100] (5);

[0101] in, It is the attention formula;

[0102] is the weight;

[0103] is the hidden layer size;

[0104] S32, the visual information and natural language instruction information are integrated through the attention mechanism, as shown in the following formula 6:

[0105] (6);

[0106] S33, by stacking 1 layer, finally get the output , as shown in Formula 7:

[0107] (7);

[0108] in, is the weight, is the activation function, is the normalization layer.

[0109] S4, based on the fusion of context information, predicts the 3D position of the action position t+1 step by predicting the heat map and offset.

[0110] Specifically, the heat map of the point cloud is generated by using the PointNext decoder , and the offset of each point The point cloud contains the entire workspace, so the 3D position at step t+1 can be predicted as shown in Equation 8:

[0111] ;

[0112] At the same time, the rotation state and opening and closing state represented by the point cloud at step t+1 can be predicted, as shown in Formula 9:

[0113] (9);

[0114] in, is the rotation state, It is open and closed. is the output of the vision module, is the output of the fusion module.

[0115] The loss function is calculated as shown in Formula 10:

[0116] (10);

[0117] in, is a dataset, is a natural language instruction, is included Point cloud information of key steps and actions A collection of . is the original value, , is the predicted value. is the mean square error loss function, is the cross entropy loss function.

[0118] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:

[0119] By converting RGB images from a multi-view camera into 3D point clouds for modeling, a Transformer-based encoder was introduced to represent the features of the point cloud data. An attention mechanism was used to fuse point cloud features with language embeddings, achieving a fusion of visual and language information. This enabled the model to comprehensively consider both visual and language information, improving the robot's ability to understand and execute complex instructions. An action decoding module was designed to estimate the robot's position by predicting heat maps and offsets, improving the accuracy of predicting the robot's future actions.

[0120] A system for guiding 3D point cloud robots based on natural language, including:

[0121] See also Figure 2 、 3 ,Data preprocessing module, based on the Transformer architecture robot, is used to set the action position t, convert the visual image data of the action position t into a three-dimensional point cloud and standardized input, and perform downsampling to complete data preprocessing;

[0122] The visual module and the text module are based on data preprocessing and are used to encode the point cloud of the generated preprocessed data, extract the spatial features of the point cloud, and generate visual information. They also embed natural language instructions into vectors that the model can understand and process, generating text information.

[0123] The fusion module is based on visual information and text information and is used to fuse the generated visual information and contextual information through the attention mechanism;

[0124] The action module, based on the fusion of contextual information, is used to predict the 3D position of the action position t+1 steps by predicting the heatmap and offset.

[0125] Furthermore, the data preprocessing module includes:

[0126] The point cloud generation unit is used to extract useful features from the point cloud, including XYZ coordinates, RGB colors, and normals, and merge the point clouds to generate a merged point cloud;

[0127] A new point cloud generation unit is used to generate a new point cloud by uniformly downsampling the merged point cloud through the Open3D toolkit and estimating the normal of each point;

[0128] The electric cloud clipping unit is used to clip the new point cloud to keep only the points of the object and the robotic arm;

[0129] The electric cloud sampling unit is used to sample 2048 points by random sampling method, retaining only the points of the object and the robotic arm.

[0130] Furthermore, in the visual module, it includes:

[0131] By using the farthest point sampling method, from the point cloud Collect N points ;

[0132] Set Radius , collect the corresponding radius for each point All neighbor nodes within;

[0133] Learn each point through MLP Corresponding point cloud information ; Aggregate all neighborhood information of each point through maximum pooling to obtain all point cloud information , where i is a point in the point cloud.

[0134] Furthermore, in the text module, include:

[0135] By freezing the encoder and adding a linear layer to the frozen encoder to obtain the embedding vector ,

[0136] The following formula 3 is shown:

[0137] (3);

[0138] in, is the weight, is a linear layer, is the input natural language instruction, and CLIP is the CLIP model.

[0139] Furthermore, the fusion module includes:

[0140] The attention mechanism learning unit takes point cloud information and natural language instruction information as input; it learns visual information through the attention mechanism, as shown in the following formulas 4 and 5:

[0141] (4);

[0142] (5);

[0143] in, It is the attention formula;

[0144] is the weight;

[0145] is the hidden layer size;

[0146] The attention mechanism fusion unit is used to fuse visual information and natural language instruction information through the attention mechanism, as shown in the following formula 6:

[0147] (6);

[0148] Attention mechanism output unit, by stacking 1 layer, finally get the output , as shown in Formula 7:

[0149] (7);

[0150] in, is the weight, is the activation function, is the normalization layer.

[0151] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0152] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0153] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0155] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0156] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A method for guiding a three-dimensional point cloud robot based on natural language, characterized in that: include, S1, a robot based on the Transformer architecture, sets an action position t, converts the visual image data of the action position t into a 3D point cloud and normalized input, and performs downsampling to complete data preprocessing; extracts useful features from the point cloud, including XYZ coordinates, RGB color, and normals, and merges the point clouds to generate a merged point cloud; The merged point cloud is uniformly downsampled and the normal of each point is estimated through the Open3D toolkit to generate a new point cloud; Crop the new point cloud to keep only the points of the object and the robotic arm; Only the points of the object and the robotic arm are retained and 2048 points are sampled through random sampling; S2, based on data preprocessing, encodes the point cloud of the generated preprocessed data, extracts the spatial features of the point cloud, and generates visual information; By embedding natural language instructions into vectors, the natural language instructions are represented as vectors that the model can understand and process, generating text information; The visual information is based on the PointNext model to learn point cloud information. By using the farthest point sampling method, from the point cloud Collect N points ; Set Radius , collect the corresponding radius for each point All neighbor nodes within; Learn each point through MLP Corresponding point cloud information ; Aggregate all neighborhood information of each point through maximum pooling to obtain all point cloud information ; The i is a point in the point cloud; S3, based on visual information and text information, fuses the generated visual information and contextual information through the attention mechanism; S4, based on the fusion of context information, predicts the 3D position of the action position t+1 step by predicting the heat map and offset.

2. The method for guiding a three-dimensional point cloud robot based on natural language according to claim 1, characterized in that: In step S2, the generation of the text information includes: By freezing the encoder and adding a linear layer to the frozen encoder to obtain the embedding vector , The following formula 3 shows: (3); in, is the weight, is a linear layer, is the input natural language instruction, and CLIP is the CLIP model.

3. The method for guiding a three-dimensional point cloud robot based on natural language according to claim 1, wherein: In step S3, it includes: S31, the input is point cloud information and natural language instruction information; the visual information is learned through the attention mechanism, as shown in the following formulas 4 and 5: (4); (5); in, It is the attention formula; is the weight; is the hidden layer size; S32, the visual information and natural language instruction information are integrated through the attention mechanism, as shown in the following formula 6: (6); S33, by stacking 1 layer, finally get the output , as shown in Formula 7: (7); in, is the weight, is the activation function, is the normalization layer.

4. A system for guiding a 3D point cloud robot based on natural language, characterized in that: include, The data preprocessing module, based on the Transformer architecture of the robot, is used to set the action position t, convert the visual image data of the action position t into a 3D point cloud and normalized input, and perform downsampling to complete the data preprocessing. It extracts useful features from the point cloud, including XYZ coordinates, RGB color, and normals, and merges the point clouds to generate a merged point cloud. The merged point cloud is uniformly downsampled and the normal of each point is estimated through the Open3D toolkit to generate a new point cloud; Crop the new point cloud to keep only the points of the object and the robotic arm; Only the points of the object and the robotic arm are retained and 2048 points are sampled through random sampling; The visual module and text module are based on data preprocessing and are used to encode the point cloud of the generated preprocessed data, extract the spatial features of the point cloud, and generate visual information; And by embedding the natural language instructions into vectors that the model can understand and process, the natural language instructions are represented as vectors to generate text information; the visual information is based on the PointNext model to learn point cloud information, By using the farthest point sampling method, from the point cloud Collect N points ; Set Radius , collect the corresponding radius for each point All neighbor nodes within; Learn each point through MLP Corresponding point cloud information ; Aggregate all neighborhood information of each point through maximum pooling to obtain all point cloud information ; The i is a point in the point cloud; The fusion module is based on visual information and text information and is used to fuse the generated visual information and contextual information through the attention mechanism; The action module, based on the fusion of contextual information, is used to predict the 3D position of the action position at step t+1 by predicting the heatmap and offset.

5. The system for guiding a three-dimensional point cloud robot based on natural language according to claim 4, characterized in that: In the text module, include: By freezing the encoder and adding a linear layer to the frozen encoder to obtain the embedding vector , The following formula 3 shows: (3); in, is the weight, is a linear layer, is the input natural language instruction, and CLIP is the CLIP model.

6. The system for guiding a three-dimensional point cloud robot based on natural language according to claim 4, characterized in that: In the fusion module, including: The attention mechanism learning unit takes point cloud information and natural language instruction information as input; it learns visual information through the attention mechanism, as shown in the following formulas 4 and 5: (4); (5); in, is the attention formula: is the weight; is the hidden layer size; The attention mechanism fusion unit is used to fuse visual information and natural language instruction information through the attention mechanism, as shown in the following formula 6: (6); The attention mechanism output unit is stacked in 1 layer to obtain the output. , as shown in Formula 7: (7); in, is the weight, is the activation function, is the normalization layer.

Citation Information

Patent Citations

  • Visual language navigation method based on double semantic comprehension and fusion

    CN116429111A

  • Robot motion skill learning method fusing text instruction and motion information

    CN117428780A

  • Human body point cloud analysis method based on millimeter wave radar

    CN119418071A