Three-dimensional scene understanding method based on large language model
Through a three-dimensional scene understanding method based on a large language model, semantic information and feature vectors are used to process three-dimensional scene data, the problem of lack of semantic relationship understanding and robustness in the understanding of three-dimensional scenes in the prior art is solved, and more efficient three-dimensional scene understanding and interaction are achieved.
Patent Information
- Application Number
- CN202510104278.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
AI Technical Summary
The existing three-dimensional scene understanding technology mainly relies on the three-dimensional coordinate information of objects, lacks understanding of the semantic relationships between objects, and is sensitive to instance segmentation noise, affecting the robustness of the model.
A three-dimensional scene understanding method based on a large language model is adopted, scene data is collected through lidar and cameras, point cloud data is processed using denoising algorithms and point cloud adaptive interpolation modules, three-dimensional semantic scene graphs and semantic relationship feature vectors are generated, feature vectors are extracted in combination with pre-trained encoders, and semantic rich scene representations are generated through a large language model.
Effectively utilize semantic information to improve the understanding ability and robustness of the three-dimensional scene understanding model, and better understand and interact with the three-dimensional environment.
Smart Images

Figure CN119942547A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of three-dimensional scene understanding, and in particular to a three-dimensional scene understanding method based on a large language model. Background Art
[0002] With the booming development of technologies such as virtual reality, robot navigation, and autonomous driving, the demand for 3D scene understanding is growing. 3D scene understanding technology can enable computers to better understand and interact with 3D environments. This is of great significance for improving the level of automation, enhancing user experience, and improving decision-making capabilities. However, existing 3D scene understanding technologies mainly rely on the 3D coordinate information of objects, lack understanding of the semantic relationship between objects, and are sensitive to instance segmentation noise, which affects the robustness of the model.
[0003] Therefore, it has become an urgent problem to propose a 3D scene understanding method based on a large language model to effectively utilize semantic information and improve the model's understanding ability and robustness. Summary of the invention
[0004] In view of this, the present invention provides a three-dimensional scene understanding method based on a large language model to solve the problems existing in the prior art.
[0005] The present invention provides a three-dimensional scene understanding method based on a large language model, comprising:
[0006] Use laser radar and camera to scan and capture images of the scene to obtain three-dimensional point cloud data and multi-view images of the scene;
[0007] Use a denoising algorithm to preprocess the point cloud data, remove the noise in the point cloud, and obtain preprocessed point cloud data;
[0008] Use the point cloud adaptive interpolation module to refine the pre-processed point cloud data to obtain a more refined point cloud representation;
[0009] Use the deep learning model to perform instance segmentation on the refined point cloud to obtain the point cloud data of each object in the scene;
[0010] Generate a three-dimensional semantic scene graph based on the point cloud data of each object using the VL-SAT method, wherein the nodes in the scene graph represent objects in the scene, each node includes a unique identifier and a feature vector of the object, and the edges in the scene graph represent the relationship between the nodes, each edge includes a relationship category and a relationship strength;
[0011] Extracting feature vectors of semantic relationships from a three-dimensional semantic scene graph using a VL-SAT method, wherein the VL-SAT method uses a graph neural network to learn semantic relationships between objects and predict the category and strength of the semantic relationship;
[0012] Use the pre-trained Uni3D encoder and DINOv2 encoder to extract the 3D geometric information feature vector and 2D semantic information feature vector of the object;
[0013] The object’s 3D geometric information feature vector, 2D semantic information feature vector, and semantic relationship feature vector are projected into the embedding space of the large language model through a trainable linear layer to obtain a semantically rich scene representation.
[0014] Select k nearest neighbors of each object according to the scene representation using a k-nearest neighbor algorithm, and construct a subgraph containing the object and its k nearest neighbors, wherein the object relationship in the subgraph is encoded as a triple (object_a, relation_ab, object_b), wherein object_a and object_b represent identifiers of object a and object b, respectively, and relation_ab represents the semantic relationship category between object a and object b;
[0015] Connecting all subgraphs into a flat subgraph sequence, wherein the flat subgraph sequence is used as input to the large language model for generating natural language descriptions and answering questions;
[0016] The large language model is trained using a self-supervised learning or unsupervised learning method according to the flat subgraph sequence to obtain a language model that can understand the three-dimensional scene.
[0017] Preferably, the denoising method is Gaussian filtering, median filtering or bilateral filtering.
[0018] Further preferably, the specific steps of using the point cloud adaptive interpolation module to refine the pre-processed point cloud data are as follows:
[0019] Create two empty collections, one for storing the indexes of the points that have been processed, and the other for storing the points generated after interpolation;
[0020] Traverse each point in the initial point cloud and interpolate. The interpolation method for any point is as follows:
[0021] Find the K nearest neighbor points of the current point;
[0022] Perform 3D Voronoi partitioning using an incremental algorithm based on the K nearest neighbor points to obtain Voronoi polygons, where the vertices of the Voronoi polygons are potential interpolation points;
[0023] Use Wasserstein distance to evaluate the topological difference between the set of K nearest neighbor points and the set containing the Voronoi polygon vertices. If the evaluation result shows that these vertices will not destroy the topological structure, the vertices of the Voronoi polygon are added to the set as interpolation points. Otherwise, perform 2D Voronoi interpolation.
[0024] The points generated after interpolation are merged with the original point cloud to obtain the refined point cloud.
[0025] Further preferably, the number of K nearest neighbor points of the current point is determined according to the sparsity of the point cloud, wherein the higher the curvature of the region, the more nearest neighbor points are selected.
[0026] Further preferably, the specific steps of 2D Voronoi interpolation are as follows:
[0027] Use principal component analysis to project the original K nearest neighbor points onto a two-dimensional plane;
[0028] Use the incremental algorithm to perform Voronoi division on the two-dimensional plane to obtain Voronoi polygons;
[0029] Map the vertices of the two-dimensional Voronoi polygon back to three-dimensional space and add the mapped vertices to the collection as interpolation points.
[0030] Further preferably, the deep learning model is Mask3D or OneFormer3D.
[0031] Further preferably, the specific steps of extracting the 3D geometric information feature vector and the 2D semantic information feature vector of the object are as follows:
[0032] Extracting feature vectors from the point cloud of each object using a pre-trained Uni3D encoder, wherein the Uni3D encoder uses a PointNet++ network to extract point cloud features and generates a feature vector of 3D geometric information of the object through a pooling operation;
[0033] A pre-trained DINOv2 encoder is used to extract a feature vector from a multi-view image of an object, wherein the DINOv2 encoder uses a Transformer network to extract image features and generates a 2D semantic information feature vector of the object through a pooling operation.
[0034] The three-dimensional scene understanding method based on a large language model provided by the present invention can effectively utilize semantic information and improve the understanding ability and robustness of the three-dimensional scene understanding model. The method first collects point cloud data and multi-view image information of the scene, and then processes the point cloud data through a denoising algorithm and a point cloud adaptive interpolation module to obtain more refined point cloud data, wherein the point cloud adaptive interpolation module can solve the problem of sparse point clouds in low curvature areas while keeping the topological structure of the scene unchanged, and then, the VL-SAT method is used to obtain the feature vectors of the three-dimensional semantic scene graph and the semantic relationship, and the pre-trained Uni3D encoder and DINOv2 encoder are used to obtain the 3D geometric information feature vector and the 2D semantic information feature vector of the object, and the feature vector of the semantic relationship, the 3D geometric information feature vector and the 2D semantic information feature vector of the object are projected to the embedding space of the large language model LLM, so that a semantically rich scene representation can be obtained, and then the k-nearest neighbor algorithm is used to obtain a subgraph sequence containing all objects, and finally, the LLM is trained according to the subgraph sequence to obtain a language model that can understand the three-dimensional scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0036] Figure 1 It is a flow chart of a three-dimensional scene understanding method based on a large language model provided by the present invention. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in order to avoid blurring the present invention due to unnecessary details, only the processing steps closely related to the scheme of the present invention are shown in the accompanying drawings, and other details that are not closely related to the present invention are omitted.
[0038] like Figure 1 As shown, the present invention provides a three-dimensional scene understanding method based on a large language model, comprising the following steps:
[0039] S1: Use laser radar and camera to scan and capture images of the scene to obtain 3D point cloud data and multi-view images of the scene;
[0040] S2: Use a denoising algorithm (such as Gaussian filtering, median filtering, bilateral filtering) to preprocess the point cloud data to remove noise in the point cloud and obtain preprocessed point cloud data;
[0041] S3: Use the point cloud adaptive interpolation module to refine the preprocessed point cloud data to obtain a more refined point cloud representation. The point cloud adaptive interpolation module can solve the problem of sparse point clouds in low curvature areas while keeping the topological structure of the scene unchanged.
[0042] The specific steps of S3 are as follows:
[0043] S31: Create two empty sets, one for storing the indexes of the points that have been processed, and the other for storing the points generated after interpolation;
[0044] S32: traverse each point in the initial point cloud and perform interpolation, wherein the interpolation method of any point is as follows:
[0045] S321: Find the K nearest neighbor points of the current point. The selection of the K nearest neighbor points needs to consider the sparsity of the point cloud. Fewer nearest neighbor points can be selected in low curvature areas, while more nearest neighbor points are required in high curvature areas.
[0046] S322: Perform 3D Voronoi partitioning using an incremental algorithm based on K nearest neighbor points to obtain Voronoi polygons, where the vertices of the Voronoi polygons are potential interpolation points;
[0047] S323: Use Wasserstein distance to evaluate the topological structure difference between the set of K nearest neighbor points and the set containing the Voronoi polygon vertices. If the evaluation result shows that these vertices will not destroy the topological structure, the vertices of the Voronoi polygon are added to the set as interpolation points. Otherwise, perform 2D Voronoi interpolation.
[0048] Among them, the specific steps of 2D Voronoi interpolation are as follows:
[0049] S3231: Project the original K nearest neighbor points onto a two-dimensional plane using principal component analysis (PCA);
[0050] S3232: Perform Voronoi partitioning on a two-dimensional plane using an incremental algorithm to obtain Voronoi polygons;
[0051] S3233: Map the vertices of the two-dimensional Voronoi polygon back to three-dimensional space and add the mapped vertices as interpolation points to the set;
[0052] S33: merging the points generated after interpolation with the original point cloud to obtain a refined point cloud;
[0053] S4: Use a deep learning model (such as Mask3D or OneFormer3D) to perform instance segmentation on the refined point cloud to obtain point cloud data for each object in the scene;
[0054] S5: Generate a 3D semantic scene graph based on the point cloud data of each object using the VL-SAT (Visual-Linguistic Semantics Assisted Training) method, wherein the nodes in the scene graph represent objects in the scene, each node contains a unique identifier and a feature vector of the object, and the edges in the scene graph represent the relationship between nodes, such as "on...", "near...", "belongs to...", etc., and each edge contains a relationship category and relationship strength;
[0055] S6: extracting feature vectors of semantic relationships from the three-dimensional semantic scene graph using a VL-SAT method, wherein the VL-SAT method uses a graph neural network to learn semantic relationships between objects and predict the category and strength of the semantic relationship;
[0056] S7: Use the pre-trained Uni3D encoder and DINOv2 encoder to extract the 3D geometric information feature vector and 2D semantic information feature vector of the object;
[0057] S7 specific steps are as follows:
[0058] S71: extracting a feature vector from the point cloud of each object using a pre-trained Uni3D encoder, wherein the Uni3D encoder extracts point cloud features using a PointNet++ network and generates a 3D geometric information feature vector of the object through a pooling operation;
[0059] S72: extracting a feature vector from a multi-view image of the object using a pre-trained DINOv2 encoder, wherein the DINOv2 encoder extracts image features using a Transformer network and generates a 2D semantic information feature vector of the object through a pooling operation;
[0060] S8: The object’s 3D geometric information feature vector, 2D semantic information feature vector, and semantic relationship feature vector are projected into the embedding space of the large language model (LLM) through a trainable linear layer to obtain a semantically rich scene representation;
[0061] S9: Use a k-nearest neighbor algorithm to select k nearest neighbors of each object according to the scene representation, and construct a subgraph containing the object and its k nearest neighbors, wherein the object relationship in the subgraph is encoded as a triple (object_a, relation_ab, object_b), wherein object_a and object_b represent identifiers of object a and object b respectively, and relation_ab represents the semantic relationship category between object a and object b;
[0062] S10: connect all subgraphs into a flat subgraph sequence, wherein the flat subgraph sequence is used as input to a large language model (LLM) for generating natural language descriptions and answering questions;
[0063] S11: Train the large language model (LLM) using a self-supervised learning or unsupervised learning method according to the flat subgraph sequence to obtain a language model that can understand the three-dimensional scene.
[0064] The three-dimensional scene understanding method based on a large language model provided by the present invention can effectively utilize semantic information and improve the understanding ability and robustness of the three-dimensional scene understanding model. The method first collects point cloud data and multi-view image information of the scene, and then processes the point cloud data through a denoising algorithm and a point cloud adaptive interpolation module to obtain more refined point cloud data, wherein the point cloud adaptive interpolation module can solve the problem of sparse point clouds in low curvature areas while keeping the topological structure of the scene unchanged, and then, the VL-SAT method is used to obtain the feature vectors of the three-dimensional semantic scene graph and the semantic relationship, and the pre-trained Uni3D encoder and DINOv2 encoder are used to obtain the 3D geometric information feature vector and the 2D semantic information feature vector of the object, and the feature vector of the semantic relationship, the 3D geometric information feature vector and the 2D semantic information feature vector of the object are projected to the embedding space of the large language model LLM, so that a semantically rich scene representation can be obtained, and then the k-nearest neighbor algorithm is used to obtain a subgraph sequence containing all objects, and finally, the LLM is trained according to the subgraph sequence to obtain a language model that can understand the three-dimensional scene, and then the three-dimensional scene can be understood using the above-mentioned language model.
[0065] It should be noted that the purpose of publishing the embodiments is to help further understand the present invention, but those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the contents disclosed in the embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined in the claims.
Claims
1. A three-dimensional scene understanding method based on a large language model, characterized in that: include: Use laser radar and camera to scan and capture images of the scene to obtain three-dimensional point cloud data and multi-view images of the scene; Use a denoising algorithm to preprocess the point cloud data, remove the noise in the point cloud, and obtain preprocessed point cloud data; Use the point cloud adaptive interpolation module to refine the pre-processed point cloud data to obtain a more refined point cloud representation; Use the deep learning model to perform instance segmentation on the refined point cloud to obtain the point cloud data of each object in the scene; Generate a three-dimensional semantic scene graph based on the point cloud data of each object using the VL-SAT method, wherein the nodes in the scene graph represent objects in the scene, each node includes a unique identifier and a feature vector of the object, and the edges in the scene graph represent the relationship between the nodes, each edge includes a relationship category and a relationship strength; Extracting feature vectors of semantic relationships from a three-dimensional semantic scene graph using a VL-SAT method, wherein the VL-SAT method uses a graph neural network to learn semantic relationships between objects and predict the category and strength of the semantic relationship; Use the pre-trained Uni3D encoder and DINOv2 encoder to extract the 3D geometric information feature vector and 2D semantic information feature vector of the object; The object’s 3D geometric information feature vector, 2D semantic information feature vector, and semantic relationship feature vector are projected into the embedding space of the large language model through a trainable linear layer to obtain a semantically rich scene representation. Select k nearest neighbors of each object according to the scene representation using a k-nearest neighbor algorithm, and construct a subgraph containing the object and its k nearest neighbors, wherein the object relationship in the subgraph is encoded as a triple (object_a, relation_ab, object_b), wherein object_a and object_b represent identifiers of object a and object b, respectively, and relation_ab represents the semantic relationship category between object a and object b; Connecting all subgraphs into a flat subgraph sequence, wherein the flat subgraph sequence is used as input to the large language model for generating natural language descriptions and answering questions; The large language model is trained using a self-supervised learning or unsupervised learning method according to the flat subgraph sequence to obtain a language model that can understand the three-dimensional scene.
2. The three-dimensional scene understanding method based on a large language model according to claim 1, characterized in that: The denoising method is Gaussian filtering, median filtering or bilateral filtering.
3. The three-dimensional scene understanding method based on a large language model according to claim 1, characterized in that: The specific steps to use the point cloud adaptive interpolation module to refine the preprocessed point cloud data are as follows: Create two empty collections, one for storing the indexes of the points that have been processed, and the other for storing the points generated after interpolation; Traverse each point in the initial point cloud and interpolate. The interpolation method for any point is as follows: Find the K nearest neighbor points of the current point; Perform 3D Voronoi partitioning using an incremental algorithm based on the K nearest neighbor points to obtain Voronoi polygons, where the vertices of the Voronoi polygons are potential interpolation points; Use Wasserstein distance to evaluate the topological difference between the set of K nearest neighbor points and the set containing the Voronoi polygon vertices. If the evaluation result shows that these vertices will not destroy the topological structure, the vertices of the Voronoi polygon are added to the set as interpolation points. Otherwise, perform 2D Voronoi interpolation. The points generated after interpolation are merged with the original point cloud to obtain the refined point cloud.
4. The three-dimensional scene understanding method based on a large language model according to claim 3, characterized in that: The number of K nearest neighbor points of the current point is determined according to the sparsity of the point cloud. The higher the curvature of the region, the more nearest neighbor points are selected.
5. The three-dimensional scene understanding method based on a large language model according to claim 3, characterized in that: The specific steps of 2DVoronoi interpolation are as follows: Use principal component analysis to project the original K nearest neighbor points onto a two-dimensional plane; Use the incremental algorithm to perform Voronoi division on the two-dimensional plane to obtain Voronoi polygons; Map the vertices of the two-dimensional Voronoi polygon back to three-dimensional space and add the mapped vertices to the collection as interpolation points.
6. The three-dimensional scene understanding method based on a large language model according to claim 1, characterized in that: The deep learning model is Mask3D or OneFormer3D.
7. The three-dimensional scene understanding method based on a large language model according to claim 1, characterized in that: The specific steps of extracting the 3D geometric information feature vector and 2D semantic information feature vector of the object are as follows: Extracting feature vectors from the point cloud of each object using a pre-trained Uni3D encoder, wherein the Uni3D encoder uses a PointNet++ network to extract point cloud features and generates a feature vector of 3D geometric information of the object through a pooling operation; A pre-trained DINOv2 encoder is used to extract a feature vector from a multi-view image of an object, wherein the DINOv2 encoder uses a Transformer network to extract image features and generates a 2D semantic information feature vector of the object through a pooling operation.