Indoor service robot navigation planning method based on open semantic mapping and large language model
By constructing a TSDF voxel map based on open semantic mapping and a large language model, and combining it with a semantic embedding vector dictionary, the adaptability and real-time performance issues of indoor service robot navigation methods in complex environments are solved, achieving accurate response to user commands and efficient navigation planning.
Patent Information
- Application Number
- CN202511210038.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-12-09
AI Technical Summary
Existing indoor service robot navigation methods cannot generalize to adapt to the diversity of complex indoor environments, cannot correctly understand and respond to user commands that do not contain predefined keywords, and have poor real-time performance in semantic mapping and semantic retrieval.
We employ an open semantic mapping and large language model approach. We use the SEEM model for semantic segmentation, construct a TSDF voxel map, combine the large language model to output the target object corresponding to the user command, plan the path on the two-dimensional grid map, and realize real-time semantic reconstruction and navigation using the TSDF voxel map and semantic embedding vector dictionary.
It has improved its adaptability to complex indoor environments, can correctly understand and respond to user commands, realizes real-time 3D semantic reconstruction and navigation planning, and improves the efficiency of object query and the real-time performance of navigation.
Smart Images

Figure CN121095337A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of path planning technology, and in particular to an indoor service robot navigation planning method based on open semantic mapping and large language models. Background Technology
[0002] Existing indoor service robot navigation methods mainly include four steps: map building, available navigation endpoint marking, robot global localization, and path planning. Map building utilizes LiDAR and cameras to collect environmental data and constructs an environmental geometric map based on SLAM (Simultaneous Localization and Mapping) technology. Available navigation endpoint marking involves manually defining a series of reachable poses for the indoor service robot, calibrating these poses, and recording them as coordinates on the environmental geometric map. To reduce reliance on manual labeling, some advanced methods introduce semantic mapping technology. Typical semantic maps are based on point cloud maps, obtaining object category labels from RGB images using object detection technology. Multi-view fusion methods are used to store the category semantics as point cloud attributes, enabling 3D semantic labeling of the scene, which serves as prior information for subsequent navigation planning. The introduction of semantic maps allows target selection in navigation planning to evolve from purely manual setting to using keyword matching technology to match user commands with scene objects, achieving a semi-automatic, mechanical response. However, due to the limitations of object detection and keyword matching technologies, on the one hand, the semantic labels in semantic maps need to be manually defined. For object types not included in the training set of the object detection model, the detector will be completely unable to recognize the object. On the other hand, the matching of fixed semantic labels and user commands requires plaintext correspondence. When the user command is not an explicit object type name or the object description is not within the semantic label set, a match is basically impossible. Furthermore, since semantic maps are typically based on point cloud maps for storing semantic information, the spatial density and retrieval complexity of point clouds limit the real-time performance of semantic mapping and semantic queries, making it difficult to meet the requirements of users for rapid response from indoor service robots in service scenarios. Summary of the Invention
[0003] To address the problems of existing indoor service robot navigation methods, such as their inability to generalize and adapt to the diversity of object types and attributes in complex indoor environments, their inability to correctly understand and respond to user commands that do not contain predefined keywords, and their poor real-time performance in semantic mapping and semantic retrieval, this invention proposes an indoor service robot navigation planning method based on open semantic mapping and a large language model.
[0004] The technical solution adopted in this invention is:
[0005] It includes the following steps:
[0006] S1. Given a three-dimensional indoor scene, command an indoor service robot equipped with a depth camera to take pictures of the indoor scene to obtain a video containing the structure of the indoor scene. Based on the video, obtain multiple frames of RGB images and multiple frames of depth images, with each frame of RGB image and each frame of depth image covering the same scene.
[0007] Ten consecutive RGB images are grouped together to obtain multiple sets of RGB images, and ten consecutive depth images are grouped together to obtain multiple sets of depth images.
[0008] S2. Use the SEEM model to perform semantic segmentation on the first frame of RGB images in each group of RGB images, and obtain the CLIP semantic embedding vector and the confidence distribution of semantic embedding vector corresponding to each segmented region in the first frame of RGB images in each group of RGB images.
[0009] S3. Based on the content of the first frame RGB image and all CLIP semantic embedding vectors in the first frame RGB image, associate a semantic embedding vector dictionary key-value pair with the CLIP semantic embedding vector of each object in the first frame RGB image. Construct a semantic embedding vector dictionary based on all CLIP semantic embedding vectors and all semantic embedding vector dictionary key-value pairs in the first frame RGB image. ;
[0010] S4. Based on the first set of RGB images and the first set of depth images, construct a TSDF voxel map using TSDF voxel grids. ;
[0011] S5. Based on S4, update the TSDF voxel map sequentially with each subsequent RGB image and depth image from the first set of RGB images and the first set of depth images. and semantic embedding vector dictionary ;
[0012] Simultaneously, starting from the second group of RGB images, the depth camera pose and depth camera intrinsics corresponding to the first frame of RGB images in each group of RGB images are obtained. Based on the first frame of RGB images, as well as the depth image, depth camera pose, depth camera intrinsics, and all CLIP semantic embedding vectors and semantic embedding vector confidence distributions on the first frame of RGB images, the TSDF voxel map and semantic embedding vector dictionary obtained from all groups of RGB images and depth images before the current group of RGB images are updated.
[0013] Obtain the TSDF voxel map corresponding to the indoor scene in S1. and semantic embedding vector dictionary ;
[0014] S6. Based on the TSDF voxel map Construct a topological graph of the scene's object distribution using all object instances;
[0015] S7. Obtain user commands, and based on the user commands and the topological relationship diagram of scene object distribution, output the target object corresponding to the user commands using a large language model.
[0016] S8. Determine the TSDF voxel map of the target object output in S7 obtained in S5. The center coordinates of the target object are converted into coordinates on a two-dimensional grid map, and the path between the indoor service robot and the target object is planned within the two-dimensional grid map.
[0017] Furthermore, in S2, the SEEM model is used to perform semantic segmentation on the first frame of the RGB image within each group of RGB images, obtaining the CLIP semantic embedding vector and the confidence distribution of the semantic embedding vector corresponding to each segmented region on the first frame of the RGB image within each group of RGB images. The specific process is as follows:
[0018]
[0019] in, For the confidence distribution of semantic embedding vectors, , This indicates the number of custom segmented regions. For the height of the RGB image, The width of the RGB image. For CLIP semantic embedding vectors, , This represents the dimension of the embedding vector. For the SEEM model, This is the first frame of the RGB image.
[0020] Furthermore, the TSDF voxel map in S4 for:
[0021]
[0022] in, This represents the total number of voxel blocks that include the effective object surfaces. For the first The voxel block containing the effective object surface, also known as the first voxel block. One active voxel block;
[0023]
[0024] in, For the first Individual factors, For voxel block resolution:
[0025]
[0026] in, For the first The three-channel color values of a single pixel. For the first TSDF update weights for individual elements For the first TSDF value of individual units For the first Semantic embedding vector dictionary key-value pairs of individual elements For the first Confidence level of individual elements.
[0027] Further, in step S5, starting from the second group of RGB images, the depth camera pose and depth camera intrinsics corresponding to the first frame of RGB images within each group of RGB images are obtained. Based on the first frame of RGB images, and the depth image, depth camera pose, depth camera intrinsics, and all CLIP semantic embedding vectors and semantic embedding vector confidence distributions on the first frame of RGB images, the TSDF voxel map and semantic embedding vector dictionary obtained from all previous groups of RGB images and depth images are updated. The specific process is as follows:
[0028] S51, Obtain the... Based on the depth image, depth camera pose, and depth camera intrinsics corresponding to the first RGB image in the group of RGB images, and using the... TSDF voxel map updated from group RGB images Render the depth image to obtain the rendering semantic embedding vector distribution mask and the rendering confidence distribution mask within the current depth camera's field of view;
[0029] S52, Calculate the first The semantic embedding vector confidence distribution of the first RGB image in the group of RGB images is maximized with the rendering confidence distribution mask within the field of view of the depth camera obtained by S51 to complete the pairing of the semantic embedding vector confidence distribution and the rendering confidence distribution mask.
[0030] S53, Based on the pairing results of S52, with the first... The first RGB image within the group of RGB images, along with the corresponding depth image, CLIP semantic embedding vector, and semantic embedding vector confidence distribution, are used to update the TSDF voxel map. and semantic embedding vector dictionary TSDF voxel map obtained and semantic embedding vector dictionary .
[0031] Furthermore, in S51, the first... Based on the depth image, depth camera pose, and depth camera intrinsics corresponding to the first RGB image in the group of RGB images, and using the... TSDF voxel map updated from group RGB images Render the depth image to obtain the rendering semantic embedding vector distribution mask and the rendering confidence distribution mask within the current depth camera's field of view. The specific process is as follows:
[0032] (1) Obtain the first The first RGB image within the group of RGB images Corresponding depth camera pose and depth camera internal parameters Based on the first frame RGB image Corresponding depth image Depth camera pose and depth camera internal parameters Depth image Convert to three-dimensional spatial coordinate distribution:
[0033]
[0034] in, These are the 3D coordinates of pixels after depth image transformation. and These are the poses of the depth camera. The rotation and translation matrices in the matrix. For depth camera intrinsic parameters, These are the two-dimensional coordinates of a pixel in the depth image;
[0035] (2) Determine the TSDF voxel map based on the field of view of the depth camera in (1). The range of the active voxel block corresponding to the field of view is used to determine the three-dimensional coordinates of the pixels after the depth image transformation obtained in (1). Is it located within the active voxel block area?
[0036] If located, the TSDF voxel map will be automatically activated. middle Corresponding active voxel blocks ;
[0037] If not located, continue to determine the 3D coordinates of the next pixel after the depth image transformation. Whether it is within the range of the active voxel block, up to the 3D coordinates of all pixels after depth image transformation. The judgment is complete, and the depth image is obtained. TSDF voxel map All corresponding active voxel blocks ;
[0038] (3) Based on S4, each active voxel block The semantic embedding vector dictionary key-value pairs and confidence scores of all voxels in (2) are projected onto the planar image formed by the depth camera's field of view:
[0039]
[0040] in, These are the coordinates of the projection point on the planar image. The scale value of the projection point. It is a representation of the SEEM model The intrinsic parameters of the output image resolution are used to ensure the three-dimensional coordinates of pixels after the depth image transformation. After being projected onto a planar image, the following constraints must be met:
[0041]
[0042] in, Indicates and, The width of the planar image, The height of the planar image;
[0043] Obtain all active voxel blocks in (2) The corresponding projection point on the planar image Among them, projection points belonging to the same object instance have the same semantic embedding vector dictionary key value, resulting in A dictionary of distinct semantic embedding vector keys and values. Take the integer part;
[0044] Based on all projection points and the first frame RGB image in S2 For the corresponding segmented regions, obtain the rendering semantic embedding vector distribution mask for each segmented region on the projection plane. and rendering confidence distribution mask That is, the rendering semantic embedding vector distribution mask within the field of view of the depth camera in (2) is obtained. and rendering confidence distribution mask .
[0045] Further, in S52, the calculation of the first... The semantic embedding vector confidence distribution of the first RGB image in the group of RGB images is paired with the rendering confidence distribution mask within the depth camera's field of view obtained by S51 by maximizing the flexible cross-union ratio. The specific process is as follows:
[0046] Calculate the first The first RGB image within the group of RGB images Semantic embedding vector confidence distribution The rendering confidence distribution mask within the depth camera's field of view obtained by S51 Maximize the flexible crossover ratio:
[0047]
[0048] in, The soft-IoU value is used to determine... and The pairing results The first frame of RGB image The total number of segmented regions The first frame of RGB image Upper Each segmented region For the first A semantic embedding vector dictionary key-value pair, Used for calculation and The corresponding soft-IoU values between each segmented region, matrix For all paired matching values The set of matching values satisfies the following numerical conditions:
[0049]
[0050] The formula is solved using the Jonker-Volgenant algorithm. The soft-IoU values of all segmented regions are obtained, which gives the confidence distribution of the semantic embedding vector for each segmented region. and rendering confidence distribution mask The pairing results are processed by removing all pairs with soft-IoU values less than 0.1, resulting in the remaining pairing results. This completes the first step. The first RGB image within the group of RGB images Semantic embedding vector confidence distribution The rendering confidence distribution mask within the corresponding depth camera's field of view. The pairing.
[0051] Furthermore, in S53, based on the pairing result of S52, the first... The first RGB image within the group of RGB images, along with the corresponding depth image, CLIP semantic embedding vector, and semantic embedding vector confidence distribution, are used to update the TSDF voxel map. and semantic embedding vector dictionary TSDF voxel map obtained and semantic embedding vector dictionary The specific process is as follows:
[0052] The first The first RGB image within the group of RGB images The depth image corresponding to the first frame RGB image, CLIP semantic embedding vector, and semantic embedding vector confidence distribution are mapped to the TSDF voxel map. Within, the TSDF voxel map is updated using a weighted average method. Each active voxel block The three-channel color values stored in each voxel and TSDF value :
[0053]
[0054]
[0055]
[0056] in, For the updated three-channel color values, These are the three-channel color values before the update. Update the weights for the TSDF before the update. Indicates multiplication. For the first The coordinates of the first RGB image within the group of RGB images are pixels, For the updated TSDF value, The TSDF value before the update. This indicates a truncation operation obtained from the SDF to obtain the TSDF. For the first The coordinates of the first frame depth image within the group depth image are pixels, The distance from the voxel to the camera. To truncate the threshold, Update the weights for the updated TSDF;
[0057] Meanwhile, regarding the TSDF voxel map The first segmented region in the existing data was successfully paired. The confidence level of each voxel stored within the segmented region of the first RGB image in the group of RGB images. Perform a weighted average and update the TSDF voxel map with the weighted average result. Confidence level of corresponding voxels If a new semantic embedding vector is added to the same object instance, then all semantic embedding vectors of the object instance are updated with weights, but the semantic embedding vector dictionary key-value pairs are updated. constant;
[0058] Regarding TSDF voxel maps The existing segmented regions that were not successfully paired For the segmented regions in the first frame of the RGB image group, the object instances and their corresponding attributes within the unmatched segmented regions in the first frame of the RGB image are mapped to the TSDF voxel map. Simultaneously, the CLIP semantic embedding vectors and semantic embedding vector dictionary keys of object instances within the unpaired segmentation regions of the first frame RGB image are used. Update to semantic embedding vector dictionary Within this context, the attributes include the object instance's shape, size, three-channel color values, TSDF value, CLIP semantic embedding vector, semantic embedding vector dictionary key value, and semantic embedding vector confidence.
[0059] Obtain TSDF voxel map and semantic embedding vector dictionary .
[0060] Furthermore, in S6, based on the TSDF voxel map... The process of constructing a scene object distribution topology graph from all object instances is as follows:
[0061] (1) Based on the TSDF voxel map Each object instance Semantic embedding vector dictionary key-value pairs Based on voxel spatial distribution, obtain object instances. voxels in TSDF voxel map Middle axis, shaft and The minimum and maximum positions of the axis are used to calculate the object instance. The center coordinates and the dimensions of each side of the 3D bounding box are recorded as object instances. Similarly, we can obtain the node spatial attributes of all object instances.
[0062] (2) For each object instance obtained in S2 The semantic embedding vectors are mapped to the embedding vector space of a large language model through the mapping network of the image captioning method ClipCap. The large language model decoder then converts the semantic embedding vectors into text units, and generates and outputs each object instance based on the text units. The textual expression;
[0063] (3) Construct spatial relationship edges for all object instances:
[0064] a. Set an auxiliary threshold;
[0065] b. Add the dimensions of each side of each 3D bounding box in (1) to the auxiliary threshold to obtain the extended 3D bounding box;
[0066] c. Calculate the intersection-union ratio (IUR) of the extended 3D bounding boxes between any two object instances. Determine whether the two object instances are adjacent based on the IUR of the extended 3D bounding boxes. If the IUR of the extended 3D bounding boxes is greater than 0.01, the two object instances are considered to be adjacent; otherwise, the two object instances are considered not to be adjacent.
[0067] d. For two object instances that are adjacent, determine the inclusion relationship between the two object instances based on the intersection of the three-dimensional bounding boxes of the two object instances relative to the size of each object instance. If the size of the intersection of the three-dimensional bounding boxes of the two object instances is greater than or equal to the size of object instance 1, it is determined that object instance 2 includes object instance 1; otherwise, there is no inclusion relationship.
[0068] e. For two object instances that are adjacent but not contained, calculate the vertical relationship of the intersection of the extended 3D bounding boxes of the two object instances relative to each object instance, and determine whether there is a vertical relationship between the two object instances based on the vertical relationship.
[0069] Based on ae, the TSDF voxel map was determined and obtained. The relationships between all object instances are determined, and an edge between any two object instances is constructed based on these relationships.
[0070] (4) Based on the node spatial attributes, text descriptions and edges between all object instances, construct and obtain the scene object distribution topology graph.
[0071] Furthermore, in step S7, user commands are acquired, and based on the user commands and the scene object distribution topology graph, the target object corresponding to the user commands is output using a large language model. The specific process is as follows:
[0072] Obtain user instructions;
[0073] 1) For user commands that directly extract the description of the target object, use a large language model to extract the text describing the target object. Using the SEEM model The text encoder in the code will describe the target object in text. Convert to query semantic embedding vector :
[0074]
[0075] Compute semantic embedding vector dictionary CLIP semantic embedding vectors of all object instances With query vector Cosine similarity between The calculation results are sorted top-k, and the semantic embedding vector with the highest similarity is selected and added to the semantic embedding vector dictionary. The corresponding object instance is taken as the target object, and the semantic embedding vector dictionary key-value pair of the target object is output.
[0076] 2) For user instructions whose descriptions of the target object cannot be directly extracted, the user instructions and the topological relationship diagram of scene object distribution are parsed simultaneously using the thinking chain technology of the large language model, and the semantic embedding vector dictionary key-value of the target object is output.
[0077] Furthermore, in S8, the target object output in S7 is determined in the TSDF voxel map obtained in S5. The center coordinates of the target object are converted into coordinates on a 2D grid map. The path between the indoor service robot and the target object is then planned within the 2D grid map. The specific process is as follows:
[0078] Based on the semantic embedding vector dictionary key values of the target object output by S7, the center coordinate attributes of the corresponding nodes of the target object in the scene object distribution topology map are extracted, that is, the TSDF voxel map of the target object obtained by S5. The center coordinates in Set the center coordinates of the target object Coordinates converted to a 2D raster map Based on the coordinates of the indoor service robot in the two-dimensional grid map coordinate system Calculate the vector between the center coordinates of the target object and the coordinates of the indoor service robot:
[0079]
[0080] Along the vector The algorithm searches backward in the 2D grid map, using the resolution of the 2D grid map as the step size and combining the shape parameters of the indoor service robot, until the nearest available expected stopping pose of the indoor service robot is found. Then, a graph search global path planning algorithm is used to generate a global path from the current position of the indoor service robot to the expected stopping pose. The graph search global path planning algorithm includes Dijkstra's algorithm, breadth-first search, and Floyd-Warshall algorithm.
[0081] The beneficial effects of this invention are as follows:
[0082] This invention constructs a TSDF voxel map using TSDF voxel grids based on RGB and depth images of a 3D indoor scene. During the TSDF voxel map construction process, the map is updated progressively according to traditional image frames. Simultaneously, based on TSDF rendering, the 3D semantic matching problem is transformed into a 2D candidate region matching problem. Efficient feature matching and updating are achieved by combining 2D flexible intersection-over-union (IoU) and extended Hungarian algorithms. Furthermore, the scene semantic distribution is rendered and stored using the TSDF voxel map combined with object semantic embedding vector dictionary key-value pairs. Real-time, online 3D semantic reconstruction of indoor scenes is achieved through parallel computation. This invention synchronously and in real-time updates the TSDF voxel map from both its basic and semantic attributes. Leveraging the ease of parallel processing of TSDF voxel maps, it can process sensor information streams in real-time, enabling online 3D semantic reconstruction.
[0083] This invention uses the Open Semantic Segmentation Model (SEEM), which can directly output object instance regions and corresponding CLIP semantic embedding vectors, to perform semantic segmentation on skip-frame RGB images of 3D indoor scenes. The CLIP semantic embedding vectors are used as object semantics, expanding the semantic set from a closed set to an open set. Simultaneously, leveraging the open semantic characteristics enhances adaptability to diverse objects in complex indoor environments, effectively improving the consistency of object semantic distribution extraction results. It possesses characteristics such as querying uncommon object categories and distinguishing between similar objects with different visual attributes, solving the problem that existing indoor service robot navigation methods cannot generalize to adapt to the diversity of object types and attributes in complex indoor environments.
[0084] This invention constructs a semantic embedding vector dictionary based on CLIP semantic embedding vectors and their associated key values, realizing discrete and sparse distributed storage of object semantics. The actual updating, storage, and retrieval of CLIP semantic embedding vectors are all completed through an independent semantic embedding vector dictionary. For object-descriptive user commands, this significantly improves object query efficiency compared to point cloud semantic maps.
[0085] This invention constructs a scene object distribution topology graph based on TSDF voxel maps, using formatted text descriptions. It uses the ClipCap method to directly generate explicit text descriptions from the semantic embedding vectors of objects corresponding to user commands. This transforms the three-dimensional semantic distribution information from dense TSDF voxel maps into formatted text that can be directly used for large language model inference. By using a large language model as the decision model, not only is there no need to train the decision model additionally, but it also fully utilizes prior environmental information, allowing for the identification of the target object of the user command in one step, thus avoiding inefficient exploration behavior.
[0086] Finally, an object-oriented global path planner is proposed, which finds the nearest available pose before applying graph search for global path planning, thus solving the problem that conventional global path planning algorithms cannot directly use object coordinates as navigation targets. Attached Figure Description
[0087] Figure 1 This is a schematic diagram of the basic idea of the CLIP model;
[0088] Figure 2 This is a schematic diagram of the CLIP-based zero-shot classification method;
[0089] Figure 3 This is a schematic diagram illustrating the principle of reconstructing an object's surface using TSDF;
[0090] Figure 4 This is a diagram illustrating the camera's field of view;
[0091] Figure 5 This is a diagram illustrating the ClipCap image captioning method;
[0092] Figure 6 This is a flowchart of S2 and S5;
[0093] Figure 7 This is a flowchart of S7 to S9;
[0094] Figure 8 This is an example of prompt words for the large-scale navigation goal planning reasoning model;
[0095] Figure 9 Pose search is available for indoor service robots; Detailed Implementation
[0096] Specific implementation method one: Combining Figures 1-9 This embodiment describes an indoor service robot navigation planning method based on open semantic mapping and a large language model, which includes the following steps:
[0097] S1. Given a 3D indoor scene, command an indoor service robot equipped with a depth camera to film the indoor scene, obtaining a video containing the structure of the indoor scene, and obtain multiple frames of RGB images from the video. and multi-frame depth images ,in, For the height of the image, The width of the image. Each frame of the RGB image and each frame of the depth image cover the same scene; both the RGB image and the depth image are two-dimensional images. The number of frames is greater than two.
[0098] Ten consecutive RGB images are grouped together to obtain multiple sets of RGB images. Ten consecutive depth images are grouped together to obtain multiple sets of depth images. The number of sets is greater than or equal to two.
[0099] S2. Using the SEEM model Perform semantic segmentation on the first RGB image within each group of RGB images:
[0100]
[0101] Obtain the CLIP semantic embedding vector corresponding to each segmented region in the first frame of the RGB image within each group of RGB images. and semantic embedding vector confidence distribution ,in, This indicates a custom SEEM object query quantity hint, i.e., the number of segmented regions. This represents the dimension of the embedding vector. The CLIP model deployed by the SEEM model of this invention is a clip-vit-base-patch32 model. .
[0102] SEEM (Segment Everything Everywhere with Multi-modal prompts allat once) is an open set instance segmentation method based on CLIP. Compared to CLIP, which only yields a general semantic embedding vector after encoding an image, SEEM can adaptively segment object instances in a single image and extract the corresponding CLIP semantic features. CLIP (Contrastive Language-Image Pre-Training) is a multimodal pre-trained neural network, whose basic idea is as follows: Figure 1As shown, text-image data pairs with the same semantics are encoded into high-dimensional semantic embedding vectors of uniform dimension using a text encoder and an image encoder. A cross-entropy loss function is constructed based on vector cosine similarity, and the model is trained to learn the alignment relationship between images and text. The CLIP model, pre-trained on 400 million pairs of internet text-image pairings, can be used to construct zero-shot image classifiers, such as... Figure 2 As shown, a series of possible object categories are defined according to a specific task. These categories are then converted into CLIP semantic embedding vectors using a CLIP text encoder. For the image to be classified, its CLIP semantic embedding vector is extracted using an image encoder. The cosine similarity between the category text feature vector and the image feature vector is calculated, and the highest result is taken as the prediction result. Because the SEEM model of this invention uses internet-scale training data in the pre-training stage, feature extraction and prediction can cover various everyday scenarios. The category query method based on vector cosine similarity calculation also gives it the characteristics of an open vocabulary set.
[0103] S3. Regarding the storage of semantic information, compared to the traditional method of storing individual semantic information for each point, in order to improve the semantic consistency of the same object and reduce storage and retrieval complexity, the advantages of S2 are fully utilized. Based on the content of the first frame RGB image and all CLIP semantic embedding vectors on the first frame RGB image, For each object in the first frame of the RGB image, a semantic embedding vector dictionary key is associated with its CLIP semantic embedding vector. A semantic embedding vector dictionary is constructed based on all CLIP semantic embedding vectors and all semantic embedding vector dictionary key values in the first frame of the RGB image. The actual CLIP semantic embedding vector updates and storage are both handled through a separate semantic embedding vector dictionary. Complete. The characteristic of TSDF that only reconstructs the object surface without storing the complete object distribution also reduces the storage requirements of semantic embedding vectors, which is beneficial for more efficient subsequent queries.
[0104] TSDF (Truncated Signed Distance Function) is a 3D surface reconstruction algorithm based on SDF (Signed Distance Function). Compared to the completely dense nature of traditional point cloud maps, TSDF voxel raster maps are spatially sparse but have dense object surfaces. While ensuring the quality of object surface reconstruction, the algorithm highly supports parallel computing deployment and can be significantly accelerated using GPUs. The physical meaning of SDF is to represent the shortest distance from every point in space to any plane; the positive or negative sign of the sign represents the relative position of the point and the plane. For a point in 3D space... The general form of expression is:
[0105]
[0106] in, Represents the distance function. Representing spatial entities, This represents the boundary of a solid surface. For points outside the solid, the distance is positive; for points inside the solid, the distance is negative. TSDF, based on SDF, truncates the distance values to a specific threshold range. Generally speaking, The value is taken as four times the voxel size. For example... Figure 3 As shown, when the normalized distance function is taken... At that time, the surface of the reconstructed object can be determined based on the positive and negative boundaries of the TSDF values stored in the voxels in space.
[0107] S4. This invention uses TSDF voxel grids as the representation of indoor scene maps. Based on the first set of RGB images and the first set of depth images, a TSDF voxel map is constructed using TSDF voxel grids. :
[0108]
[0109] in, This represents the total number of voxel blocks that include the surfaces of valid objects. For the first The voxel block containing the effective object surface, also known as the first voxel block. A set of active voxel blocks. For voxel blocks far from the object surface, the TSDF value is the initial value or null. For voxel blocks containing the object surface, the voxel block contains... A valid voxel value, This is the voxel block resolution, used to adjust the number of voxels in a single voxel block, and is generally set to... Each effective voxel block is internally dense, and all voxels contained within it have valid TSDF values. The information stored within an effective voxel block can be represented as the set of information stored in all voxels within it.
[0110]
[0111] in, For the first Individual factors, The information stored in the internal storage is:
[0112]
[0113] in, For the first The three-channel color values of a single pixel. For the first TSDF update weights for individual elements For the first TSDF value of individual units For the first Semantic embedding vector dictionary key-value pairs of individual elements For the first Confidence level of individual elements.
[0114] S5. Based on S4, update the TSDF voxel map sequentially with each subsequent RGB image and depth image from the first set of RGB images and the first set of depth images. and semantic embedding vector dictionary .
[0115] Simultaneously, starting from the second group of RGB images, the depth camera pose corresponding to the first frame of the RGB image within each group of RGB images is obtained. and depth camera internal parameters Based on the first frame RGB image, and the corresponding depth image, depth camera pose, depth camera intrinsics, and all CLIP semantic embedding vectors and semantic embedding vector confidence distributions on the first frame RGB image, the TSDF voxel map and semantic embedding vector dictionary obtained from all previous groups of RGB images and depth images are updated. The specific process is as follows:
[0116] S51, Obtain the... The first RGB image within the group of RGB images Corresponding depth camera pose and depth camera internal parameters Based on the first frame RGB image Corresponding depth image Depth camera pose and depth camera internal parameters , using the TSDF voxel map updated from group RGB images Rendering depth images This yields the rendering semantic embedding vector distribution mask within the current depth camera's field of view. and rendering confidence distribution mask The specific process is as follows:
[0117] (1) Obtain the first The first RGB image within the group of RGB images Corresponding depth camera pose and depth camera internal parameters Based on the first frame RGB image Corresponding depth image Depth camera pose and depth camera internal parameters Depth image Convert to three-dimensional spatial coordinate distribution:
[0118]
[0119] in, These are the two-dimensional coordinates of pixels in the depth image. These are the 3D coordinates of pixels after depth image transformation. and These are the poses of the depth camera. The rotation and translation matrices in the matrix. This is the intrinsic parameter of the depth camera. and They all exist in matrix form.
[0120] (2) Based on the field of view of the depth camera in (1), the TSDF voxel map is preliminarily determined. The active voxel block range corresponding to the field of view is defined as all voxel blocks within the field of view of the depth camera. Each voxel block is composed of multiple local voxels. The three-dimensional coordinates of the pixels obtained after the depth image transformation are determined by (1). Is it located within the active voxel block area? If so, the TSDF voxel map will be automatically activated. middle Corresponding active voxel blocks If not located, continue to determine the 3D coordinates of the next depth image-transformed pixel. Whether it is within the range of the active voxel block, up to the 3D coordinates of all pixels after depth image transformation. The judgment is complete, and the depth image is obtained. TSDF voxel map All corresponding active voxel blocks Active voxel blocks The number is greater than or equal to 2.
[0121] like Figure 4 As shown, the camera's field of view is all three-dimensional spatial regions that can be imaged onto the image plane directly in front of the camera, forming a three-dimensional cone. Theoretically, the three-dimensional cone extends infinitely forward, but in practical applications, it is limited by the camera's resolution. Objects that are too far away cannot be reliably imaged. Therefore, a depth image can only render a series of discrete point clouds contained within this three-dimensional cone (each pixel corresponds to one).
[0122] (3) Based on S4, each active voxel block The semantic embedding vector dictionary key-value pairs and confidence scores of all voxels in (2) are projected onto the planar image formed by the depth camera's field of view:
[0123]
[0124] in, These are the coordinates of the projection point on the planar image. The scale value of the projection point. It is a representation of the SEEM model The intrinsic parameters of the output image resolution are used to ensure the three-dimensional coordinates of pixels after the depth image transformation. After being projected onto a planar image, the following constraints must be met:
[0125]
[0126] in, Indicates and, The width of the planar image, The height of the planar image;
[0127] Obtain all active voxel blocks in (2) The corresponding projection point on the planar image Among them, projection points belonging to the same object instance have the same semantic embedding vector dictionary key-value attributes, resulting in A dictionary of distinct semantic embedding vector keys and values. Take the integer part.
[0128] Based on all projection points and the first frame RGB image in S2 For the corresponding segmented regions, obtain the rendering semantic embedding vector distribution mask for each segmented region on the projection plane. and rendering confidence distribution mask That is, the rendering semantic embedding vector distribution mask within the field of view of the depth camera in (2) is obtained. and rendering confidence distribution mask Rendering confidence distribution mask It represents the distribution of semantic information of all object instances accumulated within the current depth camera's field of view.
[0129] This step will involve each active voxel block in the 3D map. The semantic embedding vector dictionary key and confidence score are projected onto a two-dimensional plane. That is, the semantic embedding vector dictionary key and confidence score of the two-dimensional projection point are taken from the corresponding three-dimensional space point. For example, a red point in three-dimensional space is projected onto a two-dimensional plane and the resulting point is also red. This invention still retains the point's attributes after projection.
[0130] S52, Calculate the first The first RGB image within the group of RGB images Semantic embedding vector confidence distribution The rendering confidence distribution mask within the depth camera's field of view obtained by S51 Maximize the flexible crossover ratio to complete the confidence distribution of semantic embedding vectors. With rendering confidence distribution mask The pairing process is as follows:
[0131] To find the first RGB image This invention proposes a fusion candidate method for the CLIP semantic embedding vector of an object instance and its position in the TSDF voxel space. It utilizes temporal feature matching of segmented regions obtained based on S2 to complete the pairing of CLIP semantic embedding vectors with the distribution mask of rendered semantic embedding vectors, and semantic embedding vector confidence distribution with the mask of rendered confidence distribution. This invention models the pairing problem as a two-dimensional rectangle assignment problem, with the optimization objective being to maximize the confidence distribution of semantic embedding vectors in each segmented region of the RGB image. and rendering confidence distribution mask soft-IoU (soft Intersection over Union):
[0132]
[0133] in, The soft-IoU value is used to determine... and The pairing results The first frame of RGB image The total number of segmented regions The first frame of RGB image Upper Each segmented region For the first A semantic embedding vector dictionary key-value pair, function Used for calculation and The corresponding soft-IoU values between each segmented region, matrix For all paired matching values The set of matching values satisfies the following numerical conditions:
[0134]
[0135] The above formula is solved using the Jonker-Volgenant algorithm (an extended version of the Hungarian algorithm). The soft-IoU values of all segmented regions are obtained, which gives the confidence distribution of the semantic embedding vector for each segmented region. and rendering confidence distribution mask The pairing results. To avoid merging low-quality masks caused by occlusion or blurring, all soft-IoU values are discarded. For pairings less than 0.1, obtain the remaining pairings and complete the process. The first RGB image within the group of RGB images Semantic embedding vector confidence distribution The rendering confidence distribution mask within the corresponding depth camera's field of view. The pairing.
[0136] S53, Based on the pairing results of S52, with the first... The first RGB image within the group of RGB images The TSDF voxel map is updated by updating the depth image corresponding to the first frame RGB image, CLIP semantic embedding vector, and semantic embedding vector confidence distribution. and semantic embedding vector dictionary TSDF voxel map obtained and semantic embedding vector dictionary The specific process is as follows:
[0137] Regarding the basic properties of TSDF, the first The first RGB image within the group of RGB images The depth image corresponding to the first frame RGB image, CLIP semantic embedding vector, and semantic embedding vector confidence distribution are mapped to the TSDF voxel map. Internally, following the basic TSDF integration method, the TSDF voxel map is updated using a weighted average approach. Each active voxel block The three-channel color values stored in each voxel and TSDF value :
[0138]
[0139]
[0140]
[0141] in, For the updated three-channel color values, These are the three-channel color values before the update. Update the weights for the TSDF before the update. Indicates multiplication. For the first The coordinates of the first RGB image within the group of RGB images are pixels, For the updated TSDF value, The TSDF value before the update. This indicates a truncation operation obtained from the SDF to obtain the TSDF. For the first The coordinates of the first frame depth image within the group depth image are pixels, The distance from the voxel to the camera is determined by the camera pose and the voxel's coordinates in the voxel grid map. To truncate the threshold, Update the weights for the updated TSDF.
[0142] Meanwhile, updating the TSDF voxel map At that time, for the TSDF semantic attributes (the dictionary keys of semantic embedding vectors in voxels and the CLIP semantic embedding vectors in the dictionary), for the TSDF voxel map The first semantic segmentation region that has been successfully paired in the existing semantic segmentation region The confidence level stored for each voxel within the segmented region of the first RGB image in the group of RGB images. Perform a weighted average and update the TSDF voxel map with the weighted average result. Confidence level of corresponding voxels If a new semantic embedding vector is added to the same object instance, then all semantic embedding vectors of the object instance are updated with weights, but the semantic embedding vector dictionary key-value pairs are updated. constant;
[0143] Regarding TSDF voxel maps The semantic segmentation regions that have not been successfully paired in the existing semantic segmentation region For the segmented regions in the first frame of the RGB image group, the object instances and their corresponding attributes within the unmatched segmented regions in the first frame of the RGB image are mapped to the TSDF voxel map. Simultaneously, the CLIP semantic embedding vectors and semantic embedding vector dictionary keys of object instances within the unpaired segmentation regions of the first frame RGB image are used. Update to semantic embedding vector dictionary Within this scope, it is registered as a new object instance. The attributes include the object instance's shape, size, CLIP semantic embedding vector, semantic embedding vector dictionary key-value pairs, and semantic embedding vector confidence, etc.
[0144] Obtain TSDF voxel map and semantic embedding vector dictionary A TSDF voxel map can be represented as a globally sparse set of voxel blocks.
[0145] Based on S1-S4, repeat S51-S53 until all RGB images and depth images have been processed, and obtain the TSDF voxel map corresponding to the indoor scene in S1. and semantic embedding vector dictionary .
[0146] S6. Based on the TSDF voxel map Construct a topological graph of the scene's object distribution based on all object instances.
[0147] A 3D scene graph is an efficient and concise way to abstract scene representation. Each node in the graph structure typically represents an object instance, and the parameters of the nodes record various attributes of the object. The edges in the graph structure represent the relative spatial relationships between objects, usually categorized into five types: objects A and B are adjacent, object A is above B, object A is below B, object A is inside B, and object A is contained within B. Compared to completely dense 3D scene reconstruction and sparse discrete object dictionaries, scene graphs retain as much object attribute and distribution information as possible while simplifying scene representation. At the cost of reduced query and inference real-time performance, they can effectively improve adaptability to complex user commands. Furthermore, their ease of use with formatted text representations allows them to be well integrated with large language models, further improving the human-like intelligence and interactivity of navigation target planning. This invention constructs an instance-level scene object distribution topology graph using the following three steps:
[0148] (1) Extraction of spatial attributes of nodes (object instances)
[0149] For the extraction of spatial attributes of nodes (object instances), this invention uses semantic embedding vector dictionary key-value pairs. For indexing, query the distribution of all object instances within the new TSDF voxel map. Based on each object instance... Semantic embedding vector dictionary key-value pairs Based on voxel spatial distribution, obtain object instances. voxels in TSDF voxel map Middle axis, shaft and Minimum and maximum positions of the axes. This is because object instances are located on the TSDF voxel map. The bounding box is represented as a series of adjacent voxels, each with its own x, y, and z coordinates. Taking the six largest x, y, and z values, and the smallest x, y, and z values, we form two sets of coordinates. These correspond to the two opposite vertices of the 3D cube bounding box, where each side is parallel to each coordinate axis. A 3D bounding box is used to enclose an instance of an object in 3D space. A cuboid object. Calculate object instances based on minimum and maximum positions. The center coordinates and the dimensions of each side of the 3D bounding box are recorded as object instances. Similarly, we can obtain the node space attributes of all object instances.
[0150] (2) Explicit textualization of object semantics
[0151] To endow object nodes with explicit textual semantic representations for subsequent parsing by large language models, this invention introduces the image captioning method ClipCap. ClipCap mainly includes a CLIP image encoder, a mapping network, and a large language model (such as GPT2), aiming to generate high-quality textual descriptions for images. Its basic idea is as follows: Figure 5 As shown. Each object instance obtained in S2 The semantic embedding vectors are mapped to the embedding vector space of a large language model through the mapping network of the image captioning method ClipCap. The large language model decoder then converts the semantic embedding vectors into text units, and generates and outputs each object instance based on the text units. The textual expression.
[0152] (3) Construct the spatial relationship edges of all object instances
[0153] The construction of edges in the topological relationship graph of object distribution in the scene according to the present invention includes the following steps:
[0154] a. Set an auxiliary threshold, which is set manually based on experience.
[0155] b. Add the dimensions of each side of each three-dimensional bounding box in (1) to the auxiliary threshold to obtain the extended three-dimensional bounding box.
[0156] c. Calculate the intersection-union ratio (IUR) of the extended 3D bounding boxes between any two object instances. Determine whether the two object instances are adjacent based on the IUR of the extended 3D bounding boxes. If the IUR of the extended 3D bounding boxes is greater than 0.01, the two object instances are considered to be adjacent; otherwise, the two object instances are considered not to be adjacent.
[0157] d. For two object instances that are adjacent, determine the inclusion relationship between the two object instances based on the intersection of their 3D bounding boxes relative to the size of each object instance. If the size of the intersection of the 3D bounding boxes of the two object instances is greater than or equal to the size of object instance 1, then object instance 2 is considered to contain object instance 1; otherwise, no inclusion relationship exists. For example, if object instance 1 and object instance 2 are adjacent, and the intersection of object instance 1 and object instance 2 is greater than or equal to the size of object instance 2, then object instance 1 is considered to contain object instance 2.
[0158] e. For two object instances that are adjacent but not contained, calculate the top-bottom relationship of the intersection of the extended 3D bounding boxes of the two object instances relative to each object instance, and determine whether there is a clear top-bottom relationship between the two object instances based on the top-bottom relationship.
[0159] Based on the above ae, the TSDF voxel map was determined and obtained. The relationships between all object instances are determined, and an edge between any two object instances is constructed based on these relationships.
[0160] (4) Based on the node spatial attributes, text descriptions and edges between all object instances, construct and obtain the scene object distribution topology graph.
[0161] S7. Obtain user commands. Based on the user commands and the topological relationship diagram of scene object distribution, use a large language model to output the target object corresponding to the user commands.
[0162] The large language model is introduced via an online API (Application Programming Interface) call. The prompt words for the large language model are written as follows: Figure 8 As shown, a large language model is used to comprehensively analyze the topological relationship of scene objects and complex user commands to obtain the target object that best matches the user's expectations, and return the corresponding object node of the target object. The prompt words are repeatedly debugged and optimized to improve the large language model's adherence to instruction specifications and the rationality of the inference results.
[0163] 1) For user commands from which the description of the target object can be directly extracted, leverage the semantic analysis and tool usage capabilities of the large language model to extract concise and clear text describing the target object. Using the SEEM model The text encoder in the code will describe the target object in text. Convert to query semantic embedding vector :
[0164]
[0165] Compute semantic embedding vector dictionary CLIP semantic embedding vectors of all object instances With query vector Cosine similarity between The calculation results are sorted top-k, and the semantic embedding vector with the highest similarity is selected and added to the semantic embedding vector dictionary. The corresponding object instance is used as the target object, and the semantic embedding vector dictionary key-value pair of the target object is output. This enables fast and simple command response.
[0166] 2) For user instructions whose target object description cannot be directly extracted (abstract and cryptic user needs), the thinking chain technology of the large language model is used to perform multi-step reasoning, simultaneously parsing the topological relationship between user instructions and scene object distribution, understanding the correspondence between user instructions and available objects in the scene, generating human-like suggestions about the target object corresponding to the user instruction that conform to common sense cognition, and outputting the semantic embedding vector dictionary key-value of the target object.
[0167] LLM (Language Learning Model) is a large neural network model built using learning techniques. It learns language patterns, structures, and semantics through unsupervised learning on massive amounts of text data. Its architecture typically employs advanced architectures such as the Transformer, enabling it to process and generate natural language text. During training, the model predicts the next word or phrase based on context, continuously optimizing its parameters. LLM possesses powerful language understanding and generation capabilities, applicable to various fields such as text generation, machine translation, and question-answering systems, representing a significant breakthrough in natural language processing. Leading large language model agents, such as DeepSeek-R1 and ChatGPT, both domestically and internationally, already possess the ability to perform logical and semantic reasoning based on natural language descriptions. They can be used to parse user natural language instructions and text formatting context information to generate navigation target candidates.
[0168] S8. Determine the TSDF voxel map of the target object output in S7 obtained in S5. The center coordinates of the target object are converted into coordinates on a 2D grid map. The path between the indoor service robot and the target object is then planned within the 2D grid map. The specific process is as follows:
[0169] like Figure 9 As shown, based on the semantic embedding vector dictionary key values of the target object output by S7, the center coordinate attributes of the nodes corresponding to the target object in the scene object distribution topology map are extracted, which is the TSDF voxel map of the target object obtained in S5. The center coordinates in Set the center coordinates of the target object Coordinates converted to a 2D raster map Based on the coordinates of the indoor service robot in the two-dimensional grid map coordinate system Calculate the vector between the center coordinates of the target object and the coordinates of the indoor service robot:
[0170]
[0171] Along the vector The algorithm searches backward in the 2D grid map, using the resolution of the 2D grid map as the step size and incorporating the shape parameters of the indoor service robot, until the nearest available expected stopping pose of the indoor service robot is found. A global path planning algorithm based on graph search is then used to generate a global path from the current position of the indoor service robot to the expected stopping pose. Graph search global path planning algorithms include Dijkstra's algorithm, Breadth-First Search (BFS), and the Floyd-Warshall algorithm.
[0172] Since this invention does not specify the source of the depth camera pose input, although there is always a physical correspondence and transformation relationship between 2D and 3D map coordinates, the specific calibration form depends on the depth camera pose acquisition method. Therefore, their origins may not coincide, and the z-axis coordinate of the object in the 2D map may not be 0. The resolution of the 2D grid map is the size of each grid cell (described by one-dimensional length). In this invention, using it as a step size means that when searching along the vector method, a position is taken at every step size distance to determine the availability of the position until a suitable position is found. Since different indoor service robot chassis have different forms, indoor service robots can be abstracted into various shape representations, namely the indoor service robot shape parameters mentioned above. During the movement of the indoor service robot, the search for the nearest available pose and the global path replanning are periodically performed to continuously optimize the relative distance between the indoor service robot and the target object until the indoor service robot stops at the position closest to the target object.
[0173] The above examples of the present invention are merely illustrative of the computational model and process of the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is impossible to exhaustively list all possible implementations here. Any obvious variations or modifications derived from the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A navigation planning method for indoor service robots based on open semantic mapping and large language models, characterized by: It includes the following steps: S1. Given a three-dimensional indoor scene, command an indoor service robot equipped with a depth camera to take pictures of the indoor scene to obtain a video containing the structure of the indoor scene. Based on the video, obtain multiple frames of RGB images and multiple frames of depth images, with each frame of RGB image and each frame of depth image covering the same scene. Ten consecutive RGB images are grouped together to obtain multiple sets of RGB images, and ten consecutive depth images are grouped together to obtain multiple sets of depth images. S2. Use the SEEM model to perform semantic segmentation on the first frame of RGB images in each group of RGB images, and obtain the CLIP semantic embedding vector and the confidence distribution of semantic embedding vector corresponding to each segmented region in the first frame of RGB images in each group of RGB images. S3. Based on the content of the first frame RGB image and all CLIP semantic embedding vectors in the first frame RGB image, associate a semantic embedding vector dictionary key-value pair with the CLIP semantic embedding vector of each object in the first frame RGB image. Construct a semantic embedding vector dictionary based on all CLIP semantic embedding vectors and all semantic embedding vector dictionary key-value pairs in the first frame RGB image. ; S4. Based on the first set of RGB images and the first set of depth images, construct a TSDF voxel map using TSDF voxel grids. ; S5. Based on S4, update the TSDF voxel map sequentially with each subsequent RGB image and depth image from the first set of RGB images and the first set of depth images. and semantic embedding vector dictionary ; Simultaneously, starting from the second group of RGB images, the depth camera pose and depth camera intrinsics corresponding to the first frame of RGB images in each group of RGB images are obtained. Based on the first frame of RGB images, as well as the depth image, depth camera pose, depth camera intrinsics, and all CLIP semantic embedding vectors and semantic embedding vector confidence distributions on the first frame of RGB images, the TSDF voxel map and semantic embedding vector dictionary obtained from all groups of RGB images and depth images before the current group of RGB images are updated. Obtain the TSDF voxel map corresponding to the indoor scene in S1. and semantic embedding vector dictionary ; S6. Based on the TSDF voxel map Construct a topological graph of the scene's object distribution using all object instances; S7. Obtain user commands, and based on the user commands and the topological relationship diagram of scene object distribution, output the target object corresponding to the user commands using a large language model. S8. Determine the TSDF voxel map of the target object output in S7 obtained in S5. The center coordinates of the target object are converted into coordinates on a two-dimensional grid map, and the path between the indoor service robot and the target object is planned within the two-dimensional grid map.
2. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 1, characterized in that: In step S2, the SEEM model is used to perform semantic segmentation on the first frame of the RGB image within each group of RGB images, obtaining the CLIP semantic embedding vector and the confidence distribution of the semantic embedding vector corresponding to each segmented region on the first frame of the RGB image within each group of RGB images. The specific process is as follows: in, For the confidence distribution of semantic embedding vectors, , This indicates the number of custom segmented regions. For the height of the RGB image, The width of the RGB image. For CLIP semantic embedding vectors, , This represents the dimension of the embedding vector. For the SEEM model, This is the first frame of the RGB image.
3. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 2, characterized in that: TSDF voxel map in S4 for: in, This represents the total number of voxel blocks that include the surfaces of valid objects. For the first The voxel block containing the effective object surface, also known as the first voxel block. One active voxel block; in, For the first Individual factors, For voxel block resolution: in, For the first The three-channel color values of a single pixel. For the first TSDF update weights for individual elements For the first TSDF value of individual units For the first Semantic embedding vector dictionary key-value pairs of individual elements For the first Confidence level of individual elements.
4. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 3, characterized in that: In step S5, starting from the second group of RGB images, the depth camera pose and depth camera intrinsics corresponding to the first frame of RGB images within each group of RGB images are obtained. Based on the first frame of RGB images, and the depth image, depth camera pose, depth camera intrinsics, and all CLIP semantic embedding vectors and semantic embedding vector confidence distributions on the first frame of RGB images, the TSDF voxel map and semantic embedding vector dictionary obtained from all previous groups of RGB images and depth images are updated. The specific process is as follows: S51, Obtain the... Based on the depth image, depth camera pose, and depth camera intrinsics corresponding to the first RGB image in the group of RGB images, and using the... TSDF voxel map updated from group RGB images Render the depth image to obtain the rendering semantic embedding vector distribution mask and the rendering confidence distribution mask within the current depth camera's field of view; S52, Calculate the first The semantic embedding vector confidence distribution of the first RGB image in the group of RGB images is maximized with the rendering confidence distribution mask within the field of view of the depth camera obtained by S51 to complete the pairing of the semantic embedding vector confidence distribution and the rendering confidence distribution mask. S53, Based on the pairing results of S52, with the first... The first RGB image within the group of RGB images, along with the corresponding depth image, CLIP semantic embedding vector, and semantic embedding vector confidence distribution, are used to update the TSDF voxel map. and semantic embedding vector dictionary TSDF voxel map obtained and semantic embedding vector dictionary .
5. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 4, characterized in that: The S51 process obtains the first Based on the depth image, depth camera pose, and depth camera intrinsics corresponding to the first RGB image in the group of RGB images, and using the... TSDF voxel map updated from group RGB images Render the depth image to obtain the rendering semantic embedding vector distribution mask and the rendering confidence distribution mask within the current depth camera's field of view. The specific process is as follows: (1) Obtain the first The first RGB image within the group of RGB images Corresponding depth camera pose and depth camera internal parameters Based on the first frame RGB image Corresponding depth image Depth camera pose and depth camera internal parameters Depth image Convert to three-dimensional spatial coordinate distribution: .
6. Among them, These are the 3D coordinates of pixels after depth image transformation. and These are the poses of the depth camera. The rotation and translation matrices in the matrix. For depth camera intrinsic parameters, These are the two-dimensional coordinates of a pixel in the depth image; (2) Determine the TSDF voxel map based on the field of view of the depth camera in (1). The range of the active voxel block corresponding to the field of view is used to determine the three-dimensional coordinates of the pixels after the depth image transformation obtained in (1). Is it located within the active voxel block area? If located, the TSDF voxel map will be automatically activated. middle Corresponding active voxel blocks ; If not located, continue to determine the 3D coordinates of the next pixel after the depth image transformation. Whether it is within the range of the active voxel block, up to the 3D coordinates of all pixels after depth image transformation. The judgment is complete, and the depth image is obtained. TSDF voxel map All corresponding active voxel blocks ; (3) Based on S4, each active voxel block The semantic embedding vector dictionary key-value pairs and confidence scores of all voxels in (2) are projected onto the planar image formed by the depth camera's field of view: in, These are the coordinates of the projection point on the planar image. The scale value of the projection point. It is a representation of the SEEM model The intrinsic parameters of the output image resolution are used to ensure the three-dimensional coordinates of pixels after the depth image transformation. After being projected onto a planar image, the following constraints must be met: in, Indicates and, The width of the planar image, The height of the planar image; Obtain all active voxel blocks in (2) The corresponding projection point on the planar image Among them, projection points belonging to the same object instance have the same semantic embedding vector dictionary key value, resulting in A dictionary of distinct semantic embedding vector keys and values. Take the integer part; Based on all projection points and the first frame RGB image in S2 For the corresponding segmented regions, obtain the rendering semantic embedding vector distribution mask for each segmented region on the projection plane. and rendering confidence distribution mask That is, the rendering semantic embedding vector distribution mask within the field of view of the depth camera in (2) is obtained. and rendering confidence distribution mask .
7. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 5, characterized in that: The calculation of the first in S52 The semantic embedding vector confidence distribution of the first RGB image in the group of RGB images is paired with the rendering confidence distribution mask within the depth camera's field of view obtained by S51 by maximizing the flexible cross-union ratio. The specific process is as follows: Calculate the first The first RGB image within the group of RGB images Semantic embedding vector confidence distribution The rendering confidence distribution mask within the depth camera's field of view obtained by S51 Maximize the flexible crossover ratio: in, The soft-IoU value is used to determine... and The pairing results The first frame of RGB image The total number of segmented regions The first frame of RGB image Upper Each segmented region For the first A semantic embedding vector dictionary key-value pair, Used for calculation and The corresponding soft-IoU values between each segmented region, matrix For all paired matching values The set of matching values satisfies the following numerical conditions: The formula is solved using the Jonker-Volgenant algorithm. The soft-IoU values of all segmented regions are obtained, which gives the confidence distribution of the semantic embedding vector for each segmented region. and rendering confidence distribution mask The pairing results are processed by removing all pairs with soft-IoU values less than 0.1, resulting in the remaining pairing results. This completes the first step. The first RGB image within the group of RGB images Semantic embedding vector confidence distribution The rendering confidence distribution mask within the corresponding depth camera's field of view. The pairing.
8. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 6, characterized in that: The pairing result in S53 based on S52 is used as the first... The first RGB image within the group of RGB images, along with the corresponding depth image, CLIP semantic embedding vector, and semantic embedding vector confidence distribution, are used to update the TSDF voxel map. and semantic embedding vector dictionary TSDF voxel map obtained and semantic embedding vector dictionary The specific process is as follows: The first The first RGB image within the group of RGB images The depth image corresponding to the first frame RGB image, CLIP semantic embedding vector, and semantic embedding vector confidence distribution are mapped to the TSDF voxel map. Within, the TSDF voxel map is updated using a weighted average method. Each active voxel block The three-channel color values stored in each voxel and TSDF value : in, For the updated three-channel color values, These are the three-channel color values before the update. Update the weights for the TSDF before the update. Indicates multiplication. For the first The coordinates of the first RGB image within the group of RGB images are pixels, For the updated TSDF value, The TSDF value before the update. This indicates a truncation operation obtained from the SDF to obtain the TSDF. For the first The coordinates of the first frame depth image within the group depth image are pixels, The distance from the voxel to the camera. To truncate the threshold, Update the weights for the updated TSDF; Meanwhile, regarding the TSDF voxel map The first segmented region in the existing data was successfully paired. The confidence level of each voxel stored within the segmented region of the first RGB image in the group of RGB images. Perform a weighted average and update the TSDF voxel map with the weighted average result. Confidence level of corresponding voxels If a new semantic embedding vector is added to the same object instance, then all semantic embedding vectors of the object instance are updated with weights, but the semantic embedding vector dictionary key-value pairs are updated. constant; Regarding TSDF voxel maps The existing segmented regions that were not successfully paired For the segmented regions in the first frame of the RGB image group, the object instances and their corresponding attributes within the unmatched segmented regions in the first frame of the RGB image are mapped to the TSDF voxel map. Simultaneously, the CLIP semantic embedding vectors and semantic embedding vector dictionary keys of object instances within the unpaired segmentation regions of the first frame RGB image are used. Update to semantic embedding vector dictionary Within this context, the attributes include the object instance's shape, size, three-channel color values, TSDF value, CLIP semantic embedding vector, semantic embedding vector dictionary key value, and semantic embedding vector confidence. Obtain TSDF voxel map and semantic embedding vector dictionary .
9. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 7, characterized in that: In S6, based on the TSDF voxel map The process of constructing a scene object distribution topology graph from all object instances is as follows: (1) Based on the TSDF voxel map Each object instance Semantic embedding vector dictionary key-value pairs Based on voxel spatial distribution, obtain object instances. voxels in TSDF voxel map Middle axis, shaft and The minimum and maximum positions of the axis are used to calculate the object instance. The center coordinates and the dimensions of each side of the 3D bounding box are recorded as object instances. Similarly, we can obtain the node spatial attributes of all object instances. (2) For each object instance obtained in S2 The semantic embedding vectors are mapped to the embedding vector space of a large language model through the mapping network of the image captioning method ClipCap. The large language model decoder then converts the semantic embedding vectors into text units, and generates and outputs each object instance based on the text units. The textual expression; (3) Construct spatial relationship edges for all object instances: a. Set an auxiliary threshold; b. Add the dimensions of each side of each 3D bounding box in (1) to the auxiliary threshold to obtain the extended 3D bounding box; c. Calculate the intersection-union ratio (IUR) of the extended 3D bounding boxes between any two object instances. Determine whether the two object instances are adjacent based on the IUR of the extended 3D bounding boxes. If the IUR of the extended 3D bounding boxes is greater than 0.01, the two object instances are considered to be adjacent; otherwise, the two object instances are considered not to be adjacent. d. For two object instances that are adjacent, determine the inclusion relationship between the two object instances based on the intersection of the three-dimensional bounding boxes of the two object instances relative to the size of each object instance. If the size of the intersection of the three-dimensional bounding boxes of the two object instances is greater than or equal to the size of object instance 1, it is determined that object instance 2 includes object instance 1; otherwise, there is no inclusion relationship. e. For two object instances that are adjacent but not contained, calculate the vertical relationship of the intersection of the extended 3D bounding boxes of the two object instances relative to each object instance, and determine whether there is a vertical relationship between the two object instances based on the vertical relationship. Based on ae, the TSDF voxel map was determined and obtained. The relationships between all object instances are determined, and an edge between any two object instances is constructed based on these relationships. (4) Based on the node spatial attributes, text descriptions and edges between all object instances, construct and obtain the scene object distribution topology graph.
10. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 8, characterized in that: In step S7, user commands are obtained, and based on the user commands and the scene object distribution topology graph, the target object corresponding to the user commands is output using a large language model. The specific process is as follows: Obtain user instructions; 1) For user commands that directly extract the description of the target object, use a large language model to extract the text describing the target object. Using the SEEM model The text encoder in the code will describe the target object in text. Convert to query semantic embedding vector : Compute semantic embedding vector dictionary CLIP semantic embedding vectors of all object instances With query vector Cosine similarity between The calculation results are sorted top-k, and the semantic embedding vector with the highest similarity is selected and added to the semantic embedding vector dictionary. The corresponding object instance is taken as the target object, and the semantic embedding vector dictionary key-value pair of the target object is output. 2) For user instructions whose descriptions of the target object cannot be directly extracted, the user instructions and the topological relationship diagram of scene object distribution are parsed simultaneously using the thinking chain technology of the large language model, and the semantic embedding vector dictionary key-value of the target object is output.
11. The indoor service robot navigation planning method based on open semantic mapping and large language model according to claim 9, characterized in that: In S8, the target object output in S7 is determined in the TSDF voxel map obtained in S5. The center coordinates of the target object are converted into coordinates on a 2D grid map. The path between the indoor service robot and the target object is then planned within the 2D grid map. The specific process is as follows: Based on the semantic embedding vector dictionary key values of the target object output by S7, the center coordinate attributes of the corresponding nodes of the target object in the scene object distribution topology map are extracted, that is, the TSDF voxel map of the target object obtained by S5. The center coordinates in Set the center coordinates of the target object Coordinates converted to a 2D raster map Based on the coordinates of the indoor service robot in the two-dimensional grid map coordinate system Calculate the vector between the center coordinates of the target object and the coordinates of the indoor service robot: Along the vector The algorithm searches backward in the 2D grid map, using the resolution of the 2D grid map as the step size and combining the shape parameters of the indoor service robot, until the nearest available expected stopping pose of the indoor service robot is found. Then, a graph search global path planning algorithm is used to generate a global path from the current position of the indoor service robot to the expected stopping pose. The graph search global path planning algorithm includes Dijkstra's algorithm, breadth-first search, and Floyd-Warshall algorithm.
Citation Information
Cited By
Robot operation method and system based on space-time constraint enhancement
CN121374640A