A zero-shot open vocabulary scene understanding method based on semantic instance generation
By using a semantic instance generation method, we have solved the problems of inaccurate 3D instance segmentation and low spatial representation efficiency in zero-shot open-vocabulary 3D scene understanding. This method achieves more accurate 3D instance merging and efficient spatial representation, thereby improving the performance of downstream tasks.
Patent Information
- Application Number
- CN202411362732.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-09-27
AI Technical Summary
Existing technologies for zero-shot open-vocabulary 3D scene understanding suffer from problems such as inaccurate 3D instance segmentation, inaccurate feature aggregation, and low spatial representation efficiency. In particular, in complex scenes, there are issues such as over- or under-segmentation of objects, neglect of feature relationships, and large storage space consumption of point clouds.
By employing semantic instance generation methods, including 2D mask information extraction, semantic feature projection, undersegmentation filtering and merging threshold linear decay strategy, feature aggregation and construction of hybrid octree-graph structures, accurate merging and efficient spatial representation of 3D instances are achieved.
It effectively alleviates the problem of inaccurate 3D instance segmentation, improves the performance of downstream tasks, and enhances the accuracy and storage efficiency of scene understanding through efficient spatial representation methods.
Smart Images

Figure CN119478387B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, and in particular to a zero-shot open-vocabulary scene understanding method based on semantic instance generation. BACKGROUND
[0002] The zero-shot open-vocabulary 3D scene understanding task aims to model the semantics and spatial relationships of objects in a scene, thereby enabling downstream tasks such as object retrieval and target navigation. Some methods achieve fixed-vocabulary scene understanding by training on specific scenes, but lack the ability to generalize to new scenes and rich semantics.
[0003] In recent years, the development of various visual base models and multi-modal large models has promoted the progress of zero-shot open-vocabulary 3D scene understanding. However, the existing technology mainly has the following shortcomings:
[0004] (1) Inaccurate 3D instance segmentation. Current visual language models (VLMs) often over-segment or under-segment objects when dealing with complex scenes. Additionally, there are a large number of partially observed objects in scan sequences, which ultimately affect the accurate modeling of 3D instances.
[0005] (2) Inaccurate feature aggregation. Existing methods typically aggregate a set of 2D semantic features for each instance, which ignores the feature relationships between different instances, making it difficult to distinguish fine-grained semantics.
[0006] (3) Low spatial representation efficiency. Existing methods typically use point clouds to represent 3D space, which occupies a large amount of storage space in large-scale scenes and requires additional computation to support spatial occupancy queries and path planning. Some methods use grids or voxels to simplify spatial representation, but cannot accurately describe the scene.
[0007] In summary, there is currently a lack of a zero-shot open-vocabulary scene understanding method to address or partially address the aforementioned problems. SUMMARY
[0008] The present application provides a zero-shot open-vocabulary scene understanding method based on semantic instance generation to overcome the deficiencies of the prior art and address or partially address the problems of inaccurate 3D instance modeling and low spatial representation efficiency.
[0009] The object of the present application can be achieved by the following technical solutions:
[0010] In one aspect of the present application, a zero-shot open vocabulary scene understanding method based on semantic instance generation is provided, comprising the following steps:
[0011] In step S1, multi-frame image data including depth information is obtained, for each frame of image data, 2D mask information is obtained by segmentation, 3D point cloud segments are obtained by performing semantic feature extraction on the 2D mask and projecting it to 3D space, and 3D instances are obtained by grouping and merging the 3D point cloud segments into 3D instances through a time sequence grouping and merging strategy with linear decay of merging threshold and under the guidance of semantics, thereby obtaining a 3D instance set.
[0012] In step S2, the semantic features of each instance in the 3D instance set are aggregated under the premise of considering intra-class similarity and inter-class difference.
[0013] In step S3, based on the 3D instance set after feature aggregation, each 3D instance is taken as a graph node, the aggregated features in step S2 are taken as the semantic information of the 3D instance, the spatial occupancy information of the 3D instance is represented by a hybrid octree with adaptive target shape and size, the spatial relationship between different 3D instances is represented by the edges between nodes, a hybrid octree-graph structure is constructed, and zero-shot open vocabulary scene understanding is realized.
[0014] As a preferred technical solution, in step S1, for each frame of image data, the process of obtaining 3D point cloud segments by performing semantic feature extraction on the 2D mask and projecting it to 3D space includes the following steps:
[0015] In step S101, for each frame of image data, a plurality of 2D masks are obtained by using a pre-trained segmentation model.
[0016] In step S102, the semantic features corresponding to each 2D mask are generated by using a pre-trained visual language model.
[0017] In step S103, the 2D mask is projected to 3D space based on camera parameters and camera pose to obtain 3D point cloud segments.
[0018] As a preferred technical solution, in step S1, the process of grouping and merging the 3D point cloud segments into 3D instances through a time sequence grouping and merging strategy with linear decay of merging threshold and under the guidance of semantics includes the following steps:
[0019] In step S104, time sequence grouping: the 3D point cloud segments to be merged are divided into N groups according to the time frame they are in, and the first group of 3D point cloud segments is taken as the set of point cloud segments to be merged in the current round.
[0020] Step S105, for the set of point cloud segments to be merged in the current round, the spatial similarity and semantic similarity between any two point cloud segments are calculated;
[0021] Step S106, semantic-guided under-segmentation filtering: for each parent 3D point cloud segment containing a child 3D point cloud segment in the set of point cloud segments to be merged in the current round, if the variance between the semantic similarity of the parent 3D point cloud segment and the child 3D point cloud segment is greater than a preset value, the parent 3D point cloud segment is deleted;
[0022] Step S107, linear decay strategy of merging threshold: for the set of point cloud segments to be merged after semantic-guided under-segmentation filtering, 3D point cloud segments with spatial similarity and / or semantic similarity higher than a threshold are selected for iterative merging to obtain a set of relay instances after merging in the current round, wherein the threshold decreases with the increase of the iteration round.
[0023] Step S108, the union of instances in the current set of relay instances and the next set of 3D point cloud segments is taken as a new set of point cloud segments to be merged, and step S105 is executed to obtain a final set of 3D instances after N iterations.
[0024] As a preferred technical solution, the spatial similarity includes the intersection-over-union of the two 3D point cloud segments, and the semantic similarity is the cosine similarity of the corresponding semantic features of the two 3D point cloud segments.
[0025] As a preferred technical solution, in step S2, the process of obtaining the merged 3D instance set through feature aggregation under the premise of considering intra-class similarity and inter-class difference includes the following steps:
[0026] Step S201, for each instance in the set of 3D instances, the instance-associated 2D mask set is obtained by back-projection, and the semantic features are recalculated;
[0027] Step S202, for each instance-associated 2D mask set, the main feature cluster is extracted, the center feature corresponding to the main feature cluster is calculated, and the neighbor instances of each instance are calculated according to the center feature;
[0028] Step S203, for each 2D mask associated with each instance, the difference between the semantic similarity of the 2D mask corresponding to the semantic feature and the center feature of the instance to which it belongs and the center features of all neighbor instances is taken as a dynamic weight, and the aggregated 3D instance set is obtained by aggregation.
[0029] As a preferred technical solution, in the step S3, the graph structure comprises nodes and edges, wherein the nodes comprise semantic information of the object, the target shape and size adaptive octree and the center coordinates of the object, and the edges comprise semantic of spatial relationship between nodes, Euclidean distance between nodes and three-dimensional direction vector between nodes.
[0030] As a preferred technical solution, in the target shape and size adaptive octree, the shape of the voxel of the tree structure matches the shape of the object.
[0031] Another aspect of the present application provides a target query method, comprising the following steps:
[0032] Step S1, obtaining the graph structure obtained by the zero-shot open vocabulary scene understanding method based on semantic instance generation;
[0033] Step S2, obtaining user instruction information, extracting semantic information and orientation information by using a large language model, querying a target by using the graph structure, and returning information of the target.
[0034] Another aspect of the present application provides a path search method, comprising the following steps:
[0035] Step S1, obtaining the graph structure obtained by the zero-shot open vocabulary scene understanding method based on semantic instance generation;
[0036] Step S2, obtaining information of a target query point, performing coarse-grained query in nodes of the graph structure, and obtaining nodes comprising the query point.
[0037] Step S3, recursively querying whether there is an obstacle at the position of the target query point in the octree included in the node, and returning the query result, to realize path search.
[0038] Another aspect of the present application provides an electronic device, comprising one or more processors and a memory, wherein the memory stores one or more programs, and the one or more programs comprise instructions for executing the zero-shot open vocabulary scene understanding method based on semantic instance generation.
[0039] Compared with the prior art, the present application has at least one of the following beneficial effects:
[0040] (1) effectively alleviating the problem of inaccurate object segmentation: by adopting the semantic-guided under-segmentation filtering and merging threshold linear decay strategy, and by combining 3D point cloud fragments into 3D instance sets through time sequence grouping, the present application alleviates the problem of inaccurate instance segmentation compared with the prior art.
[0041] (2) Downstream task performance is good: under the premise of considering the intra-class similarity and inter-class difference, the 3D instance set after aggregation is obtained through feature aggregation, and the intra-class similarity and inter-class difference of the instance features are considered, the more distinguishable semantic features are aggregated for each instance, and the performance of the downstream task is improved.
[0042] (3) High spatial representation efficiency: based on the aggregated 3D instance set, the graph structure including the target shape and size adaptive octree is constructed, the zero-shot open vocabulary scene understanding is realized, the semantic and spatial occupation of the instance in the 3D scene are represented by constructing an efficient octree-graph structure, and the spatial representation efficiency is effectively improved. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 A flowchart of the zero-shot open vocabulary scene understanding method based on semantic instance generation in the embodiment;
[0044] Figure 2 A scheme diagram of the zero-shot open vocabulary scene understanding in the embodiment;
[0045] Figure 3 A diagram of the qualitative experimental results of the scheme in the embodiment. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be described clearly and completely in the embodiments of the present application combined with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor shall belong to the scope of protection of the present application.
[0047] Embodiment 1
[0048] In view of the problems in the prior art, the embodiment provides a zero-shot open vocabulary scene understanding method based on semantic instance generation, which aims to:
[0049] (1) Provide a three-step instance construction method without training, including 2D mask generation, time grouping 3D segment merging and distinguishable instance feature aggregation, to reduce the instance over-segmentation and under-segmentation problems caused by inaccurate segmentation model, so as to generate more accurate 3D instances and more distinguishable feature representation.
[0050] (2) Provide an efficient spatial representation method, that is, a hybrid octree-graph structure is used, which can more accurately represent the 3D space while reducing the storage requirement. Not only improves the accuracy of spatial representation, but also facilitates efficient deployment on real machines with limited computing and storage resources.
[0051] See also Figure 1 This method is roughly divided into 3D instance collection, instance feature aggregation, scene octree-graph construction and downstream application. The specific method includes the following steps:
[0052] Step S1: establishing an instance set.
[0053] The 3D instance set construction process is as follows Figure 2 As shown in (a), this step first uses a visual basic model such as SAM (Segment Anything Model) to extract 2D masks from RGB-D (i.e., multi-frame image data including depth information) input, and obtains a set of category-independent 2D masks for each frame of RGB image. Subsequently, a pre-trained visual language model is used to generate semantic features for each 2D mask. The features specifically include visual features extracted from the image region and text description features of this region. Next, the 2D mask is projected into 3D space according to the camera parameters and the current camera pose to obtain a 3D point cloud fragment.
[0054] In order to accurately merge all 3D point cloud fragments into 3D instance point clouds, this embodiment provides a temporal grouping merging algorithm. First, each 2D mask is equally divided into N groups according to the time frame in which it is located, and each group contains I frame. Then, this embodiment merges all 2D mask groups N times to obtain the final 3D instance set. Specifically, the 0th group of point clouds is first merged to obtain the relay instance set M0, and then when performing the kth merge, M is used k-1 The instances contained in and the union of the k-th group of 3D point cloud fragments are merged as input until the final set of instances is obtained.
[0055] In order to merge each group of 3D point cloud fragments, this embodiment provides a semantic-guided under-segmentation filtering and merging threshold linear attenuation strategy. Specifically, the spatial similarity and semantic similarity between any two point cloud fragments are first calculated. The spatial similarity includes the IoU and the inclusion rate IoR of the two point cloud fragments. The inclusion rate is calculated as a prerequisite for under-segmentation processing, where the specific meaning of IoR(i, j) is the proportion of the common point cloud of point cloud fragment i and point cloud fragment j to point cloud fragment j, that is, when fragment i contains fragment j, IoR(i, j) takes the maximum value of 1. The cosine similarity value of the semantic features of the two point cloud fragments is used as the semantic similarity. Subsequently, for any parent point cloud fragment containing a child point cloud fragment, if the variance of its semantic similarity with all the child point cloud fragments it contains is greater than the similarity threshold, it is regarded as an under-segmented object and filtered. After filtering the under-segmented fragments, all point cloud fragments with high spatial similarity and semantic similarity are iteratively merged. In addition, in order to merge the point cloud segments obtained from partial observations in RGB-D, the similarity threshold is linearly decayed according to the number of iterations, so that the high-similar point cloud segments of the complete observation are merged when the threshold is high, and the point cloud segments with low spatial similarity are merged when the threshold is low in the later stage.
[0056] Step S2: instance feature aggregation.
[0057] The process of instance feature aggregation is as follows Figure 2 As shown in (b), after obtaining the 3D instance set, this embodiment provides a feature aggregation method that simultaneously considers intra-class similarity and inter-class differences. Specifically, for a set of 2D masks associated with each instance, the reconstructed 3D instance is back-projected to correct each mask and recalculate the semantic features of the 2D mask, thereby reducing the noise impact of over-segmentation on the features. Next, this embodiment uses the DBSCAN clustering algorithm to extract the main feature clusters of the 2D features associated with each instance to filter out features of partial observations and imperfect perspectives. Subsequently, the central feature of the main feature cluster of each instance is calculated, and the neighbor instances of each instance are calculated based on the central feature. Based on this, a dynamic weight is assigned to the different features of each instance to aggregate the final representative instance features. Each dynamic weight is the difference between the similarity between the feature and the central feature of its own instance minus the similarity between the feature and the central features of all neighboring instances, and is normalized by the Softmax function. Specifically, the dynamic weight is calculated using the following formula:
[0058]
[0059] in, is the central feature of the i-th instance, obtained by averaging all semantic features of the i-th instance, is the jth feature of the i-th instance, is the set of neighbor instances of the i-th instance, is the central feature of the neighbor instance k of the i-th instance.
[0060] In this step, only the semantic features of each instance in the 3D instance set are aggregated, and the spatial information of the instance itself will not change.
[0061] Step S3: establishing a scene octree graph.
[0062] The process of establishing the scene octree graph is as follows Figure 2 As shown in (c), to better describe the spatial positional relationships and occupancy information of objects in the scene, this embodiment provides a new structure that combines a graph structure with an adaptive octree. The graph structure, as the upper layer, stores the semantics of the objects in the scene and the spatial relationships between similar objects. For each semantic segmentation result, an adaptive octree is embedded in each node of the graph to accurately describe its spatial occupancy information.
[0063] Specifically, the upper-level graph consists of nodes and edges. Each node contains three attributes: semantic information about the object, its adaptive octree, and its center. Regarding edge construction, a distance threshold is used to determine whether to generate an edge between the current node and another node. When the distance between nodes is less than a certain threshold, the semantics of the spatial relationship of the other node relative to the current node are calculated. The Euclidean distance and the three-dimensional spatial vector from the current node to the other node are used as the three edge attributes.
[0064] An octree is a tree-like data structure that can efficiently describe three-dimensional space. When constructing an octree, the root node is the smallest square bounding box containing a point cloud. This box is a square voxel with a center of c and a side length of d. This box is then split into eight octants of length d / 2, each serving as child nodes, using axis-aligned planes. Each child node is then recursively split until the octree reaches the target depth. Traditional octrees use square voxels to store object occupancy information, which has certain limitations. When using square voxels to enclose objects with large horizontal and vertical dimensions, large amounts of empty space remain within the voxels. Only when the octree is deep enough does the size of the child nodes approach the smaller aspect ratio of the object. To address these shortcomings, the present invention provides an adaptive octree structure that adapts to the shape and size of the object to store object occupancy information. Each voxel in the tree structure is similar to the object's bounding box, allowing the voxel size and shape to be adaptively adjusted to the scale of the object. This embodiment constructs an octree using the point cloud after instance segmentation, and uses the minimum rectangular bounding box of the point cloud as the root node of the adaptive octree for segmentation until the target depth is reached. In this way, each child node of the adaptive octree is similar to the minimum bounding box of the object, and less space can be used to describe the object's occupancy information.
[0065] For the constructed hybrid octree-graph structure, each graph node represents a 3D instance, its spatial occupancy information is efficiently represented by an adaptive octree, its semantic information is represented by aggregated features, and the edges between nodes represent the spatial relationship between different 3D instances.
[0066] Step S4, application.
[0067] Object retrieval: based on octree-graph, the embodiment provides target query and path search functions. In terms of target query, a large language model is used to understand and translate the user's natural semantic query statement, such as Figure 2 In (c) of the above, the user instruction is: "Please help me find the table on the right side of the bookshelf". The large language model can extract the main semantic information "table" and "bookshelf" and the orientation information "right side" from the statement, and the large model can further query the target by calling the API Query LLM ("bookshelf", "right side", "table") of the octree-graph. In addition to this, the graph structure provides the function of directly querying the semantic target Query LLM ("semantic information"), and the large model can access the location relationship of objects in the graph through these basic queries.
[0068] Path search: in the navigation task, the occupancy information query is the basis of path search, and the embodiment based on octree-graph provides a coarse and fine-grained combined occupancy information finding method. Coarse-grained query is performed in the nodes of the graph structure, and the node containing the query point is found, and further recursive query is performed in the octree contained in the node to determine whether the target point is located in the position with obstacle information. Through this method, the classic path search algorithm such as A* can be well supported, as shown in (c) of the above. Figure 2
[0069] To verify the effectiveness of the method, 3D semantic segmentation and 3D instance segmentation tasks are selected for quantitative and qualitative experiments. Specifically, the technical solution provided by the embodiment is used for quantitative experiments of zero-shot semantic segmentation and zero-shot instance segmentation tasks on two commonly used datasets Replica and ScanNet, as shown in Figure 3 The experimental results are shown in the schematic diagram. Referring to Tables 1 and 2, the results show that the method of the embodiment has obvious advantages over existing methods: in the quantitative experiment part, the method of the embodiment is significantly ahead of the existing methods in all evaluation indexes; in the qualitative experiment, through the visualization of the results, we find that the method of the embodiment has achieved good results in fine-grained concept discrimination, small object instance modeling and panoramic modeling.
[0070] Table 1 Quantitative experimental results of the semantic segmentation of the embodiment
[0071]
[0072] Table 2 Example segmentation quantitative experiment results
[0073]
[0074] The method has the following beneficial effects:
[0075] (1) The embodiment provides a semantic instance generation method without training, including three-step strategies: 2D segmentation mask generation, time sequence grouping segment merging, and instance feature aggregation, which can effectively solve the influence of the prior art due to inaccurate segmentation on the 3D instance map.
[0076] Specifically, for 3D segment merging, a time sequence grouping segment merging strategy is provided, segments are grouped by time sequence, and each group is merged by a semantic guided under-segmentation filtering and merging threshold linear decay strategy, and finally an accurate 3D instance set is obtained; for feature aggregation, a dynamic feature aggregation method considering both intra-class similarity and inter-class difference is provided, and representative semantic features are extracted for each instance.
[0077] (2) The embodiment provides an efficient hybrid octree-graph structure for efficiently and accurately representing the spatial occupancy information of 3D instances. It represents each instance by an adaptive octree, which can adaptively adjust the size and shape of the voxel according to the geometric shape of the instance, and can realize more accurate description with less storage occupancy compared with the traditional octree. In addition, by introducing mutual spatial relationship in the global graph, it can directly serve downstream tasks such as object retrieval and path planning.
[0078] Embodiment 2
[0079] The embodiment provides an electronic device, including one or more processors and a memory, the memory has one or more programs stored therein, and the one or more programs include instructions for executing the zero-shot open-vocabulary scene understanding method based on semantic instance generation as described in embodiment 1.
[0080] Embodiment 3
[0081] The embodiment provides a computer-readable storage medium, including one or more programs for an electronic device to execute, and the one or more programs include instructions for executing the zero-shot open-vocabulary scene understanding method based on semantic instance generation as described in embodiment 1.
[0082] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements shall be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A zero-shot open vocabulary scene understanding method based on semantic instance generation, characterized in that, The method comprises the following steps: Step S1, obtaining multi-frame image data comprising depth information, for each frame of image data, obtaining 2D mask information by segmentation, obtaining 3D point cloud segments by projecting the 2D mask to 3D space after semantic feature extraction, adopting a semantic-guided under-segmentation filtering and merging threshold linear decay strategy, merging the 3D point cloud segments into 3D instances by time sequence grouping and merging, and obtaining a 3D instance set; Step S2, performing feature aggregation processing on the semantic features of each instance in the 3D instance set under the premise of considering intra-class similarity and inter-class difference; Step S3, based on the 3D instance set after feature aggregation, taking each 3D instance as a graph node, taking the aggregated features in step S2 as the semantic information of the 3D instance, using a target shape and size adaptive hybrid octree to represent the spatial occupancy information of the 3D instance, using the edges between nodes to represent the spatial relationship between different 3D instances, constructing a hybrid octree-graph structure, and realizing zero-shot open vocabulary scene understanding.
2. The zero-shot open-vocabulary scene understanding method based on semantic instance generation of claim 1, wherein, In step S1, for each frame of image data, the process of obtaining 2D mask information by segmentation, projecting the 2D mask to 3D space after semantic feature extraction, and obtaining 3D point cloud segments comprises the following steps: Step S101, for each frame of image data, obtaining a plurality of 2D masks by using a pre-trained segmentation model; Step S102, generating semantic features corresponding to each 2D mask by using a pre-trained visual language model; Step S103, projecting the 2D mask to 3D space based on camera parameters and camera poses to obtain 3D point cloud segments.
3. The zero-shot open-vocabulary scene understanding method based on semantic instance generation of claim 1, wherein, In step S1, the process of adopting a semantic-guided under-segmentation filtering and merging threshold linear decay strategy, merging the 3D point cloud segments into 3D instances by time sequence grouping and merging, and obtaining a 3D instance set comprises the following steps: Step S104, time sequence grouping: dividing the 3D point cloud segments to be merged into N groups according to the time frames they are in, and taking the first group of 3D point cloud segments as the set of point cloud segments to be merged in the current round; Step S105, for the set of point cloud segments to be merged in the current round, calculating the spatial similarity and semantic similarity between any two point cloud segments; Step S106, semantic-guided under-segmentation filtering: for each parent 3D point cloud segment containing a child 3D point cloud segment in the set of point cloud segments to be merged in the current round, if the variance between the semantic similarity of the parent 3D point cloud segment and the child 3D point cloud segment it contains is greater than a preset value, the parent 3D point cloud segment is deleted; Step S107, merging threshold linear decay strategy: for the set of point cloud segments to be merged after semantic-guided under-segmentation filtering, selecting 3D point cloud segments with spatial similarity and / or semantic similarity higher than a threshold for iterative merging to obtain a set of relay instances after merging in the current round, wherein the threshold decreases with the increase of the iteration round; Step S108, taking the union of the instances in the current set of relay instances and the next group of 3D point cloud segments as a new set of point cloud segments to be merged, executing step S105, and obtaining a final 3D instance set after N iterations.
4. The zero-shot open vocabulary scene understanding method based on semantic instance generation of claim 3, wherein, The spatial similarity includes an intersection-over-union of two 3D point cloud segments, and the semantic similarity is a cosine similarity of corresponding semantic features of the two 3D point cloud segments.
5. The zero-shot open vocabulary scene understanding method based on semantic instance generation of claim 1, wherein, In step S2, the feature aggregation process considering the intra-class similarity and inter-class difference comprises the following steps: In step S201, for each instance in the 3D instance set, the instance-associated 2D mask set is obtained by back projection, and the semantic feature is recalculated; In step S202, the main feature cluster extraction is performed for each instance-associated 2D mask set, the center feature corresponding to the main feature cluster is calculated, and the neighbor instances of each instance are calculated according to the center feature; In step S203, for each 2D mask associated with each instance, the difference between the semantic feature corresponding to the 2D mask and the center feature of the instance and the similarity between the semantic feature and the center features of all neighbor instances is calculated as a dynamic weight, and the aggregated 3D instance set is obtained by aggregation.
6. The zero-shot open vocabulary scene understanding method based on semantic instance generation of claim 1, wherein, In step S3, the mixed octree-graph structure includes nodes and edges, wherein the nodes include semantic information of the object, an octree adaptive to the target shape and size, and a center coordinate of the object, and the edges include semantic information of spatial relationships between nodes, Euclidean distances between nodes, and three-dimensional direction vectors between nodes.
7. The zero-shot open vocabulary scene understanding method based on semantic instance generation of claim 1, wherein, In the octree adaptive to the target shape and size, the shape of the voxel of the tree structure matches the shape of the minimum bounding box of the object.
8. A target query method characterized by comprising: The method comprises the following steps: Step S1, obtaining the graph structure obtained by the zero-shot open vocabulary scene understanding method based on semantic instance generation according to any one of claims 1-7; Step S2, obtaining user instruction information, extracting semantic information and orientation information using a large language model, querying a target using the graph structure, and returning information of the target.
9. A route search method characterized by comprising: The method comprises the following steps: Step S1, obtaining the graph structure obtained by the zero-shot open vocabulary scene understanding method based on semantic instance generation according to any one of claims 1-7; Step S2, obtaining information of a target query point, performing coarse-grained query in nodes of the graph structure, and obtaining nodes including the query point; Step S3, recursively querying whether there is an obstacle at the position of the target query point in the octree included in the node, and returning the query result to realize path search.
10. An electronic device, comprising: The method comprises: One or more processors and a memory, the memory storing one or more programs, the one or more programs including instructions for executing the zero-shot open vocabulary scene understanding method based on semantic instance generation according to any one of claims 1-7.
Citation Information
Patent Citations
Large-scale point cloud registration method based on deep semantic graph matching
CN117455967A
Three-dimensional image generation method and device, equipment and storage medium
CN117576303A