3D Semantic Segmentation Method for Complex Open Scenes
By integrating multimodal data of point cloud, image and text description, and using contrast learning and geometric contrast loss terms, the data inadequate and semantic unclearity of three-dimensional semantic segmentation in complex open scenes are solved, and precise segmentation and adaptive segmentation of complex scenes are achieved.
Patent Information
- Application Number
- CN202410997936.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-07-24
AI Technical Summary
The prior art faces the problems of insufficient data volume, complex scenes and unclear semantic definitions in the three-dimensional semantic segmentation task in complex open scenarios, and the model has limited understanding and adaptability to natural real-world scenarios.
Integrate multimodal data of point cloud, two-dimensional images and text descriptions, extract features and perform comparison learning through the pre-training stage, and use comparison learning and geometric comparison loss terms to achieve robustness and semantic segmentation of three-dimensional feature expression.
Accurate three-dimensional semantic segmentation is realized in complex open scenarios, and it is adaptable to unlabeled and dynamically changing scenarios, providing good promotion and application value.
Smart Images

Figure CN118968060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and image processing, and particularly to a three-dimensional semantic segmentation method for complex open scenes. Background Art
[0002] Three-dimensional scene segmentation technology refers to the technology of dividing a scene in three-dimensional space into different parts or regions and identifying and classifying each part or region. As a core research field in the field of computer vision, three-dimensional scene segmentation technology plays a crucial role in understanding and analyzing the information of physical space scenes, and can provide technical support for cutting-edge technologies such as virtual / augmented reality, autonomous driving, smart cities, and digital twins. With the development of deep learning methods based on data supervision, good results have been achieved in three-dimensional semantic segmentation tasks, but many challenges still exist in practical applications; especially in open and complex environments, problems such as insufficient data volume, complex scenes, and unclear semantic definitions are often faced. In addition, most existing studies are currently limited to limited and predefined categories, which restricts the model's understanding and adaptation ability to the diversity of natural real-world scenes. Summary of the Invention
[0003] To solve the above problems, the purpose of the present invention is to provide a three-dimensional semantic segmentation method for complex open scenes, which processes complex three-dimensional scene segmentation tasks by integrating multi-modal data including point clouds, two-dimensional images, and text descriptions.
[0004] To achieve the above invention purpose, the present invention adopts the following technical solutions:
[0005] A three-dimensional semantic segmentation method for complex open scenes, which includes a pre-training stage and an inference stage. The pre-training stage includes the following steps:
[0006] S11. In the pre-training stage, collect the object point cloud, text vocabulary, two-dimensional image, and scene point cloud in the original scene, and extract features from the object point cloud, text vocabulary, two-dimensional image, and three-dimensional scene point cloud;
[0007] S12. Align the features of the multi-modal data in step S11, and perform knowledge distillation through contrastive learning to obtain a robust three-dimensional feature extractor;
[0008] The inference stage includes the following steps:
[0009] S2. Extract features from the "prompt" and the three-dimensional scene point cloud; calculate the similarity between the three-dimensional scene point cloud features and the prompt features, and the three-dimensional points with similarity values greater than the set threshold are the selected regions.
[0010] Further, in the above step S11, the process of extracting features from the two-dimensional image is as follows: Set the resolution of the two-dimensional image as H×W, and use a pre-trained image semantic segmentation model for feature extraction to obtain pixel-level feature embeddings, denoted as where W and H respectively represent the width and height of the two-dimensional image, C represents the dimension of the features, and i represents the i-th image.
[0011] Further, in the above step S11, the process of extracting features from the text vocabulary is as follows: Each text vocabulary uses the "a" + label + "in a scene" template and the CLIP text encoder for feature extraction to obtain a C-dimensional feature vector, denoted as
[0012] Further, in the above step S11, the process of extracting features from the object point cloud is as follows: The object point cloud uses the PointMAE network for feature extraction to obtain a C-dimensional feature vector, denoted as
[0013] Further, in the above step S11, the process of extracting features from the three-dimensional scene point cloud is as follows: Use the MinkUNet18 network for feature extraction to obtain the feature representation of each three-dimensional point in the scene point cloud, denoted as where j represents the j-th three-dimensional point.
[0014] Further, in the above step S12, the alignment of the image pixels and the three-dimensional point features includes the following sub-steps:
[0015] (1) 2D-3D matching association: Given the three-dimensional scene point cloud P and the multi-view image set I collected in this scene, collect the pre-calibrated internal and external camera parameters of each image, including the camera internal parameter K, the camera rotation parameter R, and the camera translation parameter T; Use the perspective projection transformation to achieve the correspondence between the two-dimensional image pixel μ i =(u, v) and the three-dimensional point p j The transformation formula is as follows:
[0016]
[0017] where respectively represent the homogeneous coordinates of μ i , p j ;
[0018] For the three-dimensional point p j in the three-dimensional scene, obtain a set of 2D pixels aligned with it, denoted as
[0019] After obtaining the 2D-3D matching pairs, for the three-dimensional point p in the point cloud model j , aggregate the features of all matching pixels, denoted as where avg represents average pooling;
[0020] (2) Self-supervised pre-training of the three-dimensional scene point cloud feature extractor: The feature representation of the three-dimensional point p j obtained by the three-dimensional feature extractor is The feature representation of the three-dimensional point obtained through 2D feature fusion is Then, in a contrastive learning manner, distill the rich semantic information extracted from the two-dimensional image into the three-dimensional point cloud encoder, and calculate the loss using cosine similarity measurement. The loss L pcs refers to the difference between the three-dimensional feature representation obtained from multi-view image feature fusion and the three-dimensional feature representation obtained through the three-dimensional feature extractor,
[0021]
[0022] where cosin represents cosine similarity determination; K represents the number of three-dimensional points visible from the multi-view image that coincide with the points in the scene point cloud.
[0023] Furthermore, in the above step S12, a geometric contrast loss term is introduced during the model pre-training process to make the feature representations of points belonging to the same geometric cluster more similar and the features between different clusters more differentiated; specifically: First, use an unsupervised method to over-segment the three-dimensional point cloud and divide the entire point cloud into geometrically homogeneous bodies; then, set the feature vector of each superpoint cluster as the anchor point, take the points belonging to the same geometric cluster as positive samples, and the points in different geometric clusters as negative samples,
[0024]
[0025] where τ is the temperature coefficient used to control the smoothness of the output distribution; M represents the number of geometric clusters; Ω(m) represents the set of negative samples of the m-th cluster; represents the positive sample feature in the same geometric cluster as the superpoint m; represents the negative sample feature from different geometric clusters.
[0026] Further, in the above step S12, the object point cloud is aligned with the text word features, including the following sub-steps:
[0027] (1) Object point cloud and text matching association: Obtain 3D object-text matching pairs based on an object classification dataset; for a 3D object and a text description, first, respectively pass through a 3D object feature extractor and a text feature extractor to obtain the feature representation of the 3D object and the feature representation of the text description
[0028] (2) Pre-training of the 3D object point cloud feature extractor: For the object point cloud feature extractor, perform contrastive learning between the feature representation of the 3D object and the feature representation of the text description to make the feature representation extracted by the object point cloud feature extractor and the CLIP text feature in the same feature space;
[0029]
[0030] where, L obj represents the contrastive loss between the object point cloud feature and the text feature; cosin represents the cosine similarity determination; N represents the number of 3D object point clouds; represents the feature representation of the nth 3D object; represents the feature representation of the nth text description.
[0031] Furthermore, in the above step S2, feature extraction is performed on the "prompt" and the 3D scene point cloud, and query prompt-based segmentation is adopted, including the following sub-steps:
[0032] According to the given query prompt, that is, a 2D image, text, or object point cloud, first, obtain the feature representation of the prompt through the corresponding feature extractor, and this feature is a C-dimensional vector; for the scene point cloud, through the scene point cloud feature encoder, extract the feature representation of each point, and the feature of each point is a C-dimensional vector; since these feature representations have been aligned to the same feature space, cosine similarity calculation is performed between the feature of the query prompt and the 3D point feature to obtain a similarity value, and segmentation is carried out relying on this similarity value; this similarity value is distributed between 0 and 1, and the larger this similarity value is, the higher the similarity degree of the two features. Set a relatively high threshold in the similarity value. If the calculated similarity coefficient exceeds this threshold, it is considered that the corresponding area belongs to the target segmentation area.
[0033] Furthermore, in the above step S2, feature extraction is performed on the "prompt" and the 3D scene point cloud, and closed-set semantic segmentation based on text vocabulary is adopted, including the following sub-steps:
[0034] Given all the categories to be segmented out, first, use the text encoder of CLIP to calculate a feature vector for each category to obtain an N×C feature matrix F text, where N represents the number of categories and C represents the feature dimension; then, the feature F of the scene point cloud pcs is dot - producted with the text feature set and normalized. The dimension of the feature F of the scene point cloud pcs is K×C, where K is the number of point clouds, to obtain a similarity matrix S of K×N. Among them, S ij represents the probability that the point p i belongs to the category j, and the category with the largest probability is selected as the semantic category of the current 3D point;
[0035]
[0036] Among them, represents the feature expression of the point p in the scene point cloud i ; represents the feature expression of the j - th text word.
[0037] Due to the adoption of the above - mentioned technical solution, the present invention has the following advantages:
[0038] The 3D semantic segmentation method for complex open scenes takes object point clouds, text words, and 2D images as cues to segment corresponding regions from complex 3D scenes. At the same time, this 3D semantic segmentation method also has the ability to perform closed - set semantic segmentation on the entire 3D scene; the point cloud data provides accurate geometric information for the model, while 2D images and text are rich in context and semantic information, and the above - mentioned information is crucial for understanding the overall situation of the scene; it enables the model to not only learn deep - level features that are difficult to obtain from a single modality, but also show unique adaptability and accuracy when facing unlabeled and dynamically changing real - world scenes, and has good popularization and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is the flowchart of the first embodiment of the 3D semantic segmentation method for complex open scenes of the present invention;
[0040] Figure 2 is the scene segmentation comparison diagram of the first embodiment of the 3D semantic segmentation method for complex open scenes of the present invention; Figure 2 a is the original point cloud scene with RGB information; Figure 2 b is the ground - truth map of semantic segmentation;
[0041] Figure 3 is Figure 2 the complete semantic segmentation result diagram of the original point cloud scene with RGB information in a;
[0042] Figure 4 is the semantic segmentation result diagram queried with the text word "table" as an example in the first embodiment; Figure 4a is a result graph that visually presents similarity values. Among them, the redder region 1 indicates stronger similarity, and the bluer region 2 indicates lower similarity; Figure 4 b is a segmentation region graph where the similarity value of the red region 3 is greater than the set threshold;
[0043] Figure 5 is a semantic segmentation result graph for querying a two-dimensional image using a table as an example; Figure 5 a is a two-dimensional image of a table; Figure 5 b is a result graph that visually presents similarity values. Among them, the redder region 1 indicates stronger similarity, and the bluer region 2 indicates lower similarity; Figure 5 c is a segmentation region graph where the similarity value of the red region 3 is greater than the set threshold;
[0044] Figure 6 is a semantic segmentation result graph for querying the object point cloud using a table as an example; Figure 6 a is a point cloud image of a table; Figure 6 b is a result graph that visually presents similarity values. Among them, the redder region 1 indicates stronger similarity, and the bluer region 2 indicates lower similarity; Figure 6 c is a segmentation region graph where the similarity value of the red region 3 is greater than the set threshold. Detailed implementation manners
[0045] The technical solutions of the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.
[0046] As Figure 1 shown, this three-dimensional semantic segmentation method for complex open scenes includes a pre-training stage and an inference stage. The pre-training stage includes the following steps:
[0047] S11. In the pre-training stage, collect the object point cloud, text vocabulary, two-dimensional image, and scene point cloud in the original scene, and extract features from the object point cloud, text vocabulary, two-dimensional image, and three-dimensional scene point cloud;
[0048] The process of extracting features from a two-dimensional image is as follows: Set the resolution of the two-dimensional image to H×W, and use a pre-trained image semantic segmentation model to extract features to obtain pixel-level feature embeddings, denoted as where W and H respectively represent the width and height of the two-dimensional image, C represents the dimension of the feature, and i represents the i-th image; The image semantic segmentation model has been aligned with the CLIP language model. Therefore, the image features and language features share a semantic space;
[0049] The process of extracting features from text vocabulary is as follows: Each text vocabulary uses the "a"+label+"in a scene" template and the CLIP text encoder to extract features to obtain a C-dimensional feature vector, denoted as For example, if the text vocabulary is "table", the "a" + label + "in a scene" template is adopted, that is, the single word is converted to "a table in a scene" for feature extraction;
[0050] The process of extracting features from the object point cloud is as follows: The object point cloud uses the PointMAE network for feature extraction to obtain a C-dimensional feature vector, denoted as
[0051] The process of extracting features from the three-dimensional scene point cloud is as follows: The MinkUNet18 network is used for feature extraction to obtain the feature representation of each three-dimensional point in the scene point cloud, denoted as where j represents the j-th three-dimensional point;
[0052] S12. Align the auxiliary information such as camera parameters among the multi-modal data features in step S11, and perform knowledge distillation through contrastive learning to obtain a robust three-dimensional feature extractor;
[0053] Align the image pixels with the three-dimensional point features, that is Figure 1 in For Figure 1 in The CLIP model has achieved text-image alignment; for Figure 1 in Alignment will be achieved through the pre-calibrated camera parameters. This step can achieve the ability to query and segment the three-dimensional scene with the image as a prompt; for Figure 1 in Align the text and the image. Since the image and the three-dimensional scene are aligned, the text and the three-dimensional scene are indirectly aligned. This step can achieve the ability to query and segment the three-dimensional scene with the text as a prompt; specifically, it includes the following sub-steps:
[0054] (1). 2D-3D matching and association: Given the three-dimensional scene point cloud P and the multi-view image set I collected in this scene, collect the pre-calibrated internal and external camera parameters of each image, including the camera internal parameter K, the camera rotation parameter R, and the camera translation parameter T; use the perspective projection transformation to achieve the correspondence between the two-dimensional image pixel μ i =(u, v) and the three-dimensional point p j The transformation formula is as follows:
[0055]
[0056] where respectively represent μ i and p jThe homogeneous coordinates. For a three-dimensional scene with multiple captured images, for the three-dimensional point p in the three-dimensional scene j , a set of 2D pixels is obtained and aligned with it, denoted as
[0057] After obtaining the 2D-3D matching pairs, for the three-dimensional point p in the point cloud model j , aggregate the features of all matching pixels, denoted as where avg represents average pooling.
[0058] (2) Self-supervised pre-training of the three-dimensional scene point cloud feature extractor: The feature of the three-dimensional point p obtained by the three-dimensional feature extractor j is denoted as The feature representation of the three-dimensional point obtained through 2D feature fusion is denoted as Then, in the way of contrastive learning, distill the rich semantic information extracted from the two-dimensional image into the three-dimensional point cloud encoder. The core of the distillation process lies in minimizing the difference between the three-dimensional point cloud superpoint feature and its corresponding two-dimensional projection feature. The cosine similarity metric is used for loss calculation, and the loss L pcs refers to the difference between the three-dimensional feature expression obtained from multi-view image feature fusion and the three-dimensional feature expression obtained through extraction by the three-dimensional feature extractor
[0059]
[0060] where cosin represents cosine similarity determination; K represents the number of three-dimensional points visible in the multi-view image that coincide with the points in the scene point cloud, that is, the coincident points are visible in the multi-view image and are included in the point cloud scene
[0061] Align the object point cloud with the text word features, that is Figure 1 in If the three-dimensional object point cloud is used as a prompt to query and segment the scene point cloud, the most intuitive way is to align between the object point cloud feature and the scene point cloud feature. Given that after the previous step, the three-dimensional scene point cloud feature has been aligned with the text feature, it is only necessary to align the object point cloud feature with the text feature. Next, the method for realizing the feature alignment between the object point cloud and the text vocabulary will be described in detail, including the following sub-steps
[0062] (1) Object point cloud and text matching association: Based on the object classification dataset, the category to which each object belongs can be clearly known, and it is very convenient to obtain three-dimensional object-text matching pairs; for the three-dimensional object and text description, first, pass through the three-dimensional object feature extractor and the text feature extractor respectively to obtain the feature expression of the three-dimensional object and the feature expression of the text description
[0063] (2) Pre-training of the 3D object point cloud feature extractor: Similar to the training process of the 3D scene point cloud feature extractor, for the object point cloud feature extractor, contrastive learning is performed between the feature representation of the 3D object and the feature representation of the text description so that the feature representation extracted by the object point cloud feature extractor and the CLIP text features are in the same feature space;
[0064]
[0065] where L obj represents the contrastive loss between the object point cloud features and the text features; cosin represents the cosine similarity determination; N represents the number of 3D object point clouds; represents the feature representation of the nth 3D object; represents the feature representation of the nth text description;
[0066] After step S12, a scene point cloud feature extractor and an object point cloud feature extractor are obtained, and the extracted features and the CLIP text features are in the same feature space;
[0067] In the inference stage, the following two methods are included:
[0068] The specific steps of the first method are to extract features from the "prompt" and the 3D scene point cloud, and perform segmentation based on the query prompt: According to the given query prompt, that is, a 2D image, text, or object point cloud, first, obtain the feature representation of the prompt through the corresponding feature extractor, and this feature is a C-dimensional vector; for the scene point cloud, after passing through the scene point cloud feature encoder, extract the feature representation of each point, and the feature of each point is a C-dimensional vector;
[0069] Since these feature representations have been aligned to the same feature space, cosine similarity calculation is performed between the feature of the query prompt and the 3D point feature to obtain a similarity value, and segmentation is performed relying on this similarity value; this similarity value is distributed between 0 and 1, and the larger this similarity value, the higher the similarity degree of the two features. Set a relatively high threshold in the similarity value. If the calculated similarity coefficient exceeds this threshold, it is considered that the corresponding area belongs to the target segmentation area; preferably, the threshold is 0.9;
[0070] The specific steps of the second method are to extract features from the "prompt" and the 3D scene point cloud, and perform closed-set semantic segmentation based on text vocabulary:
[0071] Given all the categories to be segmented, first, use the text encoder of CLIP to calculate a feature vector for each category to obtain an N×C feature matrix Ftext , where N represents the number of categories and C represents the feature dimension; then, the feature F of the scene point cloud pcs is dot - producted with the text feature set and normalized. The dimension of the feature F of the scene point cloud pcs is K×C, where K is the number of point clouds, and a similarity matrix S of K×N is obtained. Among them, S ij represents the probability that the point p i belongs to the category j, and the category with the largest probability is selected as the semantic category of the current 3D point;
[0072]
[0073] Among them, represents the feature expression of the point p in the scene point cloud i ; represents the feature expression of the j - th text vocabulary;
[0074] The similarity between the 3D scene point cloud feature and the prompt feature is calculated, and the 3D points with similarity values greater than the set threshold are the selected regions.
[0075] In the above step S12, in order to enhance the geometric consistency of the learned 3D features, a geometric contrast loss term is introduced during the model pre - training process, making the feature expressions of points belonging to the same geometric cluster more similar and the features between different clusters more differentiated. Specifically: First, the 3D point cloud is over - segmented (including but not limited to VCCS, region growing) using an unsupervised method, and the entire point cloud is divided into geometrically homogeneous bodies; then, the feature vector of each super - point cluster is set as the anchor point, the points belonging to the same geometric cluster are used as positive samples, and the points in different geometric clusters are used as negative samples,
[0076]
[0077] Among them, τ is the temperature coefficient, used to control the smoothness of the output distribution; M represents the number of geometric clusters; Ω(m) represents the set of negative samples of the m - th cluster; represents the positive sample feature in the same geometric cluster as the super - point m; represents the negative sample feature from different geometric clusters.
[0078] Through this framework, not only is it ensured that the 3D point cloud encoder can learn semantic information corresponding to 2D images, but also the geometric consistency between the internal features of the point cloud is strengthened, providing an effective solution for the in - depth understanding and high - quality segmentation of 3D scenes.
[0079] From Figures 2 to 6It can be seen that the 3D semantic segmentation method of the present invention for complex open scenes uses text vocabulary, 2D images, and object point clouds as cues respectively, and can segment corresponding regions from complex 3D scenes. At the same time, it also has the ability to perform complete semantic segmentation on the entire scene.
[0080] Quantitative segmentation results: Quantitative evaluation metrics for complete semantic segmentation on the ScanNet dataset. Among them, Distill represents the point features obtained by the 3D feature extractor after distillation, and Ensemble represents the combined effect of the 3D point features obtained by 2D multi-view fusion and the point features extracted by the 3D feature extractor.
[0081]
[0082] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Those of ordinary skill in the art should understand that: the technical solutions recorded in the foregoing embodiments can be modified, or some of the technical features can be equivalently replaced; these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A 3D semantic segmentation method for complex open scenarios, characterized in that: It includes a pre-training stage and an inference stage. The pre-training stage includes the following steps: S11. In the pre-training stage, collect the object point cloud, text vocabulary, two-dimensional image, and scene point cloud in the original scene, and extract features from the object point cloud, text vocabulary, two-dimensional image, and three-dimensional scene point cloud; S12. Align the multi-modal data features in step S11, and obtain a robust three-dimensional feature expressor through contrastive learning; among them, the alignment of image pixels and three-dimensional point features includes the following sub-steps: (1) 2D-3D Matching Association: Given a three-dimensional scene point cloud P and a set of multi-view images I collected in this scene, collect the pre-calibrated internal and external camera parameters of each image, including the camera internal parameter K, the camera rotation parameter R, and the camera translation parameter T; use the perspective projection transformation to achieve the correspondence between the two-dimensional image pixel μ i =(u, v) and the three-dimensional point p j The transformation formula is as follows: Among them, respectively represent the homogeneous coordinates of μ i , p j ; For a 3D point p in a 3D scene j , obtain a set of 2D pixels that align with it, denoted as After obtaining the 2D-3D matching pairs, for the three-dimensional point p in the point cloud model j , aggregate the features of all matching pixels, denoted as where avg represents average pooling; (2) Self-supervised pre-training of 3D scene point cloud feature extractor: The 3D point p obtained by the 3D feature extractor j The characteristic is expressed as The feature representation of the 3D point obtained by 2D feature fusion is: Then, the rich semantic information extracted from the two-dimensional image is distilled into the three-dimensional point cloud encoder by contrastive learning, and the cosine similarity metric is used for loss calculation. pcs Refers to the difference between the 3D feature expression obtained from multi-view image feature fusion and the 3D feature expression obtained by 3D feature extractor extraction. where cosin represents the cosine similarity determination; K represents the number of overlapping points between the three-dimensional points visible from the multi-view image and the scene point cloud; The inference stage includes the following steps: S2. Extract features from the "prompt" and the three-dimensional scene point cloud; calculate the similarity between the three-dimensional scene point cloud features and the prompt features, and the three-dimensional points with a similarity value greater than the set threshold are the selected regions.
2. The three-dimensional semantic segmentation method for complex open scenarios according to claim 1, wherein: In the step S11, the process of extracting features from the two-dimensional image is as follows: set the resolution of the two-dimensional image to H×W, and use a pre-trained image semantic segmentation model to extract features, obtaining pixel-level feature embeddings, denoted as where W and H respectively represent the width and height of the two-dimensional image, C represents the dimension of the features, and i represents the i-th image.
3. The three-dimensional semantic segmentation method for complex open scenarios according to claim 1, characterized in that: In the step S11, the process of extracting features from text vocabulary is as follows: Each text vocabulary uses the "a" + label + "in a scene" template and the CLIP text encoder to extract features, obtaining a feature vector of C dimensions, denoted as 4. The three-dimensional semantic segmentation method for complex open scenarios according to claim 1, wherein: In the step S11, the process of extracting features from the object point cloud is as follows: The object point cloud is used in the PointMAE network for feature extraction to obtain a feature vector of dimension C, denoted as 5. The 3D semantic segmentation method for complex open scenarios according to claim 1, characterized in that: In the step S11, the process of extracting features from the 3D scene point cloud is as follows: The MinkUNet18 network is used for feature extraction to obtain the feature representation of each 3D point in the scene point cloud, denoted as where j represents the j-th 3D point.
6. The 3D semantic segmentation method for complex open scenarios according to claim 1, characterized in that: In the step S12, a geometric contrast loss term is introduced during the model pre-training process to make the feature representations of points belonging to the same geometric cluster more similar and the features between different clusters more differentiated. Specifically, first, the 3D point cloud is over-segmented using an unsupervised method to divide the entire point cloud into geometrically homogeneous bodies. Then, the feature vector of each superpoint cluster is set as an anchor point, the points belonging to the same geometric cluster are used as positive samples, and the points in different geometric clusters are used as negative samples. Among them, τ is the temperature coefficient, which is used to control the smoothness of the output distribution; M represents the number of geometric clusters; Ω(m) represents the set of negative samples of the m-th cluster; represents the positive sample features in the same geometric cluster as the superpoint m; represents the negative sample features from different geometric clusters.
7. The 3D semantic segmentation method for complex open scenarios according to claim 1, characterized in that: In step S12, the alignment of the object point cloud and the text word features includes the following sub-steps: (1) Object point cloud and text matching and association: Obtain three-dimensional object-text matching pairs based on an object classification dataset; for the three-dimensional object and the text description, first, respectively pass through a three-dimensional object feature extractor and a text feature extractor to obtain the feature expression of the three-dimensional object and the feature expression of the text description (2) Pre-training of the 3D object point cloud feature extractor: For the object point cloud feature extractor, contrastive learning is carried out between the feature representation of the 3D object and the feature representation of the text description so that the feature representation extracted by the object point cloud feature extractor and the CLIP text features are in the same feature space; Among them, L obj represents the contrast loss between the object point cloud features and the text features; cosin represents the cosine similarity determination; N represents the number of three-dimensional object point clouds; represents the feature expression of the nth three-dimensional object; represents the feature expression of the nth text description.
8. The 3D semantic segmentation method for complex open scenarios according to claim 1, characterized in that: In step S2, when extracting features from the "prompt" and the three-dimensional scene point cloud, the query prompt-based segmentation is adopted, including the following sub-steps: According to the given query prompt, that is, the two-dimensional image, text, or object point cloud. First, obtain the feature expression of the prompt through the corresponding feature extractor, and this feature is a C-dimensional vector; for the scene point cloud, after passing through the scene point cloud feature encoder, extract the feature expression of each point, and the feature of each point is a C-dimensional vector; since these feature expressions have been aligned to the same feature space, the cosine similarity is calculated between the feature of the query prompt and the three-dimensional point feature to obtain a similarity value, and segmentation is carried out based on this similarity value; this similarity value is distributed between 0 and 1, and the larger this similarity value is, the higher the similarity degree of the two features. Set a relatively high threshold in the similarity value. If the calculated similarity coefficient exceeds this threshold, it is considered that the corresponding region belongs to the target segmentation region.
9. The three-dimensional semantic segmentation method for complex open scenarios according to claim 1, characterized in that: In step S2, when extracting features from the "prompt" and the three-dimensional scene point cloud, the closed-set semantic segmentation based on text vocabulary is adopted, including the following sub-steps: Given all the categories to be segmented, first, use the text encoder of CLIP to calculate a feature vector for each category, obtaining an N×C feature matrix F text , where N represents the number of categories and C represents the feature dimension; then, take the dot product of the features F pcs of the scene point cloud with the text feature set and normalize it. The features F pcs of the scene point cloud have a dimension of K×C, where K is the number of points in the point cloud, obtaining a K×N similarity matrix S. Among them, S ij represents the probability that point p i belongs to category j, and select the category with the highest probability as the semantic category of the current 3D point; Among them, represents the feature expression of point p in the scene point cloud i ; represents the feature expression of the j-th text word.
Citation Information
Patent Citations
Target detection method and device, computer equipment and storage medium
CN117636330A