A method and system for unannotated point cloud panoramic segmentation driven by a differentiable Gaussian field

By constructing a differentiable 3D Gaussian field and a joint loss function, and combining 3D point cloud and 2D image data, the dependence on annotation and cross-domain adaptability of 3D point cloud panoramic segmentation are solved, achieving efficient and low-cost 3D scene understanding.

CN122115850APending Publication Date: 2026-05-29WUHAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2026-01-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing 3D point cloud panoramic segmentation technology relies on high-cost 3D labeled data and performs poorly in cross-domain applications, making it difficult to meet the needs of large-scale and diversified applications.

Method used

By acquiring 3D point cloud data and multi-view 2D image data, a pre-trained model is used to extract 2D pseudo-labels, and a differentiable 3D Gaussian field is constructed. The 3D feature extractor is then optimized by combining a joint loss function to achieve panoramic segmentation of 3D point clouds.

Benefits of technology

It achieves high-quality 3D point cloud segmentation without the need for 3D annotation, possesses strong cross-domain adaptability and robustness, reduces the technical application threshold and cost, and improves the accuracy and information richness of 3D scene understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115850A_ABST
    Figure CN122115850A_ABST
Patent Text Reader

Abstract

The application provides a Gauss field drivable label-free point cloud panorama segmentation method and system, and belongs to the technical field of three-dimensional scene understanding. The method comprises the following steps: obtaining three-dimensional point cloud data and multi-view two-dimensional image data of a target scene; extracting two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels of the multi-view two-dimensional image data respectively; constructing a differentiable three-dimensional Gauss field according to the three-dimensional point cloud data and a three-dimensional feature extractor; projecting the differentiable three-dimensional Gauss field to each view of the multi-view two-dimensional image data to obtain rendered two-dimensional semantic feature maps and two-dimensional instance embedding maps; constructing a joint loss function, and optimizing parameters of the three-dimensional feature extractor through back propagation to obtain an optimized three-dimensional feature extractor for point cloud panorama segmentation, and obtaining a panorama segmentation result. The application only uses widely available two-dimensional image data to eliminate three-dimensional labeling, realizes accurate point cloud panorama segmentation, and has wide applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and 3D scene understanding technology, and in particular to a differentiable Gaussian field-driven method and system for label-free panoramic segmentation of point clouds. Background Technology

[0002] Accurate and fine-grained 3D scene understanding is a crucial foundation for applications such as autonomous systems, mixed reality, and large-scale 3D reconstruction. As one of the core tasks for achieving this goal, point cloud panoramic segmentation, by simultaneously assigning semantic category labels and unique instance identifiers to each point in the scene, completes a comprehensive and structured analysis of the 3D environment, thereby providing directly usable element-level semantic and instance information for 3D scene understanding. This task plays a vital role in high-level understanding applications that accurately distinguish individual objects (such as vehicles, pedestrians, and streetlights) and precisely infer their spatial relationships.

[0003] Currently, mainstream 3D point cloud panoramic segmentation methods can be divided into two categories: The first category is based on fully supervised deep learning methods. These methods rely on a large amount of manually annotated 3D point cloud data (including point-by-point semantic and instance labels) for model training. Although they achieve high accuracy on standard datasets, creating panoramic segmentation annotations for large-scale point cloud data is an extremely time-consuming, labor-intensive, and specialized task, resulting in high costs and severely hindering the practical deployment and promotion of the technology. Furthermore, the models often overfit to specific training data distributions (such as specific sensors or specific urban landscapes), and their performance degrades significantly when applied to new scenarios with domain differences (such as different acquisition devices, geographical environments, and weather conditions), making it difficult to meet the needs of large-scale, diverse applications.

[0004] The second category is methods based on 2D-3D projection or geometric reasoning. These methods attempt to use readily available 2D image labels or models to aid 3D understanding, avoiding 3D annotation. Direct projection methods directly back-project the prediction results of the 2D image onto the 3D point cloud. This method is affected by severe occlusion, viewpoint limitations, and the sparse correspondence between point cloud and image pixels, resulting in a large amount of noise and inconsistent results. Rule-based methods combine 2D segmentation results with hand-designed geometric rules (such as continuity and coplanarity) for reasoning. These methods are usually fragile, inflexible, and struggle to handle complex and varied real-world scenes, and cannot be optimized end-to-end.

[0005] Therefore, developing a general point cloud panoramic segmentation technology that can alleviate the dependence on 3D ground truth annotation, while possessing strong cross-domain adaptability and accurate instance discrimination, is a key challenge that academia and industry urgently need to overcome. Summary of the Invention

[0006] This invention provides a differentiable Gaussian field-driven label-free point cloud panoramic segmentation method and system to solve the defects of existing point cloud panoramic segmentation technology that rely on ground truth labeling and have poor universality, so as to achieve accurate point cloud segmentation and meet the needs of large-scale and diversified applications.

[0007] In a first aspect, the present invention provides a label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field, comprising: Acquire 3D point cloud data of the target scene and multi-view 2D image data registered with the 3D point cloud data; Two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels are extracted from the multi-view two-dimensional image data respectively; The shape parameters of a three-dimensional Gaussian are constructed based on the local geometric features of each data point in the three-dimensional point cloud data. The semantic features and instance features of each data point are extracted by a three-dimensional feature extractor as the attribute parameters of the three-dimensional Gaussian. A differentiable three-dimensional Gaussian field is constructed based on the shape parameters and the attribute parameters. The differentiable three-dimensional Gaussian field is projected onto each viewpoint of the multi-view two-dimensional image data to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map. Based on the two-dimensional semantic pseudo-labels, the two-dimensional instance pseudo-labels, the two-dimensional semantic feature map, and the two-dimensional instance embedding map, a joint loss function is constructed, and the parameters of the three-dimensional feature extractor are optimized through backpropagation to obtain the optimized three-dimensional feature extractor. The optimized 3D feature extractor is used to perform panoramic point cloud segmentation on the point cloud data to be segmented, and the panoramic segmentation result is obtained.

[0008] According to the present invention, a differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation method is provided, wherein the extraction of two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from the multi-view two-dimensional image data includes: Initial semantic pseudo-labels are extracted from the multi-view two-dimensional image data using a pre-trained semantic segmentation model; The initial instance mask of the multi-view two-dimensional image data is extracted using a pre-trained instance segmentation model; By determining the dominant category of the initial instance mask, conflict resolution and fusion are performed on the initial semantic pseudo-label and the initial instance mask to obtain two-dimensional semantic pseudo-label and two-dimensional instance pseudo-label.

[0009] According to the present invention, a differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation method is provided, wherein constructing the shape parameters of a three-dimensional Gaussian based on the local geometric features of each data point in the three-dimensional point cloud data includes: For each data point in the 3D point cloud data, determine the neighborhood point set of the data point and calculate the covariance matrix of the neighborhood point set.

[0010] The covariance matrix is ​​decomposed into eigenvalues ​​to obtain multiple eigenvalues ​​and multiple eigenvectors; Determine the initial rotation matrix of the three-dimensional Gaussian centered on the data point based on the multiple eigenvectors; The initial scale parameters of the three-dimensional Gaussian are determined based on the multiple eigenvalues; wherein the shape parameters include the position of the data points, the initial rotation matrix, and the initial scale parameters.

[0011] According to the present invention, a method for label-free panoramic segmentation of point clouds driven by a differentiable Gaussian field is characterized in that the step of projecting the differentiable three-dimensional Gaussian field onto each viewpoint of the multi-view two-dimensional image data to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map includes: Based on each perspective of the multi-view two-dimensional image data, the differentiable three-dimensional Gaussian field is projected onto the two-dimensional image plane corresponding to each perspective of the multi-view two-dimensional image data through a three-dimensional Gaussian sputtering differentiable rendering pipeline, to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map.

[0012] According to the present invention, a differentiable Gaussian field-driven label-free point cloud panoramic segmentation method is characterized in that the step of constructing a joint loss function based on the two-dimensional semantic pseudo-labels, the two-dimensional instance pseudo-labels, the two-dimensional semantic feature map, and the two-dimensional instance embedding map includes: Construct a semantic consistency loss function between the two-dimensional semantic feature map and the two-dimensional semantic pseudo-label; Construct an instance contrast learning loss function between the two-dimensional instance embedding graph and the two-dimensional instance pseudo-labels; A joint loss function is constructed using a weighted approach based on the semantic consistency loss function and the instance contrast learning loss function.

[0013] According to the present invention, a differentiable Gaussian field-driven label-free point cloud panoramic segmentation method is provided, which constructs an instance contrast learning loss function between the two-dimensional instance embedding map and the two-dimensional instance pseudo-labels, including: In a single 2D instance embedding image, two pixels belonging to the same 2D instance pseudo-label are defined as a positive sample pair, and two pixels belonging to different 2D instance pseudo-labels are defined as a negative sample pair. Calculate the similarity between the instance embedding vectors corresponding to any two pixels in the two-dimensional instance embedding graph; To maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs, an instance contrastive learning loss function is constructed.

[0014] According to the present invention, a differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation method is provided, wherein the expression of the instance contrastive learning loss function is:

[0015] in, , Each of the two pixels embedded in the two-dimensional instance is an arbitrary two-pixel point. , The corresponding instance embedding vector, The total number of pixels in the embedded image of the degree example. For the positive sample set, For all sample sets, The similarity measurement function is defined as follows: the positive sample set is the set of all positive sample pairs, and the set of all samples is the set that includes all positive sample pairs and negative sample pairs.

[0016] According to the present invention, a differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation method is provided, wherein the panoramic segmentation result is obtained by performing point cloud panoramic segmentation on the point cloud data to be segmented using the optimized 3D feature extractor, including: The optimized 3D feature extractor is used to extract the semantic category and instance embedding vector of each data point in the point cloud data to be segmented. Based on the semantic category corresponding to each data point, data points belonging to the instantiation semantic category are selected, and the instance embedding vectors corresponding to each selected data point are clustered to obtain the panoramic segmentation result.

[0017] Secondly, the present invention also provides a differentiable Gaussian field-driven label-free point cloud panoramic segmentation system, comprising: The data acquisition module is used to acquire three-dimensional point cloud data of the target scene and multi-view two-dimensional image data registered with the three-dimensional point cloud data; The tag extraction module is used to extract two-dimensional semantic pseudo-tags and two-dimensional instance pseudo-tags from the multi-view two-dimensional image data, respectively. The Gaussian field construction module is used to construct the shape parameters of a three-dimensional Gaussian based on the local geometric features of each data point in the three-dimensional point cloud data, extract the semantic features and instance features of each data point as the attribute parameters of the three-dimensional Gaussian through a three-dimensional feature extractor, and construct a differentiable three-dimensional Gaussian field based on the shape parameters and the attribute parameters. The projection rendering module is used to project the differentiable three-dimensional Gaussian field onto each viewpoint of the multi-view two-dimensional image data to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map. The parameter optimization module is used to construct a joint loss function based on the two-dimensional semantic pseudo-label, the two-dimensional instance pseudo-label, the two-dimensional semantic feature map and the two-dimensional instance embedding map, and optimize the parameters of the three-dimensional feature extractor through backpropagation to obtain the optimized three-dimensional feature extractor. The panoramic segmentation module is used to perform panoramic segmentation of the point cloud data to be segmented using the optimized 3D feature extractor, and obtain the panoramic segmentation result.

[0018] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field as described above.

[0019] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field as described above.

[0020] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field as described above.

[0021] The beneficial effects of the technical solutions provided by some embodiments of the present invention include at least the following: 1) This invention provides a differentiable Gaussian field-driven label-free point cloud panoramic segmentation method and system. By acquiring multi-view 2D image data registered with 3D point cloud data of the target scene, high-performance 3D understanding can be achieved using only widely available 2D image data and publicly available pre-trained models. This eliminates the need for 3D annotation required only for 3D point cloud data, alleviating the dependence on extremely costly and cumbersome production of ground truth data for 3D point cloud panoramic segmentation, and greatly reducing the technical application threshold and cost. In addition, this invention constructs a differentiable 3D Gaussian field using 3D point cloud data and a 3D feature extractor. This differentiable 3D Gaussian field serves as a bridge connecting 2D visual knowledge and 3D geometric structure. By constructing a joint loss function, the semantic and instance knowledge from 2D image data is effectively fused, optimized, and "upgraded" to 3D geometric space. Thus, high-quality panoramic segmentation results can be obtained directly through the 3D feature extractor without the need for 3D annotation.

[0022] 2) This invention extracts two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from multi-view two-dimensional image data through pre-trained semantic segmentation model and pre-trained instance segmentation model, respectively, to quickly achieve three-dimensional understanding of the target scene. Furthermore, by determining the dominant category of the initial instance mask, it can resolve conflicts and fuse the initial semantic pseudo-labels and the initial instance mask, effectively improving the accuracy of instance differentiation.

[0023] 3) This invention performs eigenvalue decomposition on the neighborhood point set to which each data point belongs in the 3D point cloud data to obtain local geometric features such as eigenvalues ​​and eigenvectors. These features are then used to initialize the initial rotation matrix and initial scale parameters of the 3D Gaussian. The semantic and instance features of each data point are extracted by a 3D feature extractor as the attribute parameters of the 3D Gaussian. Based on the shape and attribute parameters of the 3D Gaussian corresponding to each data point, a differentiable 3D Gaussian field integrating geometric, semantic, and instance information is constructed, enabling flexible and continuous scene representation, improving the information richness of 3D scene representation, and providing data support for subsequent high-quality 3D scene understanding.

[0024] 4) To address the issue of inconsistent instance identifiers in 2D instance pseudo-labels across different viewpoints, this invention proposes an in-view contrastive learning strategy. Within a single rendered 2D instance embedding map, a contrastive relationship between pixel pairs is constructed based on their corresponding 2D instance pseudo-labels: pixel pairs belonging to the same 2D instance pseudo-label are defined as positive sample pairs, and pixel pairs belonging to different 2D instance pseudo-labels are defined as negative sample pairs. The goal is to maximize the similarity between positive sample pairs while minimizing the similarity between negative sample pairs. An instance contrastive learning loss function is constructed. By optimizing the instance contrastive loss function, the embedding vectors of 3D points corresponding to the same physical entity are closer to each other in the feature space, while the embedding vectors of different entities are farther apart. This cleverly avoids the problem of inconsistent instance identifiers across multiple viewpoints. Without any 3D instance annotation or complex cross-view instance matching, it can autonomously learn highly discriminative instance features, clearly and accurately separating independent object instances in the scene.

[0025] 5) This invention utilizes a semantic consistency loss function between a 2D semantic feature map and 2D semantic pseudo-labels, and an instance contrast learning loss function between a 2D instance embedding map and 2D instance pseudo-labels, to construct a joint loss function. Through backpropagation, the parameters of the 3D feature extractor are iteratively optimized, thereby driving the continuous optimization of the attributes carried by the entire 3D Gaussian field. This ultimately forms a semantic-instance model that is globally consistent with the multi-view 2D supervision signal and self-consistent in 3D space. Furthermore, since the 3D Gaussian field is initialized using the geometric information of the point cloud itself, this iterative optimization process can be accurately anchored to precisely measured 3D coordinates, ensuring that the learned semantic and instance information is closely integrated with the real 3D structure of the scene, resulting in physically reasonable and clearly defined panoramic segmentation results.

[0026] 6) Through the "knowledge enhancement" mechanism of this invention, the prior knowledge of a two-dimensional visual model, which possesses strong versatility and world knowledge, is efficiently transferred to a three-dimensional geometric space. This enables the invention to exhibit far superior adaptability and robustness compared to traditional fully supervised methods when faced with new data of different distributions. It effectively solves the problem of insufficient generalization ability of traditional models, has wide applicability, and provides a novel and practical technical solution for many fields that require deep understanding of three-dimensional scenes, such as autonomous driving, robot navigation, mixed reality, and smart cities. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0028] Figure 1 This is one of the flowcharts of the label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field provided by the present invention; Figure 2 This is a schematic diagram of the process for generating and fusing two-dimensional pseudo-tags provided in an embodiment of the present invention; Figure 3 This is the second flowchart of the label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field provided by the present invention. Figure 4 This is a schematic diagram illustrating the principle of the in-view instance comparison learning strategy provided in an embodiment of the present invention; Figure 5 This is a comparison chart of the qualitative effects of three-dimensional point cloud panoramic segmentation obtained by different methods from a point cloud data under test provided in an embodiment of the present invention; Figure 6This is a comparison of the qualitative effects of three-dimensional point cloud panoramic segmentation obtained under different methods from another point cloud data to be tested, provided in an embodiment of the present invention. Figure 7 This is a schematic diagram of the structure of the label-free point cloud panoramic segmentation system driven by a differentiable Gaussian field provided by the present invention; Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0030] Example 1 Please see Figure 1 , Figure 1 One of the flowcharts for a differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation method provided as an embodiment of the present invention includes: S101. Acquire the 3D point cloud data of the target scene and the multi-view 2D image data registered with the 3D point cloud data; S102. Extract the two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from the multi-view two-dimensional image data respectively; S103. Construct the shape parameters of a three-dimensional Gaussian based on the local geometric features of each data point in the three-dimensional point cloud data. Extract the semantic features and instance features of each data point as the attribute parameters of the three-dimensional Gaussian using a three-dimensional feature extractor. Construct a differentiable three-dimensional Gaussian field based on the shape parameters and attribute parameters. S104. Project the differentiable three-dimensional Gaussian field onto each view of the multi-view two-dimensional image data to obtain the rendered two-dimensional semantic feature map and the two-dimensional instance embedding map. S105. Based on the two-dimensional semantic pseudo-labels, two-dimensional instance pseudo-labels, two-dimensional semantic feature maps and two-dimensional instance embedding maps, construct a joint loss function, and optimize the parameters of the three-dimensional feature extractor through backpropagation to obtain the optimized three-dimensional feature extractor. S106. The optimized 3D feature extractor is used to perform panoramic point cloud segmentation on the point cloud data to be segmented, and the panoramic segmentation result is obtained.

[0031] This invention achieves high-performance 3D understanding by acquiring multi-view 2D image data registered with 3D point cloud data of the target scene, utilizing only widely available 2D image data and publicly available pre-trained models. This eliminates the need for 3D annotation required for using only 3D point cloud data, alleviating the dependence on extremely costly and cumbersome production of ground truth data for 3D point cloud panoramic segmentation, and greatly reducing the technical application threshold and cost. In addition, this invention constructs a differentiable 3D Gaussian field using 3D point cloud data and a 3D feature extractor. This differentiable 3D Gaussian field serves as a bridge connecting 2D visual knowledge and 3D geometric structure. By constructing a joint loss function, the semantic and instance knowledge from the 2D image data is effectively fused, optimized, and "upgraded" to the 3D point cloud space. Thus, high-quality panoramic segmentation results can be obtained directly through the 3D feature extractor without the need for 3D annotation.

[0032] In S101 of this embodiment, the original three-dimensional point cloud data of the target scene and the multi-view two-dimensional image data spatiotemporally synchronized and registered with it are acquired.

[0033] It is understood that the multi-view two-dimensional image data obtained in this embodiment includes known camera intrinsic and extrinsic parameters.

[0034] In S102 of this embodiment, existing machine learning models or neural network models are used to extract two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from multi-view two-dimensional image data, respectively.

[0035] For example, various pre-trained machine learning models / neural network models / vertical large models can be used to complete the task of extracting two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from multi-view two-dimensional image data.

[0036] In some possible embodiments, two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels are extracted from the multi-view two-dimensional image data, including: Initial semantic pseudo-labels are extracted from multi-view two-dimensional image data using a pre-trained semantic segmentation model; The initial instance mask of multi-view 2D image data is extracted using a pre-trained instance segmentation model; By determining the dominant category of the initial instance mask, conflict resolution and fusion are performed on the initial semantic pseudo-label and the initial instance mask to obtain two-dimensional semantic pseudo-label and two-dimensional instance pseudo-label.

[0037] Specifically, Figure 2 The diagram illustrates the process of generating and fusing two-dimensional pseudo-tags according to an embodiment of the present invention. Figure 2As shown, taking the original RGB image as an example, a 2D semantic segmentation model (such as the Mask2Former model) pre-trained on a large and diverse dataset is used to infer for each image, resulting in an initial semantic pseudo-label map, which provides rich category priors, although noise may exist. A general 2D instance segmentation model (such as the SAM model) capable of generating high-quality object masks is then used to process each image, resulting in a series of category-independent initial instance masks, which provide accurate object contour information.

[0038] Since the pseudo-label prediction results from the two sources mentioned above may conflict—for example, an instance mask obtained through SAM model segmentation might cover pixels of both "vehicle" and "road surface"—conflict resolution and knowledge fusion are necessary. This invention adheres to the principle that "the same visual instance should have a single semantic meaning," and for each instance mask, calculates the semantic pseudo-label image for all pixels within it. The semantic category distribution in the image is used to determine the dominant category and its purity, where purity represents the proportion of pixels in the dominant category to the total number of pixels. If the purity is higher than a preset threshold... If the value is 0.8, the instance is considered semantically pure, and the semantics of all pixels within it are uniformly corrected to the dominant category; otherwise, the region is iteratively re-segmented using the SAM model until the resulting sub-masks meet the purity requirement. Ultimately, semantically consistent two-dimensional pseudo-label pairs with clear instance boundaries are obtained. The two-dimensional pseudo-label pair will serve as a supervisory signal to drive the optimization of the three-dimensional feature extractor.

[0039] This invention extracts two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from multi-view two-dimensional image data using pre-trained semantic segmentation models and pre-trained instance segmentation models, respectively, to quickly achieve three-dimensional understanding of the target scene. Furthermore, by determining the dominant category of the initial instance mask, it can resolve conflicts and fuse the initial semantic pseudo-labels and the initial instance mask, effectively improving the accuracy of instance differentiation.

[0040] For example, U-Net networks, fully convolutional networks, and Transformer models can be used to extract initial semantic pseudo-labels for multi-view two-dimensional image data; Mask R-CNN models and Transformer models can also be used to extract initial semantic pseudo-labels for multi-view two-dimensional image data.

[0041] In S103 of this embodiment, the traditional three-dimensional Gaussian sputtering used for efficient color rendering is extended, and a differentiable three-dimensional Gaussian field carrying rich attributes is constructed based on three-dimensional point cloud data and a three-dimensional feature extractor.

[0042] S103 of this embodiment may specifically include the following sub-steps: S103-1. Construct the shape parameters of a three-dimensional Gaussian based on the local geometric features of each data point in the three-dimensional point cloud data; S103-2. Extract the semantic features and instance features of each data point as attribute parameters of the three-dimensional Gaussian using a three-dimensional feature extractor; S103-3. Assign the attribute parameters of the three-dimensional Gaussian corresponding to each data point to the three-dimensional Gaussian centered on the data point to form a three-dimensional Gaussian field.

[0043] In some possible embodiments, S103-1, which constructs the shape parameters of a three-dimensional Gaussian based on the local geometric features of each data point in the three-dimensional point cloud data, includes: For each data point in the 3D point cloud data, determine the neighborhood point set of the data point and calculate the covariance matrix of the neighborhood point set.

[0044] Eigenvalue decomposition of the covariance matrix yields multiple eigenvalues ​​and multiple eigenvectors. The initial rotation matrix of the three-dimensional Gaussian centered on the data point is determined based on multiple eigenvectors; The initial scale parameters of the three-dimensional Gaussian are determined based on multiple eigenvalues; among them, the shape parameters include the initial rotation matrix and the initial scale parameters.

[0045] Specifically, for each data point in the 3D point cloud data Find its M nearest neighbors to form a neighborhood point set, and calculate the covariance matrix of this neighborhood point set. , The calculation formula is:

[0046] Where m = 1, 2, ..., M, For data points The number of neighboring points in the set of neighboring points. For the first The position vectors of the neighboring points It is the mean of the position vectors of all neighboring points in the neighborhood point set.

[0047] For covariance matrix Perform eigenvalue decomposition to obtain eigenvalues. and the corresponding orthogonal eigenvectors The eigenvector matrix As a data point Initial rotation matrix of a 3D Gaussian centered at the center The square root of the eigenvalues ​​is appropriately scaled and used as the initial scale parameter. Based on the location of the data points Initial rotation matrix and initial scale parameters This initializes a 3D Gaussian geometry. This operation ensures that the Gaussian ellipsoid corresponding to each 3D Gaussian closely fits the local surface geometry from the very beginning of optimization. As a powerful geometric regularizer, it effectively prevents the generation of distorted Gaussians that violate physical structures in occluded or sparse regions during the optimization process, ensuring the geometric basis for subsequent feature learning.

[0048] In S103-2 of this embodiment, a three-dimensional feature extractor with a sparse convolutional network as its backbone is built. The 3D point cloud data is input into this 3D feature extractor, which outputs the depth features of each data point through multi-level feature encoding. Subsequently, it is processed through two lightweight multilayer perceptron decoders. and Each Mapping to semantic features and instance features Semantic features can be expressed as semantic logical value vectors, and instance features can be expressed as instance embedding vectors.

[0049] Will As an attribute parameter, it is assigned to the point. The three-dimensional Gaussian field corresponding to each data point is obtained by centered on the three-dimensional Gaussian field, which includes shape parameters and attribute parameters. The three-dimensional Gaussian fields corresponding to all data points in the three-dimensional point cloud data form a three-dimensional Gaussian field, thereby constructing a differentiable three-dimensional Gaussian field that combines accurate geometry, semantic information and instance features.

[0050] This invention performs eigenvalue decomposition on the neighborhood point set to which each data point belongs in 3D point cloud data to obtain local geometric features such as eigenvalues ​​and eigenvectors. These features are then used to initialize the initial rotation matrix and initial scale parameters of a 3D Gaussian field. Semantic and instance features of each data point are extracted using a 3D feature extractor as attribute parameters of the 3D Gaussian field. Based on the shape and attribute parameters of the 3D Gaussian field corresponding to each data point, a differentiable 3D Gaussian field integrating geometric, semantic, and instance information is constructed. This enables flexible and continuous scene representation, improves the information richness of 3D scene representation, and provides data support for subsequent high-quality 3D scene understanding.

[0051] In S104 of this embodiment, the differentiable three-dimensional Gaussian field obtained in S103 is projected onto a two-dimensional image plane, and a two-dimensional feature map is generated by differentiable rendering.

[0052] In some possible embodiments, a differentiable three-dimensional Gaussian field is projected onto various perspectives of the multi-view two-dimensional image data to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map, including: Based on the various perspectives of the multi-view 2D image data, a differentiable 3D Gaussian field is projected onto the 2D image plane corresponding to each perspective of the multi-view 2D image data through a 3D Gaussian sputtering differentiable rendering pipeline, resulting in a rendered 2D semantic feature map and a 2D instance embedding map.

[0053] Specifically, for each known camera viewpoint of the multi-view 2D image data, a differentiable 3D Gaussian field is projected onto the 2D image plane corresponding to each viewpoint of the multi-view 2D image data through a 3D Gaussian sputtering differentiable rendering pipeline. Similar to rendering RGB colors, a 2D semantic feature map can be rendered. and 2D instance embedding graph Among them, the two-dimensional semantic feature map can be expressed in the form of a two-dimensional semantic logical value map:

[0054] in, It is the opacity after Gaussian projection. It is the cumulative transmittance. Data points extracted by the 3D feature extractor semantic features, instance features, It is a set of data points for 3D point cloud data.

[0055] In S105 of this embodiment, the two-dimensional pseudo-tag pair is formed by the two-dimensional semantic pseudo-tags and two-dimensional instance pseudo-tags obtained in S102. To supervise the signal, a joint loss function is designed to drive the robust transfer of knowledge from two-dimensional pseudo-labels to three-dimensional space.

[0056] For example, a joint loss function based on the alignment of a two-dimensional semantic feature map, a two-dimensional instance embedding map, and a two-dimensional pseudo-label rendered by a differentiable three-dimensional Gaussian field can be designed, and its expression can be:

[0057] in, For semantic alignment loss function, for function, The geometric regularization loss function is... These are the weight coefficients for each loss term. Specifically, the semantic alignment loss function... It can be obtained by directly applying cross-entropy loss to the rendered 2D semantic feature map and 2D semantic pseudo-labels; function The cross-entropy loss function can be obtained by treating each instance as an independent class; while the geometric regularization loss function can be constructed by introducing spatial smoothness constraints. For example, spatial smoothness constraints... To encourage smooth feature changes in non-boundary regions, its expression can be:

[0058] in, V Number of viewpoints, viewpoint index v =1,2,..., V , Projecting a three-dimensional Gaussian field onto the viewpoint v The 2D semantic feature map obtained after post-rendering is the first k Feature vector of each pixel Projecting a three-dimensional Gaussian field onto the viewpoint v The instance embedding vector of the k-th pixel in the post-rendered 2D instance embedding graph. , They are respectively , gradient, This is the boundary mask value, which is 1 at instance boundaries and 0 at non-boundary locations.

[0059] In this embodiment, based on the constructed joint loss function, the parameters of the 3D feature extractor are iteratively optimized through backpropagation to obtain the optimized 3D feature extractor. The optimized 3D feature extractor has learned semantics and instance labels from the prior of the 2D visual model, and therefore has powerful semantic segmentation and instance segmentation capabilities.

[0060] In S106 of this embodiment, the optimized three-dimensional feature extractor performs panoramic point cloud segmentation on the point cloud data to be segmented, and obtains high-quality point cloud segmentation results.

[0061] Example 2 Please see Figure 3 , Figure 3 A second flowchart illustrating a differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation method provided as an embodiment of the present invention, the method comprising: S201. Acquire the 3D point cloud data of the target scene and the multi-view 2D image data registered with the 3D point cloud data; S202. Extract the two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from the multi-view two-dimensional image data respectively; S203. Construct the shape parameters of a three-dimensional Gaussian based on the local geometric features of each data point in the three-dimensional point cloud data. Extract the semantic features and instance features of each data point as the attribute parameters of the three-dimensional Gaussian using a three-dimensional feature extractor. Construct a differentiable three-dimensional Gaussian field based on the shape parameters and attribute parameters. S204. Project the differentiable three-dimensional Gaussian field onto each viewpoint of the multi-view two-dimensional image data to obtain the rendered two-dimensional semantic feature map and the two-dimensional instance embedding map. S205. Construct a semantic consistency loss function between the two-dimensional semantic feature map and the two-dimensional semantic pseudo-label; S206. Construct an instance contrast learning loss function between the two-dimensional instance embedding graph and the two-dimensional instance pseudo-labels; S207. Based on the semantic consistency loss function and the instance comparison learning loss function, a joint loss function is constructed using a weighted approach. S208. Optimize the parameters of the 3D feature extractor through backpropagation to obtain the optimized 3D feature extractor; S209. Extract the semantic category and instance embedding vector of each data point in the point cloud to be segmented using the optimized 3D feature extractor; S210. Based on the semantic category corresponding to each data point, select the data points belonging to the instantiation semantic category, and cluster the instance embedding vectors corresponding to each selected data point to obtain the panoramic segmentation result.

[0062] This invention addresses the issue of inconsistent instance identifiers across different viewpoints for 2D instance pseudo-labels by proposing an intra-view contrastive learning strategy. Within a single rendered 2D instance embedding map, a contrastive relationship between pixel pairs is constructed based on their corresponding 2D instance pseudo-labels: pixel pairs belonging to the same 2D instance pseudo-label are defined as positive sample pairs, and pixel pairs belonging to different 2D instance pseudo-labels are defined as negative sample pairs. The goal is to maximize the similarity between positive sample pairs while minimizing the similarity between negative sample pairs. An instance contrastive learning loss function is constructed, and by optimizing this loss function, the embedding vectors of 3D points corresponding to the same physical entity are brought closer together in the feature space, while the embedding vectors of different entities are moved further apart. This cleverly avoids the problem of inconsistent instance identifiers across multiple viewpoints. Without any 3D instance annotation or complex cross-view instance matching, it can autonomously learn highly discriminative instance features, clearly and accurately separating independent object instances in a scene.

[0063] S201~S204 of this embodiment can be referred to S101~S104 of Embodiment 1, and will not be repeated here.

[0064] In S205 of this embodiment, a semantic consistency loss function is constructed to ensure that the predicted semantics of the 3D model under different perspectives remain consistent with the 2D visual prior.

[0065] Specifically, the two-dimensional semantic feature map obtained from the rendering is calculated. The two-dimensional semantic pseudo-labels generated by S102 The semantic consistency loss function is obtained by using pixel-wise cross-entropy loss. .

[0066] In S206 of this embodiment, in order to solve the problem of inconsistent supervisory signal conflicts caused by inconsistent identifiers of two-dimensional instance pseudo-labels across different viewpoints, an in-view contrastive learning strategy is proposed to construct an instance contrastive learning loss function.

[0067] In some possible embodiments, an instance contrastive learning loss function is constructed between the two-dimensional instance embedding graph and the two-dimensional instance pseudo-labels, including: In a single 2D instance embedding graph, two pixels belonging to the same 2D instance pseudo-label are defined as a positive sample pair, and two pixels belonging to different 2D instance pseudo-labels are defined as a negative sample pair. Calculate the similarity between the instance embedding vectors corresponding to any two pixels in a two-dimensional instance embedding graph; To maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs, an instance contrastive learning loss function is constructed.

[0068] Specifically, such as Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the principle of the in-view instance comparison learning strategy provided in an embodiment of the present invention. It shows the embedding of instances within a single rendered 2D image. Internally, based on its corresponding two-dimensional instance pseudo-tag To construct the comparison relationship of pixel pairs: pixel pairs belonging to the same instance mask (with the same instance ID) are defined as positive sample pairs, for example... Figure 4 Two pixels belonging to the same street lamp are considered a positive sample pair; pixel pairs belonging to different instance masks are defined as negative sample pairs, for example... Figure 4 A pixel belonging to a street lamp and a pixel belonging to a road sign are considered negative sample pairs. Based on this, an optimized contrastive loss function is used to ensure that, in the feature space, the embedding vectors of 3D points corresponding to the same physical entity (positive sample pairs) are close to each other, while the embedding vectors of different entities (negative sample pairs) are far apart. This strategy allows the 3D feature extractor to autonomously learn highly discriminative instance representations without complex cross-viewpoint instance matching.

[0069] In some possible embodiments, the expression for the instance contrastive learning loss function is:

[0070] in, , Embed any two pixels in the two-dimensional instance graph. 、 The corresponding instance embedding vector, For example, the set of pixels in the embedded image. For the positive sample set, For all sample sets, The similarity measurement function is defined as follows: the positive sample set is the set of all positive sample pairs, and the all sample set is the set that includes all positive and negative sample pairs.

[0071] Embodiments of the present invention employ an improved dense contrastive loss calculation instance contrastive learning loss function. The core function of this loss function is to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs in the feature space. In this way, the network does not need to explicitly align instance IDs from different perspectives; instead, it implicitly learns to map all 3D points of the same physical instance into compact clusters in the feature space, and to map points of different instances into mutually distant regions. This is a data-driven, self-supervised instance feature learning mechanism.

[0072] In S207 of this embodiment, the semantic consistency loss function and the instance contrast learning loss function are weighted and summed to construct a joint loss function.

[0073]

[0074] in, Let the semantic consistency loss function be... To compare and contrast the learning loss function, This is the balance coefficient.

[0075] This joint loss function enables the joint optimization of semantic consistency and instance comparison learning.

[0076] In step S208 of this embodiment, the joint loss function relative to the parameters of the 3D feature extractor is calculated through backpropagation. The gradient of the loss function is calculated, and the parameters of the 3D feature extractor are updated based on the gradient of the loss function. This drives the continuous optimization of the properties carried by the entire three-dimensional Gaussian field, ultimately forming a three-dimensional feature extractor that is globally consistent with the multi-view two-dimensional supervision signal and self-consistent in three-dimensional space. This three-dimensional feature extractor can be directly used for semantic and instance panoramic segmentation of three-dimensional point cloud data.

[0077] Understandably, the backpropagation algorithm is used to calculate the joint loss function on the parameters of the 3D feature extractor. The gradient, and update During this process, the geometric parameters of the 3D Gaussian field (the position of data points, rotation matrix, and scale parameters) remain fixed. The focus of optimization is to enable the 3D feature extractor to learn to predict the correct attributes based on the geometric context. This process continues to iterate until the semantics and instance attributes carried by the 3D Gaussian field reach a globally optimal consistency with the multi-view 2D supervision signal.

[0078] In S209 of this embodiment, a three-dimensional feature extractor is used to directly predict the semantic category and instance embedding vector of each point in the point cloud data to be segmented.

[0079] Specifically, once the optimization process of S208 converges, the optimized 3D feature extractor possesses powerful panoramic segmentation inference capabilities. For the point cloud data to be segmented, the semantic category prediction results and instance embedding vectors for each data point in the point cloud data can be obtained directly through network forward propagation.

[0080] In S210 of this embodiment, the instance embedding vectors of points belonging to the "object" class are clustered to obtain the final point cloud panoramic segmentation result.

[0081] Specifically, firstly, based on the semantic category of each data point, data points belonging to the instantiated semantic category are selected from the prediction results of the optimized 3D feature extractor to form a subset of the point cloud. Among them, data points belonging to the instantiated semantic category can be considered as data points belonging to the "countable objects" category, while data points belonging to the "uncountable objects" category, such as the ground and sky, should be excluded.

[0082] Then, for each data point p in the point cloud subset t Corresponding high-dimensional instance embedding vector Cluster analysis is performed, and instance identifiers are automatically assigned to each data point. Finally, a complete point cloud panoramic segmentation result with semantic labels and instance identifiers is output.

[0083] For example, a density-based clustering algorithm can be used to cluster the instance embedding vectors corresponding to each data point in the point cloud subset. Since points belonging to the same instance are close to each other and points belonging to different instances are separated from each other in the optimized feature space, the density-based clustering algorithm can automatically divide them into different groups, and each group is assigned a unique instance ID.

[0084] This invention utilizes synchronously acquired 3D point cloud data and multi-view 2D image data, employing a powerful pre-trained 2D vision model to "distill" supervised knowledge from the images. Leveraging the emerging and efficient continuous scene representation of differentiable 3D Gaussian sputtering as an intermediary bridge, it seamlessly and consistently "upgrades" and solidifies the 2D knowledge into the spatial structure of the 3D point cloud through an end-to-end optimization process, forming a "knowledge enhancement" mechanism. This enables the invention to exhibit far superior adaptability and robustness compared to traditional fully supervised methods when facing new data with varying distributions, effectively addressing the problem of insufficient generalization ability in traditional models.

[0085] To verify the universality and effectiveness of this invention, experiments were conducted on several publicly available large-scale point cloud datasets with different characteristics. The experiments demonstrated that the method of this invention has the following advantages: (1) High accuracy without annotation: On the KITTI-360 dataset, this invention was trained and evaluated without using any 3D ground truth annotations, and achieved panoramic segmentation accuracy comparable to fully supervised benchmark methods (such as MinkowskiNet and KPConv methods) that require a large number of 3D annotations, which fully demonstrates the effectiveness of the "knowledge enhancement" paradigm.

[0086] (2) Excellent cross-domain robustness: To test generalization ability, a fully supervised model trained only on the KITTI-360 dataset was directly applied to the significantly different WHU-Urban3D dataset, resulting in a severe performance drop. In contrast, the method of this invention directly generates pseudo-labels on the two-dimensional images of the WHU-Urban3D dataset and runs the optimization process, achieving significantly better performance than the fully supervised cross-domain model on the WHU-Urban3D dataset, demonstrating its extraordinary robustness to domain shifts, which is crucial for practical applications.

[0087] (3) High-quality segmentation: Figure 5 , Figure 6 The images show a comparison of the qualitative effects of 3D point cloud panoramic segmentation obtained using different methods with different point cloud data provided in the embodiments of the present invention. Figure 5 , Figure 6 Different point cloud datasets were used to compare the point cloud panoramic segmentation effects of the MinkowskiNet (Minkowski Convolutional Neural Network) method, the KPConv (Kernel Point Convolution) method, the method of this invention, and the results of 3D ground truth annotation. Figure 5 , Figure 6The qualitative results show that the method of the present invention can accurately identify the correct semantics of most points in the scene, and at the same time separate a large number of independent objects with similar appearances and close arrangement (such as rows of vehicles and telephone poles) in the scene. The instance boundaries are clear, which proves the success of the strategy proposed in the present invention.

[0088] In summary, the "knowledge enhancement" framework proposed in this invention creatively utilizes differentiable 3D Gaussian sputtering as a bridge to safely and effectively transfer general knowledge from 2D visual models to 3D geometric space, achieving high-precision, highly generalized point cloud panoramic segmentation with zero 3D annotation cost. This method has broad applicability, providing a novel and practical technical solution for many fields requiring deep understanding of 3D scenes, such as autonomous driving, robot navigation, mixed reality, and smart cities.

[0089] Please see Figure 7 , Figure 7 A schematic diagram of a differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation system provided as an embodiment of the present invention, the system comprising: The data acquisition module 710 is used to acquire three-dimensional point cloud data of the target scene and multi-view two-dimensional image data registered with the three-dimensional point cloud data; The pseudo-label extraction module 720 is used to extract two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from multi-view two-dimensional image data, respectively. The Gaussian field construction module 730 is used to construct the shape parameters of a three-dimensional Gaussian field based on the local geometric features of each data point in the three-dimensional point cloud data. The semantic features and instance features of each data point are extracted by a three-dimensional feature extractor as the attribute parameters of the three-dimensional Gaussian field. A differentiable three-dimensional Gaussian field is constructed based on the shape parameters and attribute parameters. The projection rendering module 740 is used to project a differentiable three-dimensional Gaussian field onto various perspectives of multi-view two-dimensional image data to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map. The parameter optimization module 750 is used to construct a joint loss function based on the two-dimensional semantic pseudo-label, two-dimensional instance pseudo-label, two-dimensional semantic feature map and two-dimensional instance embedding map, and optimize the parameters of the three-dimensional feature extractor through backpropagation to obtain the optimized three-dimensional feature extractor. The panoramic segmentation module 760 is used to perform panoramic segmentation of the point cloud data to be segmented using an optimized 3D feature extractor, and obtain panoramic segmentation results.

[0090] The label-free point cloud panoramic segmentation system driven by differentiable Gaussian field described above and the label-free point cloud panoramic segmentation method driven by differentiable Gaussian field described above can be referred to each other.

[0091] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions stored in the memory 830 to execute the differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation method provided in the above-described method embodiments.

[0092] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0093] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the differentiable Gaussian field-driven label-free point cloud panoramic segmentation method provided in the above-described method embodiments.

[0094] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the differentiable Gaussian field-driven label-free point cloud panoramic segmentation method provided in the above-described method embodiments.

[0095] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0096] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation method, characterized in that, include: Acquire 3D point cloud data of the target scene and multi-view 2D image data registered with the 3D point cloud data; Two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels are extracted from the multi-view two-dimensional image data respectively; The shape parameters of a three-dimensional Gaussian are constructed based on the local geometric features of each data point in the three-dimensional point cloud data. The semantic features and instance features of each data point are extracted by a three-dimensional feature extractor as the attribute parameters of the three-dimensional Gaussian. A differentiable three-dimensional Gaussian field is constructed based on the shape parameters and the attribute parameters. The differentiable three-dimensional Gaussian field is projected onto each viewpoint of the multi-view two-dimensional image data to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map. Based on the two-dimensional semantic pseudo-labels, the two-dimensional instance pseudo-labels, the two-dimensional semantic feature map, and the two-dimensional instance embedding map, a joint loss function is constructed, and the parameters of the three-dimensional feature extractor are optimized through backpropagation to obtain the optimized three-dimensional feature extractor. The optimized 3D feature extractor is used to perform panoramic point cloud segmentation on the point cloud data to be segmented, and the panoramic segmentation result is obtained.

2. The label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field according to claim 1, characterized in that, The step of extracting two-dimensional semantic pseudo-labels and two-dimensional instance pseudo-labels from the multi-view two-dimensional image data includes: Initial semantic pseudo-labels are extracted from the multi-view two-dimensional image data using a pre-trained semantic segmentation model; The initial instance mask of the multi-view two-dimensional image data is extracted using a pre-trained instance segmentation model; By determining the dominant category of the initial instance mask, conflict resolution and fusion are performed on the initial semantic pseudo-label and the initial instance mask to obtain two-dimensional semantic pseudo-label and two-dimensional instance pseudo-label.

3. The label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field according to claim 1, characterized in that, The step of constructing the shape parameters of a three-dimensional Gaussian based on the local geometric features of each data point in the three-dimensional point cloud data includes: For each data point in the three-dimensional point cloud data, determine the neighborhood point set of the data point and calculate the covariance matrix of the neighborhood point set; The covariance matrix is ​​decomposed into eigenvalues ​​to obtain multiple eigenvalues ​​and multiple eigenvectors; Determine the initial rotation matrix of the three-dimensional Gaussian centered on the data point based on the multiple eigenvectors; The initial scale parameters of the three-dimensional Gaussian are determined based on the multiple eigenvalues; wherein the shape parameters include the position of the data points, the initial rotation matrix, and the initial scale parameters.

4. The label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field according to claim 1, characterized in that, The step of projecting the differentiable three-dimensional Gaussian field onto each viewpoint of the multi-view two-dimensional image data to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map includes: Based on each perspective of the multi-view two-dimensional image data, the differentiable three-dimensional Gaussian field is projected onto the two-dimensional image plane corresponding to each perspective of the multi-view two-dimensional image data through a three-dimensional Gaussian sputtering differentiable rendering pipeline, to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map.

5. The label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field according to claim 1, characterized in that, The step of constructing a joint loss function based on the two-dimensional semantic pseudo-label, the two-dimensional instance pseudo-label, the two-dimensional semantic feature map, and the two-dimensional instance embedding map includes: Construct a semantic consistency loss function between the two-dimensional semantic feature map and the two-dimensional semantic pseudo-label; Construct an instance contrast learning loss function between the two-dimensional instance embedding graph and the two-dimensional instance pseudo-labels; A joint loss function is constructed using a weighted approach based on the semantic consistency loss function and the instance contrast learning loss function.

6. The label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field according to claim 5, characterized in that, Constructing an instance contrastive learning loss function between the two-dimensional instance embedding map and the two-dimensional instance pseudo-labels includes: In a single 2D instance embedding image, two pixels belonging to the same 2D instance pseudo-label are defined as a positive sample pair, and two pixels belonging to different 2D instance pseudo-labels are defined as a negative sample pair. Calculate the similarity between the instance embedding vectors corresponding to any two pixels in the two-dimensional instance embedding graph; To maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs, an instance contrastive learning loss function is constructed.

7. The label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field according to claim 6, characterized in that, The expression for the instance contrastive learning loss function is: in, , Each of the two pixels embedded in the two-dimensional instance is an arbitrary two-pixel point. , 'Corresponding instance embedding vector, The total number of pixels in the embedded image of the degree example. For the positive sample set, For all sample sets, The similarity measurement function is defined as follows: the positive sample set is the set of all positive sample pairs, and the set of all samples is the set that includes all positive sample pairs and negative sample pairs.

8. The label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field according to claim 1, characterized in that, The step of performing panoramic point cloud segmentation on the point cloud data to be segmented using the optimized 3D feature extractor to obtain the panoramic segmentation result includes: The optimized 3D feature extractor is used to extract the semantic category and instance embedding vector of each data point in the point cloud data to be segmented. Based on the semantic category corresponding to each data point, data points belonging to the instantiation semantic category are selected, and the instance embedding vectors corresponding to each selected data point are clustered to obtain the panoramic segmentation result.

9. A differentiable Gaussian field-driven annotation-free point cloud panoramic segmentation system, characterized in that, include: The data acquisition module is used to acquire three-dimensional point cloud data of the target scene and multi-view two-dimensional image data registered with the three-dimensional point cloud data; The tag extraction module is used to extract two-dimensional semantic pseudo-tags and two-dimensional instance pseudo-tags from the multi-view two-dimensional image data, respectively. The Gaussian field construction module is used to construct the shape parameters of a three-dimensional Gaussian based on the local geometric features of each data point in the three-dimensional point cloud data, extract the semantic features and instance features of each data point as the attribute parameters of the three-dimensional Gaussian through a three-dimensional feature extractor, and construct a differentiable three-dimensional Gaussian field based on the shape parameters and the attribute parameters. The projection rendering module is used to project the differentiable three-dimensional Gaussian field onto each viewpoint of the multi-view two-dimensional image data to obtain a rendered two-dimensional semantic feature map and a two-dimensional instance embedding map. The parameter optimization module is used to construct a joint loss function based on the two-dimensional semantic pseudo-label, the two-dimensional instance pseudo-label, the two-dimensional semantic feature map and the two-dimensional instance embedding map, and optimize the parameters of the three-dimensional feature extractor through backpropagation to obtain the optimized three-dimensional feature extractor. The panoramic segmentation module is used to perform panoramic segmentation of the point cloud data to be segmented using the optimized 3D feature extractor, and obtain the panoramic segmentation result.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the label-free point cloud panoramic segmentation method driven by a differentiable Gaussian field as described in any one of claims 1 to 7.