Lightweight real-time semantic segmentation method for three-dimensional Gaussian scene
By combining the representation form of PP-LiteSeg network and 3DGS, and combining projection and optimization, semantic information segmentation and transmission from two-dimensional to three-dimensional is achieved, solving the problems of robustness and real-time requirements of three-dimensional semantic segmentation in the prior art, and achieving fast, lightweight and efficient three-dimensional Gaussian scene segmentation.
Patent Information
- Application Number
- CN202510301772.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
The existing three-dimensional semantic segmentation methods are not robust and adaptable in complex scenarios, and have high computational complexity, making it difficult to meet real-time requirements, especially on lightweight edge devices with limited resources.
A lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes is proposed. By combining the lightweight, real-time two-dimensional segmentation capabilities of the PP-LiteSeg network and the unique representation form of 3DGS, the semantic information segmentation and transmission from two-dimensional to three-dimensional is realized by combining projection and optimization.
It realizes fast, lightweight, and takes into account accuracy and real-time three-dimensional Gaussian scene segmentation. It is suitable for complex scenarios and edge devices with weak computing power performance, improving segmentation robustness and practicality.
Smart Images

Figure CN120236273A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and three-dimensional scene understanding, and in particular to a lightweight real-time semantic segmentation method for three-dimensional scenes. Background Art
[0002] With the advancement of 3D perception technology, more and more 3D reconstruction methods have been introduced, and 3D semantic segmentation has also shown important value in fields such as autonomous driving, robotics, and virtual reality. Traditional 3D segmentation methods rely on the construction of geometric models, such as geometric clustering based on point clouds or RANSAC algorithms that rely on manual features. Although they are interpretable, they are not robust and adaptable in complex scenes. In recent years, with the development of deep learning, neural network-based methods such as PointNet and VoxelNet have been introduced. They rely on a large amount of data learning and significantly improve segmentation accuracy. However, these methods have high computational complexity and are difficult to meet real-time requirements. With the advent of neural radiant fields (NeRF) and 3D Gaussian sputtering (3DGS), as new 3D representation technologies, they provide new ideas for segmentation. In particular, the 3D scene based on Gaussian function has the characteristics of "internal continuous and external discrete" in the smallest basic unit of the scene. However, NeRF has high computational cost and 3DGS has consistency problems in multi-view semantic fusion. Existing methods are difficult to balance between accuracy and efficiency, especially on resource-constrained lightweight edge devices, and a fast and efficient solution is urgently needed. Summary of the invention
[0003] In view of the shortcomings of the prior art, the present invention proposes a lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes, which makes full use of the correspondence between three-dimensional scenes and two-dimensional scenes, and integrates model-driven and digital-driven. By utilizing the lightweight and real-time two-dimensional segmentation capabilities of the PP-LiteSeg network and the unique representation of 3DGS, the segmentation and transmission of semantic information from two-dimensional to three-dimensional is achieved through a combination of projection and optimization. The purpose of the present invention is to provide a fast, lightweight, accurate and real-time three-dimensional Gaussian scene segmentation method suitable for complex scenes and edge devices with weak computing performance.
[0004] The present invention achieves this object through the following technical solutions:
[0005] A lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes, comprising the following steps:
[0006] S1: Data preparation: Obtain multi-view color images of the scene, camera pose information and initial point cloud model, and pre-process the data;
[0007] S2: Two-dimensional semantic segmentation: Input the image into the PP-LiteSeg lightweight segmentation model for semantic segmentation to obtain a pixel-level two-dimensional semantic segmentation map and semantic labels;
[0008] S3: Three-dimensional scene modeling and semantic projection: Fuse multi-view color images, camera poses, and initial point clouds for three-dimensional Gaussian reconstruction of multi-modal data to generate a set of Gaussian functions, and project the two-dimensional semantic segmentation map onto the Gaussian functions through the camera pose to initially assign semantic features to the Gaussian functions;
[0009] S4: Semantic optimization and segmentation output: Design a loss function, and optimize the semantic features of the Gaussian functions by comparing the rendered semantic map with the two-dimensional semantic segmentation map to obtain a set of three-dimensional Gaussian functions with semantic labels and perform segmentation output.
[0010] Furthermore, when the lightweight two-dimensional segmentation model PP-LiteSeg in step S2 processes the input image, it adopts a dual strategy of quantization compression and knowledge distillation, converting the floating-point operations of PP-LiteSeg into integer operations, effectively reducing memory occupancy.
[0011] Furthermore, during the distillation process, the pre-trained ResNet-50 is used as the teacher network, and high-level semantic features are transmitted through the feature imitation loss function to ensure the loss of segmentation accuracy of the compressed model.
[0012] Furthermore, in step S3, after quickly generating a three-dimensional Gaussian model from multi-modal information, the process of projecting the two-dimensional semantic segmentation map onto the Gaussian functions specifically includes:
[0013] Calculate the projected pixel coordinates of the Gaussian function in each view image through the position of the Gaussian function and the camera pose information;
[0014] Collect the semantic labels of the projected points in all views, and use the majority voting method to select the label with the largest weight as the label of this point, initially assign semantic features to the Gaussian function and convert the initial assignment result into a semantic feature vector.
[0015] Furthermore, the optimization of the semantic features of the Gaussian functions in step S4 specifically includes:
[0016] Render the semantic features of the Gaussian scene obtained in step S3 onto a two-dimensional image, compare it with the two-dimensional semantic segmentation map, and calculate the difference between the two;
[0017] In addition to using the image difference as the optimization target, introduce a smoothing constraint between adjacent Gaussian functions to ensure that Gaussian functions close in space have similar semantic features;
[0018] Considering the image reconstruction objective, semantic accuracy, and smoothness comprehensively, the semantic features of the Gaussian function are adjusted through an optimized iterative method to make them consistent under multiple perspectives.
[0019] Furthermore, the segmentation output in step S4 includes semantic segmentation and instance segmentation. The final label is generated through the optimized semantic features, and different instances of the same category are separated, specifically including:
[0020] According to the optimized semantic features, directly determine the semantic label of each Gaussian function and write it into the parameters of the Gaussian function to complete semantic segmentation;
[0021] For scenarios where instances need to be distinguished, use the position information and semantic features of the Gaussian function to separate different object instances in the same category through a clustering method.
[0022] Even further, during the clustering process, combine spatial distance and feature similarity for judgment and set clustering parameters to adapt to different object scales and densities.
[0023] Even further, the segmentation result can be selectively post-processed through optimization methods including neighborhood consistency checking and instance boundary optimization to improve the accuracy and coherence of the segmentation result.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] The present invention combines multi-modal data input, model-based methods with learning-based methods, as well as lightweight design and real-time performance, and can improve the segmentation robustness while solving the problem of inconsistent perspectives and ensuring accuracy. Specifically, the lightweight design is reflected in the combination of the efficient segmentation ability of PP-LiteSeg and the fast rendering characteristics of 3DGS, and the method can run in real time on devices with relatively low computational resource consumption; the efficient projection fusion effectively projects and integrates the semantic information in the two-dimensional space into the three-dimensional space through the preliminary assignment and optimization process of multi-perspective semantics; the real-time performance benefits from the optimized computational amount and improved rendering efficiency, making the method suitable for practical environments with high real-time requirements such as roads. In addition, for the problem of inconsistent perspectives, the present invention improves the stability of the segmentation result in cases such as missing perspectives and perspective differences through consistency optimization and smooth constraints, thereby enhancing the robustness and practicality of the method, and is particularly suitable for fields such as robot navigation and autonomous driving that require fast and efficient three-dimensional scene understanding. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is the overall flowchart of the method of the present invention;
[0027] Figure 2 is the schematic diagram of the semantic generation and output part in the present invention;
[0028] Figure 3 is the schematic diagram of the two-dimensional to three-dimensional semantic projection of the present invention;
[0029] Figure 4 is the comparison between the rendered image obtained by the multi-modal Gaussian reconstruction of the present invention and the real image. Detailed implementation manners
[0030] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0031] The present invention proposes a lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes, which makes full use of the correspondence between three-dimensional scenes and two-dimensional scenes and integrates model-driven and data-driven. By using the lightweight and real-time two-dimensional segmentation ability of the PP-LiteSeg network and the unique representation form of 3DGS, through a method combining projection and optimization, the semantic information segmentation and transmission from two-dimensional to three-dimensional are realized. A lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes provided by the present invention realizes the real-time semantic segmentation of complex scenes by collecting multi-view images and combining efficient two-dimensional segmentation and three-dimensional modeling technologies. The process is as Figure 1 shown, and the specific steps are as follows:
[0032] S1, Data preparation: Obtain cross-modal data of the original point cloud, multi-view color images and the corresponding camera poses. In a specific implementation, the initial point cloud is obtained by the SLAM (Simultaneous Localization and Mapping) method, and the adaptive downsampling result based on the voxel grid is used to complete it; the multi-view color images are collected by the camera, covering different angles of the scene to be reconstructed to ensure the integrity and accuracy of the three-dimensional Gaussian reconstruction; the camera pose parameters include the camera intrinsic matrix (including the focal length and the principal point coordinates) and the camera extrinsic parameters (including the rotation matrix and the translation vector), and these parameters are used not only for constructing the three-dimensional point cloud but also for the subsequent projection mapping from two-dimensional to three-dimensional.
[0033] In actual operation, first use the camera to take pictures of the indoor scene (such as a living room or an office) from different angles to ensure complete coverage of the perspective and capture the complete geometric information of the scene. The camera pose is obtained by the SLAM technology. The collected data is preprocessed, including image denoising and pose calibration, to improve the accuracy of subsequent segmentation.
[0034] S2, Two-dimensional semantic segmentation: In this step, the PP-LiteSeg network is used to generate a pixel-level semantic segmentation map in real time. In this step, for each input image, the PP-LiteSeg network processes it using its lightweight encoder-decoder structure. Among them, the encoder extracts image features through multiple layers of convolution, combines depthwise separable convolution technology to reduce the computational amount, and at the same time retains key object boundaries and context information; the decoder converts the feature map into a segmentation map with the same resolution as the original image through upsampling and feature fusion, and each pixel is assigned a semantic label, such as "table", "chair" or "cabinet", providing a semantic basis for subsequent 3D projection.
[0035] Input image I into PP-LiteSeg, and output a two-dimensional semantic segmentation map:
[0036]
[0037] where represents the pixel-level semantic label map, and θ represents the set of parameters of the network model.
[0038] PP-LiteSeg adopts a lightweight encoder-decoder structure to efficiently process the input image, generate a pixel-level semantic segmentation map, and ensure real-time segmentation performance; it uses depthwise separable convolution to replace standard convolution, decomposes convolution into depth convolution and point convolution, reduces the computational amount and parameter scale; introduces a lightweight feature extraction network in the encoder, uses multi-scale features to capture local details and global context of the image, and improves the segmentation accuracy; the encoder extracts image features through multiple layers of convolution, while maintaining sensitivity to object boundaries; the decoder gradually enlarges the feature map through upsampling operations and fuses multi-scale information in the encoder, and finally generates a semantic segmentation map with the same resolution as the input image.
[0039] In addition, the present invention converts the floating-point operations of PP-LiteSeg into integer operations through a dual strategy of quantization compression and knowledge distillation, effectively reducing memory occupancy. During the distillation process, a pre-trained ResNet-50 is used as the teacher network, and high-level semantic features are transmitted through a feature imitation loss function to ensure the loss of segmentation accuracy of the compressed model. The speed of the segmentation process can reach 30 frames per second on the experimental device (NVIDIA GTX 3070) in this example, meeting the real-time requirement.
[0040] S3, 3D scene modeling and semantic projection: As Figure 2 shown, this step reconstructs the scene through 3DGS and projects two-dimensional semantics.
[0041] First, using the data in the data preparation process, a fast three-dimensional Gaussian sputtering technique based on multi-modal data is adopted to generate a three-dimensional Gaussian representation of the scene, which consists of several Gaussian function structures. Each Gaussian function structure contains a three-dimensional position x i , covariance matrix Σ i , color c i , opacity α i . The set of Gaussian functions generated by reconstructing the scene through 3DGS can be expressed as
[0042] After that, calculate the projection position of each Gaussian function on the two-dimensional image through the camera pose, and correspond it to the semantic label of the pixel in the two-dimensional semantic segmentation map generated by PP-LiteSeg, and transfer it to the corresponding Gaussian function representation. Through projection Obtain the projection label where represents the image coordinates, π represents the projection function, and P k = {R k , T k , K k} represents the set of camera pose parameters, where R k ∈ SO(3) represents the rotation matrix, represents the translation matrix, represents the internal parameter matrix; x i represents the three-dimensional position information.
[0043] To initially determine the semantic features, weights are set for each projection map according to factors such as distance and perspective, and a voting mechanism is adopted to integrate the label information from various perspectives; for example, for a cross-semantic scene, the semantic label with the highest frequency after weighted average is preferentially selected, and the initial assignment result is converted into the form of a semantic feature vector and represented by a softmax distribution to provide an initial value for subsequent optimization. Specifically:[[]]
[0044] When a certain Gaussian function has projections in multiple perspectives, the set of all visible perspectives is defined as V i , collect all labels L i,k , and use the majority voting method to determine the initial label, that is, select the one with the largest weight as the label of this point. The initial semantic features are represented in the form of probability, reflecting the distribution of multi-perspective labels. Then, the preliminary semantic label of the Gaussian function with multiple perspective projections is expressed as:[[]]
[0045]
[0046] where c represents the category, is the initially determined semantic label, and δ(L i,k , c) is a judgment function, that is, if L i,k = c, then δ(Li,k , c) = 1, otherwise δ(L i,k , c) = 0.
[0047] After that, the Softmax function is used to convert the preliminary semantic labels into the preliminary semantic feature vector f i init .
[0048] S4, Semantic Optimization and Segmentation Output: As Figure 2 shown, this step optimizes the preliminary semantic features generated before and outputs the segmentation result based on the 3D model. In this step, the idea of iterative optimization to minimize the error is adopted. Using the characteristics of the Gaussian function and the corresponding relationship between 3D and 2D conversions, the semantic features of the Gaussian function are projected back to the 2D image and compared with the segmentation ground truth of PP-LiteSeg. By iterative optimization to minimize the error, the difference between the rendered image and the real image is adjusted to ensure consistency under multi-view semantics. At the same time, in order to increase the credibility of the result, a smoothing mechanism is introduced to make the semantic features of adjacent Gaussian points tend to be consistent, so as to reduce the influence brought by projection noise, error, etc. Finally, when the number of iterative optimization reaches the threshold or the iterative effect tends to be stable, according to the optimized semantic features, the semantic label of each Gaussian function is directly generated to complete semantic segmentation, or different object instances in the same category are distinguished through clustering technology, and a complete 3D Gaussian scene with semantic labels is output. The specific process is as follows:
[0049] First, use the 3DGS rendering mechanism to project the semantic features F of the Gaussian function i onto the 2D image to generate a rendered semantic map, denoted as where w i is the Gaussian weight based on Σ i ; After that, compare it with the 2D semantic segmentation map of PP-LiteSeg, and combine the image reconstruction target and the neighborhood smoothing requirement to establish a loss function
[0050] L total = L image + λ1L semantic + λ2L smooth .
[0051] where L image , L semantic , L smooth represent the reconstruction loss, semantic consistency loss, and smoothing loss respectively, and λ1, λ2 represent the weight coefficients.
[0052] The optimization process comprehensively considers three aspects: First, ensure that the rendering result is as consistent as possible with the 2D segmentation map; second, keep the color of the Gaussian function as realistic as possible; third, make the semantic features of adjacent Gaussian functions as similar as possible to avoid isolated noise. After optimization, according to the loss function, the semantic label of each optimized Gaussian function is output:
[0053]
[0054] After that, the Softmax function is used to convert the semantic label into a semantic feature vector f i 。
[0055] As Figure 3 shown, this step obtains a new 3D initialization model representation composed of Gaussian functions, and the final set of Gaussian functions is represented as Each Gaussian function not only contains information such as spatial position and color, but also has an attached semantic label for each to describe the object category in the scene, which is further used for subsequent scene understanding tasks. The comparison between the rendered image and the real image obtained by the multi-modal Gaussian reconstruction of the present invention is as Figure 4 shown, where Figure 4 (a) is the rendered image obtained by the multi-modal Gaussian reconstruction of the present invention, Figure 4 (b) is the real image.
[0056] The final segmentation output includes semantic segmentation and instance segmentation. The final label is generated through the optimized semantic features, and different instances of the same category are separated. Specifically, it includes:
[0057] According to the optimized semantic features, directly determine the semantic label of each Gaussian function and write it into the parameters of the Gaussian function to complete semantic segmentation.
[0058] For scenes that require instance segmentation, different instances are further separated by a clustering method (such as DBSCAN) in combination with position information. During the clustering process, judgments are made by combining spatial distance and feature similarity, and clustering parameters are set to adapt to different object scales and densities.
[0059] Optionally, post-processing can be performed on the segmentation result. Through optimization methods including neighborhood consistency checking and instance boundary optimization, the accuracy and coherence of the segmentation result are improved
[0060] In summary, the proposed method can achieve segmentation in test scenarios. More importantly, it meets the requirements of lightweight and real-time performance, and has broad application prospects. For example, in robot navigation applications, the method can segment the travel route, obstacles, and invalid scenarios, providing more accurate, rapid, and diverse information support for path planning. In virtual reality and augmented reality, users can utilize the segmentation results to perform real-time editing on objects in the scenario. Additionally, the advantages of the present invention lie in its lightweight design and high processing efficiency, enabling it to run on edge devices (such as embedded GPUs) while maintaining high segmentation accuracy and robustness.
[0061] The present invention has been described in detail through the above embodiments. However, the above content is only an exemplary embodiment of the present invention and should not be considered as limiting the scope of implementation of the present invention. The protection scope of the present invention is defined by the claims. Any use of the technical solutions described in the present invention, or any similar technical solutions designed by those skilled in the art under the inspiration of the technical solutions of the present invention, within the essence and protection scope of the present invention, to achieve the above technical effects, or any equivalent changes and improvements made to the scope of the application, shall still fall within the scope of patent protection of the present invention. It should be noted that, for the sake of clear description, the description of the present invention omits the expressions of some components and processes that are not directly and significantly related to the protection scope of the present invention but are known to those skilled in the art.
Claims
1. A lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes, characterized in that: The following steps are involved: S1: Data preparation: Obtain multi-view color images of the scene, camera pose information and initial point cloud model, and pre-process the data; S2: 2D semantic segmentation: Input the image into the PP-LiteSeg lightweight segmentation model for semantic segmentation to obtain a pixel-level 2D semantic segmentation map and semantic labels; S3: 3D scene modeling and semantic projection: Fusion of multi-view color images, camera poses and initial point clouds for 3D Gaussian reconstruction of multimodal data, generation of a set of Gaussian functions, and projection of the 2D semantic segmentation map onto the Gaussian function through the camera pose, preliminarily assigning semantic features to the Gaussian function; S4: Semantic optimization and segmentation output: Design a loss function, optimize the semantic features of the Gaussian function by comparing the rendered semantic map with the two-dimensional semantic segmentation map, obtain a set of three-dimensional Gaussian functions with semantic labels and perform segmentation output.
2. According to claim 1, a lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes is characterized in that: The lightweight two-dimensional segmentation model PP-LiteSeg in step S2 adopts a dual strategy of quantization compression and knowledge distillation when processing the input image, converting the floating-point operations of PP-LiteSeg into integer operations, effectively reducing memory usage.
3. The lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes according to claim 2, characterized in that: During the distillation process, a pre-trained ResNet-50 is used as the teacher network, and high-level semantic features are transferred through the feature imitation loss function to ensure the loss of model segmentation accuracy after compression.
4. The lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes according to claim 1, characterized in that: In step S3, after quickly generating a three-dimensional Gaussian model from multimodal information, the process of projecting the two-dimensional semantic segmentation map onto the Gaussian function specifically includes: Calculate the projection pixel coordinates of the Gaussian function in each perspective image through the position of the Gaussian function and the camera pose information; The semantic labels of the projection points in all viewing angles are collected, and the label with the largest weight is selected as the label of the point by majority voting. The semantic features of the Gaussian function are preliminarily assigned and the preliminary assignment results are converted into semantic feature vectors.
5. The lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes according to claim 1, characterized in that: The optimization of the semantic features of the Gaussian function in step S4 specifically includes: Rendering the semantic features of the Gaussian scene obtained in step S3 to a two-dimensional image, and comparing it with the two-dimensional semantic segmentation map, and calculating the difference between the two; In addition to using image difference as the optimization target, a smoothness constraint is introduced between adjacent Gaussian functions to ensure that Gaussian functions close in space have similar semantic features. Taking into account the image reconstruction goal, semantic accuracy and smoothness, the semantic features of the Gaussian function are adjusted through optimization iteration to make them consistent under multiple perspectives.
6. The lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes according to claim 1, characterized in that: The segmentation output of step S4 includes semantic segmentation and instance segmentation. The final label is generated through the optimized semantic features, and different instances of the same category are separated, which specifically includes: According to the optimized semantic features, the semantic label of each Gaussian function is directly determined and written into the parameters of the Gaussian function to complete semantic segmentation; For scenarios where instances need to be distinguished, the location information and semantic features of the Gaussian function are used to separate different object instances in the same category through clustering methods.
7. The lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes according to claim 6, characterized in that: In the clustering process, spatial distance and feature similarity are combined for judgment, and clustering parameters are set to adapt to different object scales and densities.
8. The lightweight real-time semantic segmentation method for three-dimensional Gaussian scenes according to claim 6, characterized in that: The segmentation results can be optionally post-processed to improve their accuracy and consistency through optimization methods including neighborhood consistency check and instance boundary optimization.
Citation Information
Cited By
Three-dimensional Gaussian model construction method based on semantic segmentation
CN120451360A
A method for constructing a 3D Gaussian model based on semantic segmentation
CN120451360B
Three-dimensional reconstruction and semantic map generation method and system based on Gaussian sputtering
CN120543794A
Three-dimensional reconstruction and semantic map generation method and system based on Gaussian sputtering
CN120543794B
Image reconstruction method and device based on 3D Gaussian model, electronic equipment and medium
CN120823325A