Three-dimensional scene segmentation method based on 3D Gaussian Splitting
Through the 3D Gaussian Splatting method based on 3D Gaussian Splatting, combined with Gaussian hybrid algorithm and PinPrompt algorithm, the problem of insufficient volume integrity and accuracy when segmenting objects in the traditional method is solved, and three-dimensional segmentation with high precision and high integrity is achieved.
Patent Information
- Application Number
- CN202510424198.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-11
AI Technical Summary
The existing three-dimensional scene segmentation algorithms are difficult to improve segmentation accuracy while ensuring volume integrity. Especially in a small number of lens scenarios, traditional methods cannot effectively solve the problem of volume integrity and accuracy when segmenting objects.
The three-dimensional scene segmentation method based on 3D Gaussian Splatting is adopted, and the 3D Gaussian Splatting scene preparation and training is prepared and trained, and the three-dimensional segmentation is performed by combining multi-views. The Gaussian hybrid algorithm and the PinPrompt algorithm are used to project and update the segmentation mask. Transmittance is introduced as the confidence factor, and efficient rendering from different perspectives is achieved using a cyclic anchoring scheme.
The accuracy and integrity of three-dimensional segmentation are improved, especially in a small number of lens scenarios, and the boundary distinction ability is enhanced and high-quality three-dimensional segmentation is achieved.
Smart Images

Figure CN120298697A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional scene segmentation, and in particular to a three-dimensional scene segmentation method based on 3D Gaussian Splatting. Background Art
[0002] With the continuous development of deep learning technology, the application scope of three-dimensional scene reconstruction technology has become increasingly extensive, and the generation effect and application scenarios have also been greatly improved, giving rise to numerous downstream tasks in the 3D vision field. These tasks usually involve local operations on 3D scenes, which makes it crucial to query and index each position in the reconstructed scene. 3D segmentation is a basic tool for dividing complex scenes into different regions, which can effectively classify and index points and provide necessary support for subsequent tasks. Through 3D segmentation, different objects and regions in the 3D scene can be accurately separated, improving the efficiency and accuracy of subsequent tasks.
[0003] Recently, the state-of-the-art in radiance fields, 3D Gaussian Splatting, has received a great deal of attention, with its reconstruction quality and real-time rendering efficiency being impressive. Our 3D segmentation technology is also developed based on 3D Gaussian Splatting. Current mainstream methods usually render the scene as a 2D image and then use a 2D segmentation model, such as the Segment Anything Model (SAM), to parse the segmentation semantics, and then refine it back to the 3D scene. Therefore, this process usually involves two key steps: first, propagating initial cues across multiple views, and second, aggregating 2D segmentation masks from these views to create a comprehensive 3D segmentation. Although feasible, achieving an effective upgrade from low dimensions to high dimensions requires meeting two main objectives: the first is volume integrity, where each element in the scene volume must be assigned a clear segmentation value. The second objective is accuracy, and these values should accurately reflect the semantics of the scene. By adhering to these objectives, the upgraded 3D segmentation semantics can accurately and comprehensively describe the entire scene, supporting more complex downstream tasks.
[0004] Existing 3D scene segmentation algorithms are mainly divided into two categories: (1) Methods based on neural network optimization: For each training view, a neural network is used to optimize the 3D mask; this method cannot guarantee volume integrity. On the one hand, their use of cues is limited to basic self-cues or direct view warping, without explicitly addressing how cues interact across different views. The lack of timely interaction leads to inconsistencies in target objects, resulting in unsatisfactory performance. On the other hand, the aggregation process from 2D masks to 3D segmentation is crucial for volume integrity and accuracy. SA3D proposes a simple paradigm, which is supervised at the pixel level through differentiable rendering, and updates the 3D segmentation using gradient descent. (2) Methods based on mapping projection: GaussianEditor uses Gaussian-centered views and calculates the 3D segmentation of each Gaussian based on all 2D segmentation values of the rays passing through that point. This method meets the requirements of volume integrity, but the segmentation accuracy is poor, resulting in blurred object edges and backgrounds. Summary of the Invention
[0005] The object of the present invention is to overcome the deficiencies of the existing technologies, and propose a 3D scene segmentation method based on 3D Gaussian Splatting, which solves the problem that the traditional methods cannot balance volume integrity and accuracy when segmenting objects.
[0006] To achieve the above object, the technical solution provided by the present invention is: A 3D scene segmentation method based on 3D Gaussian Splatting, including the following steps:
[0007] 1) 3D Gaussian Splatting scene preparation: 3D Gaussian Splatting is a 3D scene storage framework, the Chinese name of this storage framework is 3D Gaussian sputtering, which is an explicit storage structure; the colmap technology is used to perform sparse 3D reconstruction on the multi-views of the scene to obtain sparse sample points; the training of 3D Gaussian Splatting requires the above-obtained sparse sample points as input, and at the same time requires multi-views and their camera poses, and optimizes the entire scene by rendering different views;
[0008] 2) Use the trained 3D Gaussian Splatting as the base framework, and use multi-view as a bridge to perform 3D segmentation on the target object in the scene: Render a 2D image of a certain perspective of the scene, which contains the target object, and then use the 2D segmentation large model SAM to perform 2D segmentation on the 2D image to obtain the 2D mask of the target object. After segmentation, project the 2D mask result back onto the 3D Gaussian Splatting. The projection process is implemented using the Gaussian mixture algorithm. Specifically, use the rendering function of 3D Gaussian Splatting to inverse render the label for each Gaussian. The 2D mask ∈ [0,1]. Traverse all Gaussians in the 2D image using the 2D mask, and assign the semantic segmentation score as the mask mask * T and the penetration rate weight T to update the result of the 3D segmentation field and achieve 3D segmentation. Among them, the 2D prompt for the first rendered view is given manually, and the 2D prompts for subsequent perspectives are automatically generated by the PinPrompt algorithm. The PinPrompt algorithm uses a loop anchoring scheme to anchor the 2D prompt to the 3D scene, thus achieving efficient rendering from different perspectives;
[0009] 3) Step 2) will be repeated multiple times until all 2D images containing the target object from all perspectives have been processed once, and finally a 3D segmentation field is obtained as the final result.
[0010] Furthermore, in step 1), 3D Gaussian splatting is a high-quality unstructured radiance field representation that uses a set of 3D Gaussian points to represent the scene, where N represents the number of Gaussian points, and g i represents the i-th Gaussian point. Each Gaussian point is associated with learnable parameters for representing its 3D coordinates, scale, rotation angle, density, and color;
[0011] The rendering of 3D Gaussian Splatting uses the α - blending technique to project all Gaussian points onto the image plane along a ray to obtain the rendered per-pixel color. The α - blending technique can also be applied to render other attributes, unifying the representation of the α - blending technique for color and other attributes, so that for any attribute The rendering process for a given pixel p is as follows:
[0012]
[0013] where α i represents the density of the Gaussian point g i h i represents the attribute of the Gaussian point g i , represents transparency, j represents the j-th Gaussian on this ray, Represents the set of points on the light ray emitted from pixel p, Represents the set of points on the light ray emitted from pixel p;
[0014] The goal is to segment the object of interest to the user in the 3D scene, that is, to label all Gaussian points related to the object as 1 and other background points as 0, and assign an additional segmentation score s to each Gaussian point i , s i ∈[0,1], and these scores can be directly rendered into a 2D segmentation mask, that is, a 2D mask, for further supervision or evaluation, denoted as Represents the set of segmentation scores, p represents the pixel, and at the same time
[0015] Furthermore, in step 2), the specific steps of the Gaussian mixture algorithm are as follows:
[0016] First, the concepts of view-centered modeling and Gaussian-centered modeling are defined; view-centered modeling refers to the idea of modeling the pixel values of a 2D mask, which approximates the rendered view to the provided 2D mask. SA3D-GS is a view-centered modeling method that renders the 2D mask obtained by formula (1) and optimizes it through backpropagation; GaussianEditor adopts Gaussian-centered modeling and uses a more intuitive modeling formula; for each Gaussian point, the segmentation score s i Is calculated as the average of the different view rays of the 2D segmentation passing through this point;
[0017] The influence of occlusion is considered during the weighting process. From a practical point of view, among the rays emitted from different viewpoints, unoccluded rays should be given priority because the segmentation scores corresponding to occluded rays cannot be used as valid evidence. Therefore, the transmittance T is introduced as a confidence factor during the segmentation score weighting. Finally, the Gaussian mixture algorithm can be expressed as:
[0018]
[0019] In the formula, T i p Represents the transparency, Represents the 2D mask, Represents the set of Gaussian points where the ray intersecting the pixel p emits.
[0020] Furthermore, in step 2), the specific implementation of the PinPrompt algorithm is as follows:
[0021] Given a set of multi-view images Captured from a set of views v kDenote the k-th training view, N v Denote the total number of captured images, x k Denote the k-th captured image, and the goal is to generate reliable prompts for other views based only on the initial prompt provided by the user Assign a segmentation score s to each Gaussian point i , p n Denote the prompt on the n-th image, N init Denote the number of initial images. By developing a prompt propagation strategy and expanding the initial prompt, reliable prompts can be generated for other views, facilitating the calculation of 3D segmentation, as shown in Equation (2);
[0022] A new prompt sampling and propagation strategy, called the PinPrompt algorithm, is proposed to promote accurate prompt transmission across different perspectives. The PinPrompt algorithm uses a recycling scheme to anchor 2D prompts to the 3D scene, enabling efficient rendering from different viewpoints. To ensure the validity of the anchor points in the new view, the few-shot property of the Gaussian mixture algorithm is used to filter and select the anchor points. PinPrompt promotes the co-refinement between 3D segmentation and fast propagation;
[0023] To fix the 2D prompt to the 3D scene, the Gaussian points on the object surface are associated with the attribute a i ∈{0,1}, Denote the Gaussian anchor domain, a i Denote the attribute value of the i-th Gaussian point. A total of N Gaussian points are marked, indicating the existence of point prompts. If a certain Gaussian point g i is marked as a prompt, its attribute a i is assigned a value of 1, which is expressed as follows: Given the prompt P, determine the first Gaussian point intersected by the ray emitted from each pixel p in P, and mark its attribute a i as 1. Formally:
[0024]
[0025] where d p (i) represents the distance between i and p, Denote the set of distances of Gaussian points along the ray emitted from p, Denote the set of Gaussian points intersected by the ray emitted from the indexed pixel p;
[0026] In Equation (3), all prompts require a pinning-back process, and the final a i depends on the union of the flags to ensure consistency. After determining , the Gaussian anchor domain can be effectively constructed. Similarly, the set of Gaussian segmentation domains S can be interpreted as a Gaussian segmentation field;
[0027] Rendering attributes in 3D Gaussian Splatting, for the convenience of using the Gaussian segmentation domain S and the Gaussian anchor domain Define the rendering as follows: Given a set of attributes and a view v k , calculate the rendering mask according to formula (1) denotes the set of pixels on view v k where
[0028] Derive a 2D mask with high confidence through the PinPrompt algorithm, and establish a prompt library containing positive anchors and negative anchors Initially, integrate the user prompt P init into the Gaussian anchor domain, thus generating the initial anchor through formula (3) The positive anchor is initialized as while the negative anchor is set as an empty set;
[0029] Perform the following iterative steps on each view for training:
[0030] For the given view v k , derive the positive and negative anchor masks from the positive and negative anchors using formula (1) and That is and F is the derivation function, and the candidate positive and negative prompts and are extracted from the anchor masks;
[0031] Subsequently, render the segmentation mask which predicts the segmentation of the current view; Based on S k , apply a confidence filtering process to discard incorrect positive class prompts; Use the edge area of determined by the Canny operator to filter the negative prompts represents the negation result of S k , that is, take the negation on the original S k , finally, apply the farthest point sampling algorithm to the positive and negative prompt sets and to generate the final positive and negative prompt sets for SAM and Among them, for the initial view, since there are no candidate sampling prompts, the above steps will be ignored;
[0032] Next, use SAM to segment the image x k with the following sampling prompts:
[0033]
[0034] In the formula, represents the predicted segmentation mask, and represent positive and negative 2D cues. Subsequently, is non-projected, and a new segmentation score is determined through Equation (2). This score can serve as a reliable metric for evaluating new view cues. Additionally, the farthest point sampling algorithm is used to sample n from the segmentation mask supp supplementary cues and n supp represents the sampled points obtained by sampling from and is divided into positive and negative cue point sets and and the supplementary cues are fixed to the 3D scene through Equation (3) to obtain and represents the Gaussian anchor domain of positive class points, represents the Gaussian anchor domain of negative class points. Among them, in order to reduce the accumulation of false cues, only the anchors in the previous views in the cue library are retained, and it is updated according to the first-in, first-out method, while retaining the initial anchor
[0035]
[0036] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0037] The core idea of the present invention is to make full use of the connection between different views during the step-by-step segmentation process, which involves two complementary aspects: First, the goal is to enhance the algorithm for aggregating 2D segmentation masks into 3D segmentation, especially to improve accuracy and integrity in the few-shot scenario. This few-shot function is crucial for providing cues for new views because the cues rely on segmentation predictions from previous views. Therefore, an accurate estimate can serve as a faithful guide for new view cues, thereby potentially improving segmentation accuracy. Second, by retaining the reliable cues used in previous views, new views can avoid cueing inconsistent objects and improve mask quality. On this basis, a simple and effective boosting algorithm based on Gaussian central perspective is proposed, which not only ensures the integrity of 3D segmentation but also strengthens the boundary discrimination, achieving accurate 3D segmentation. In addition, a cue propagation strategy is introduced, which effectively utilizes reliable cues from adjacent views by restoring 2D cues to the scene, thereby achieving accurate segmentation.
[0038] The present invention introduces a Gaussian mixture algorithm for 3D segmentation of volume perception and proposes PinPrompt for accurate prompt selection and refinement. Based on the assumption that most of the prompt anchors in SAM prediction are reliable, we apply an aggressive (Farthest Point Sampling algorithm) FPS algorithm to maximize the utility of the current prompt and encourage prompt selection at the boundaries of objects of interest. Based on the few-shot robustness of the Gaussian mixture algorithm, we use the rendered segmentation masks to filter the existing prompts to ensure their reliability when propagated to other views. This mutually synergistic refinement process ensures the effective exploration of cross-view information, thereby improving the accuracy of 3D segmentation prediction. Brief Description of the Drawings
[0039] Figure 1 It is a framework diagram of the method of the present invention. Detailed Embodiment
[0040] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0041] As Figure 1 shown, this embodiment discloses a three-dimensional scene segmentation method based on 3D Gaussian Splatting, and the specific situation is as follows:
[0042] 1) 3D Gaussian Splatting scene preparation: 3D Gaussian Splatting is a 3D scene storage framework, and its Chinese name is 3D Gaussian splattering, which is an explicit storage structure; the colmap technology is used to perform sparse three-dimensional reconstruction on the multi-views of the scene to obtain sparse sample points; the training of 3D Gaussian Splatting requires the above-obtained sparse sample points as input, and at the same time requires multi-views and their camera poses, and optimizes the entire scene by rendering different views;
[0043] 3D Gaussian splattering is a high-quality unstructured radiance field representation that uses a set of 3D Gaussian points to represent the scene, where N represents the number of Gaussian points, and g i represents the i-th Gaussian point, and each Gaussian point is associated with learnable parameters for representing its three-dimensional coordinates, scale, rotation angle, density, and color;
[0044] The rendering of 3D Gaussian Splatting adopts the α-blending technology, projects all Gaussian points onto the image plane along a ray to obtain the rendered per-pixel color, and the α-blending technology can also be applied to render other attributes, unifying the representation of the α-blending technology for colors and other attributes, so that for any attribute The rendering process for a given pixel p is:
[0045]
[0046] where α i represents the density at the Gaussian point g i and h i represents the property at the Gaussian point g i . represents the transparency, j represents the j-th Gaussian on the ray, which represents the set of points on the ray emitted from pixel p, which represents the set of points on the ray emitted from pixel p;
[0047] In the present invention, our goal is to segment the object of interest in a 3D scene, that is, to label all Gaussian points related to the object as 1 and other background points as 0. Specifically, an additional segmentation score s i is assigned to each Gaussian point. s i ∈ [0, 1]. These scores can be directly rendered into a 2D segmentation mask, that is, a 2D mask, for further supervision or evaluation, denoted as represents the set of segmentation scores, p represents the pixel, and at the same time
[0048] 2) Use the trained 3D Gaussian Splatting as the base framework and use multi-views as a bridge to perform 3D segmentation on the target object in the scene: Render a 2D image of a certain perspective of the scene, which includes the target object. Then use the 2D segmentation large model SAM to perform 2D segmentation on the 2D image to obtain the 2D mask of the target object. After segmentation, project the 2D mask result back onto 3D Gaussian Splatting. The projection process is implemented using the Gaussian mixture algorithm. Specifically, use the rendering function of 3D Gaussian Splatting to inverse-render the label for each Gaussian. The 2D mask ∈ [0, 1]. Traverse all Gaussians in the 2D image using the 2D mask and assign the semantic segmentation score as mask * T and the penetration rate weight T to update the result of the 3D segmentation field to achieve 3D segmentation; where the 2D prompt for the first rendered view is given manually, and the 2D prompts for subsequent perspectives are automatically generated by the PinPrompt algorithm; The PinPrompt algorithm uses a loop anchoring scheme to anchor the 2D prompt to the 3D scene, thereby achieving efficient rendering from different perspectives;
[0049] The specific steps of the Gaussian mixture algorithm are as follows:
[0050] First, the concepts of view - centered modeling and Gaussian - centered modeling are defined; view - centered modeling refers to the idea of modeling the pixel values of a 2D mask, which approximates the rendered view to the provided 2D mask. SA3D - GS is a view - centered modeling method that renders the 2D mask obtained from formula (1) and optimizes it through backpropagation; compared with SA3D - GS, GaussianEditor adopts Gaussian - centered modeling and uses a more intuitive modeling formula; for each Gaussian point, the segmentation score s i is calculated as the average of the different view rays of the 2D segmentation passing through that point;
[0051] The present invention takes into account the influence of occlusion during the weighting process, rather than simply weighting the segmentation values according to the number of passing rays. From a practical perspective, among the rays emitted from different viewpoints, unoccluded rays should be given priority because the segmentation scores corresponding to occluded rays cannot be used as valid evidence. Therefore, the present invention introduces the transmittance T as a confidence factor during the segmentation score weighting. Finally, the Gaussian mixture algorithm can be expressed as:
[0052]
[0053] In the formula, T i p represents transparency, represents the 2D mask, represents the set of Gaussian points intersected by the rays emitted from the index pixel p.
[0054] Although simply combining the transmittance as a weighting factor, the Gaussian mixture method appropriately aggregates the high - confidence segmentations from different viewpoints. This transmittance - based weighting method achieves high - precision 3D volume segmentation by fully utilizing the reliable regions from different perspectives, thus reducing the potential ambiguity in the segmentation score estimation.
[0055] The specific implementation of the PinPrompt algorithm is as follows:
[0056] Given a set of multi - view images captured from a set of views v k represents the k - th training view, N v represents the total number of captured images, x k represents the k - th captured image. The goal is to assign the segmentation score s to each Gaussian point only based on the initial prompt i provided by the user, p n represents the prompt on the n - th image, N initDenote the initial number of images by developing a prompt propagation strategy and expanding the initial prompt to generate reliable prompts for other views, thus facilitating the calculation of 3D segmentation, as shown in Equation (2);
[0057] A new prompt sampling and propagation strategy, called the PinPrompt algorithm, is proposed to promote accurate prompt transmission across different viewpoints. The PinPrompt algorithm utilizes a recycling scheme to anchor 2D prompts to the 3D scene, enabling efficient rendering from different viewpoints. To ensure the validity of the anchor points in the new view, the few-shot property of the Gaussian mixture algorithm is used to filter and select the anchor points. PinPrompt promotes the co-refinement between 3D segmentation and fast propagation;
[0058] To fix the 2D prompts to the 3D scene, the Gaussian points on the object surface are associated with the attribute a i ∈{0,1}, denote the Gaussian anchor domain, and a i denote the attribute value of the i-th Gaussian point. A total of N Gaussian points are marked, indicating the existence of point prompts. Specifically, if a certain Gaussian point g i is marked as a prompt, its attribute a i is assigned the value 1, which is expressed as follows: Given a prompt P, determine the first Gaussian point intersected by the ray emitted from each pixel p in P and mark its attribute a i as 1. Formally:
[0059]
[0060] where d p (i) represents the distance between i and p, represents the set of distances of Gaussian points along the ray emitted from p, represents the set of Gaussian points intersected by the ray emitted from the indexed pixel p;
[0061] In Equation (3), all prompts require a pinning-back process, and the final a i depends on the union of the flags to ensure consistency. After determining , the Gaussian anchor domain can be effectively constructed. Similarly, the set of Gaussian segmentation domains S can be interpreted as a Gaussian segmentation field;
[0062] Render the attributes in 3D Gaussian Splatting. To facilitate the use of the Gaussian segmentation domain S and the Gaussian anchor domain define the rendering as follows: Given a set of attributes and a view v k , calculate the rendering mask according to Equation (1) Denote the view as v k The set of pixels on
[0063] Derive a 2D mask with high confidence through the PinPrompt algorithm, and establish a prompt library containing positive anchors and negative anchors Initially, integrate the user prompt P init into the Gaussian anchor domain, and thus generate the initial anchor through formula (3) The positive anchor is initialized as while the negative anchor is set as an empty set;
[0064] Perform the following iterative steps on each view for training:
[0065] For the given view v k , derive the positive and negative anchor masks from the positive and negative anchors using formula (1) and That is and F is the derivation function, and the candidate positive and negative prompts and are extracted from the anchor masks;
[0066] Subsequently, render the segmentation mask which predicts the segmentation of the current view; Based on S k , apply a confidence filtering process to discard incorrect positive class prompts; Use the edge area of determined by the Canny operator to filter the negative prompts represents the negation result of S k , that is, take the negation on the original S k . Finally, apply the farthest point sampling algorithm to the positive and negative prompt sets and to generate the final positive and negative prompt sets of SAM and Among them, for the initial view, since there are no candidate sampling prompts, the above steps will be ignored;
[0067] Next, use SAM to segment the image x k The sampling prompts are as follows:
[0068]
[0069] In the formula represents the predicted segmentation mask, and represent the positive and negative 2D prompts. Subsequently, for Perform non - projection, determine the new segmentation score through formula (2), which can be used as a reliable indicator for evaluating new view cues; in addition, use the farthest point sampling algorithm to sample n from the segmentation mask supp Supplementary cues and n supp represent the sampled points sampled from and are divided into positive and negative cue point sets and and fix the supplementary cues to the 3D scene through formula (3) to obtain and represents the Gaussian anchor domain of positive class points, represents the Gaussian anchor domain of negative class points; among them, in order to reduce the accumulation of false cues, only the anchors in the previous views in the cue library are retained, updated according to the first - in - first - out method, and the initial anchor is also retained
[0070]
[0071] 3) Step 2) will be repeated multiple times until all 2D images of the target object from all perspectives have been processed, and finally a 3D segmentation field is obtained as the final result.
[0072] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A three-dimensional scene segmentation method based on 3D Gaussian Splatting, characterized in that, Including the following steps: 1) 3D Gaussian Splatting scene preparation: 3D Gaussian Splatting is a 3D scene storage framework, whose Chinese name is 3D Gaussian Splatting, and it is an explicit storage structure; Use the colmap technology to perform sparse three-dimensional reconstruction on the multi-views of the scene to obtain sparse sample points; The training of 3D Gaussian Splatting requires the above-obtained sparse sample points as input, and at the same time requires multi-views and their camera poses, and optimizes the entire scene by rendering different views; 2) Use the trained 3D Gaussian Splatting as the base framework and multi-views as a bridge to perform three-dimensional segmentation on the target object in the scene: The scene renders a two-dimensional image of a certain perspective, which contains the target object, and then uses the 2D segmentation large model SAM to perform 2D segmentation on the two-dimensional image to obtain the 2D mask of the target object. After segmentation, project the 2D mask result back onto 3D Gaussian Splatting. The projection process is implemented using the Gaussian mixture algorithm. Specifically, use the rendering function of 3DGaussian Splatting to inverse-render the label for each Gaussian. The 2D mask ∈ [0,1]. Use the 2D mask to traverse all Gaussians in the two-dimensional image, and assign the semantic segmentation score as the mask mask*T and the transmittance weight T, and update the result of the 3D segmentation field to achieve three-dimensional segmentation; Among them, the 2D prompt when rendering the first view is given manually, and the 2D prompts for subsequent perspectives are automatically generated by the PinPrompt algorithm; The PinPrompt algorithm uses a loop anchoring scheme to anchor the 2D prompt to the 3D scene, so as to achieve efficient rendering from different viewpoints; 3) Step 2) will be repeated multiple times until all two-dimensional images containing the target object in all perspectives have been executed once, and finally a 3D segmentation field is obtained as the final result.
2. The three-dimensional scene segmentation method based on 3D Gaussian Splatting according to claim 1, characterized in that, In step 1), 3D Gaussian sputtering is a high-quality representation of an unstructured radiation field that utilizes a set of 3D Gaussian points to represent the scene, where N represents the number of Gaussian points, and g i represents the i-th Gaussian point. Each Gaussian point is associated with learnable parameters for representing its three-dimensional coordinates, scale, rotation angle, density, and color; The rendering of 3D Gaussian Splatting uses the α - blending technique. All Gaussian points are projected onto the image plane along a ray to obtain the per - pixel color for rendering. The α - blending technique can also be applied to render other properties, unifying the representation of the α - blending technique for colors and other properties, such that for any property The rendering process for a given pixel p is as follows: where α i represents the density of the Gaussian point g i and h i represents the property of the Gaussian point g i . represents the transparency, j represents the j-th Gaussian on the light ray, denotes the set of points on the light ray emitted from the pixel p, denotes the set of points on the light ray emitted from the pixel p; The set goal is to segment the objects of interest to the user in the 3D scene, that is, to label all Gaussian points related to the objects as 1, while other background points are labeled as 0, and an additional segmentation score s is assigned to each Gaussian point i , s i ∈[0,1]. These scores can be directly rendered into a 2D segmentation mask, that is, a 2D mask, for further supervision or evaluation, denoted as represents the set of segmentation scores, p represents the pixel, and at the same time 3. The three-dimensional scene segmentation method based on 3D Gaussian Splatting according to claim 2, wherein In step 2), the specific steps of the Gaussian mixture algorithm are as follows: First, the concepts of view - centered modeling and Gaussian - centered modeling are defined; view - centered modeling refers to the idea of modeling the pixel values of a 2D mask, which approximates the rendered view to the provided 2D mask. SA3D - GS is a view - centered modeling method that renders the 2D mask obtained from formula (1) and is optimized through backpropagation; GaussianEditor adopts Gaussian - centered modeling and uses a more intuitive modeling formula; for each Gaussian point, the segmentation score s i is calculated as the average of the different view rays of the 2D segmentation passing through that point; Consider the influence of occlusion during the weighting process. From a practical perspective, among the rays emitted from different viewpoints, unoccluded rays should be given priority, because the segmentation scores corresponding to occluded rays cannot be used as valid evidence. Therefore, introduce the transmittance T as a confidence factor during the segmentation score weighting. Finally, the Gaussian mixture algorithm can be expressed as: where T i p represents the transparency, represents a 2D mask, represents the set of Gaussian points where the rays emitted by the indexed pixel p intersect.
4. The three-dimensional scene segmentation method based on 3D Gaussian Splatting according to claim 3, wherein In step 2), the specific implementation of the PinPrompt algorithm is as follows: Given a set of multi-view images captured from a set of views where v k denotes the k-th training view, N v denotes the total number of captured images, and x k denotes the k-th captured image, the goal is to assign a segmentation score s to each Gaussian point based only on an initial prompt provided by the user i , p n denotes the prompt on the n-th image, and N init denotes the initial number of images, by developing a prompt propagation strategy and extending the initial prompt to generate reliable prompts for other views, thus facilitating the calculation of 3D segmentation, as shown in Equation (2); A new prompt sampling and propagation strategy is proposed, called the PinPrompt algorithm, which aims to promote accurate prompt transmission across different viewpoints. The PinPrompt algorithm uses a recycling scheme to anchor the 2D prompt to the 3D scene, so as to achieve efficient rendering from different viewpoints. To ensure the effectiveness of the anchor points under the new view, use the few-shot characteristics of the Gaussian mixture algorithm to filter and select the anchor points. The PinPrompt promotes the collaborative refinement between 3D segmentation and fast propagation; To fix the 2D hint to the 3D scene, associate the Gauss points on the object surface with the attribute a i ∈ {0, 1}, representing the Gauss anchor domain, a i representing the attribute value of the i-th Gauss point. A total of N Gauss points are marked, indicating the existence of point hints. If a certain Gauss point g i is marked as a prompt, then its attribute a i is assigned the value 1, which is expressed as follows: Given the hint P, determine the first Gauss point intersected by the ray emitted by each pixel p in P, and mark its attribute a i as 1. Formally: where d p (i) represents the distance between i and p, represents the set of distances of the Gaussian points along the ray emitted from p, represents the set of Gaussian points where the rays emitted from the indexed pixel p intersect; In Equation (3), all hints require a pinning-back process, and the final a i depends on the union of flags to ensure consistency. After determining it, a Gaussian anchor domain can be effectively constructed. Similarly, the set of Gaussian segmentation domains S can be interpreted as a Gaussian segmentation field; Rendering attributes in 3D Gaussian Splatting, for the convenience of using the Gaussian segmentation domain S and the Gaussian anchor domain Define the rendering as follows: Given a set of attributes and a view v k , calculate the rendering mask according to formula (1) represents the set of pixels on view v k , where Derive a 2D mask with high confidence through the PinPrompt algorithm and establish a prompt library containing positive anchors and negative anchors Initially, integrate the user prompt P init into the Gaussian anchor domain to generate an initial anchor through formula (3) The positive anchor is initialized as while the negative anchor is set as an empty set; Perform the following iterative steps for training on each view: For a given view v k , the positive and negative anchor masks are derived from the positive and negative anchors using Equation (1) and i.e. and where F is the derivation function, and the candidate positive and negative prompts and are extracted from the anchor masks Subsequently, render the segmentation mask It predicts the segmentation of the current view; based on S k , apply a confidence filtering process to discard false positive class cues; Use the edge area determined by the Canny operator to filter negative prompts and represents the negation result of S k , that is, take the negation on the original S k . Finally, apply the farthest point sampling algorithm to the positive and negative prompt sets and to generate the final positive and negative prompt sets of SAM and Among them, for the initial view, since there are no candidate sampling prompts, the above steps will be ignored; Next, use SAM to segment the image x k The sampling prompts are as follows: In the formula, represents the predicted segmentation mask, and represent positive and negative 2D cues. Subsequently, is unprojected, and a new segmentation score is determined through formula (2). This score can be used as a reliable metric for evaluating the new view cues. Additionally, the farthest point sampling algorithm is used to sample n from the segmentation mask supp Supplementary cues and n supp represents the sampled points obtained by sampling from and is divided into positive and negative cue point sets and and the supplementary cues are fixed to the 3D scene through formula (3) to obtain and represent the Gaussian anchor regions for positive class points, represents the Gaussian anchor region for negative class points. Among them, in order to reduce the accumulation of false cues, only the anchors in the previous views in the cue library are retained, and it is updated according to the first-in-first-out method, while retaining the initial anchor
Citation Information
Cited By
Path planning device of bronchoscope
CN121196730A