Real-time multi-instance segmentation method and device based on Gaussian splash radiation field model
Through a real-time multi-instance segmentation method based on the Gaussian splash radiation field model, the problems of multi-view consistency and training complexity in three-dimensional scene semantic segmentation are solved, and fast, semantically consistent multi-instance segmentation and tracking are achieved, which is suitable for tasks such as augmented reality and robotic perception.
Patent Information
- Application Number
- CN202510700224.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-12
AI Technical Summary
Existing 3D scene semantic segmentation and tracking methods have difficulty ensuring semantic consistency between multiple perspectives, have complex training processes, high deployment costs, and are unable to meet the needs of interactive segmentation that does not require training and is plug-and-play.
Based on the pre-trained two-dimensional Gaussian splash radiation field model, the multi-view consistent two-dimensional instance segmentation mask is obtained by rendering the image sequence and using the video segmentation model. It is mapped to the three-dimensional space and a voting strategy is used to assign unique instance labels. The Gaussian splash algorithm is used to obtain multi-view consistent segmentation masks under any perspective, and edge detection and regional connectivity restoration are performed through a lightweight post-processing network.
It achieves fast and semantically consistent multi-instance segmentation and supports continuous tracking of multiple targets. It is suitable for real-time 3D understanding tasks such as augmented reality and robotic perception, and meets the dual requirements of instance accuracy and system response speed.
Smart Images

Figure CN120635100A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer vision and three-dimensional space modeling, and in particular to a real-time multi-instance segmentation method and device based on a Gaussian splash radiation field model. Background Art
[0002] With the rapid development of 3D perception and neural rendering technologies, 3D Gaussian Splatting (3DGS), an explicit, efficient, and editable method for representing 3D scenes, has gradually become a key enabling technology in the fields of intelligent perception and spatial modeling. By explicitly modeling a large number of Gaussian primitives in 3D space and combining camera projection with a differentiable rasterization rendering mechanism, this method enables real-time rendering from any viewpoint while maintaining image detail quality. This approach holds great promise in applications such as augmented reality, robotic navigation, virtual simulation, and digital twins.
[0003] Although 3DGS has demonstrated significant advantages in modeling and rendering, research in three-dimensional semantic understanding, especially multi-target semantic segmentation and consistent instance tracking, is still in its infancy. Existing methods mostly generate semantic masks based on two-dimensional image models and extend the two-dimensional semantics to three-dimensional structures through projection mapping, feature distillation, or label alignment. However, these methods suffer from problems such as accuracy dependence on two-dimensional models, inconsistent labels, blurred boundaries, and structural drift. They are unable to ensure semantic consistency across multiple viewpoints, and their training process is complex, deployment costs are high, and their practical application effects are limited.
[0004] In recent years, several studies, such as Feature3DGS, OmniSeg3D, and SAGA, have attempted to improve the accuracy and robustness of semantic prediction through mechanisms such as multi-view consistency modeling, context fusion, and contrastive learning. However, these methods often still require modifications to the Gaussian structure itself or rely on additional training processes, making it difficult to meet the requirements of interactive segmentation that requires no training and is plug-and-play. Furthermore, existing methods also face technical bottlenecks in handling multi-instance representations, maintaining cross-view consistency, and improving segmentation boundary clarity, making it difficult to simultaneously meet the dual requirements of semantic accuracy and system response speed in real-world scenarios.
[0005] Therefore, existing 3D scene semantic segmentation and tracking methods are difficult to ensure semantic consistency between multiple perspectives, have complex training processes, high deployment costs, and are difficult to meet the needs of interactive segmentation that does not require training and is plug-and-play. Summary of the Invention
[0006] This application provides a real-time multi-instance segmentation method based on the Gaussian splash radiation field model to solve the problems in the existing technology that the existing three-dimensional scene semantic segmentation and tracking methods are difficult to ensure semantic consistency between multiple perspectives, the training process is complex, the deployment cost is high, and it is difficult to meet the interactive segmentation requirements that do not require training and are plug-and-play.
[0007] Correspondingly, the present application also provides a real-time multi-instance segmentation device based on a Gaussian splash radiation field model, an electronic device, and a computer-readable storage medium to ensure the implementation and application of the above method.
[0008] In order to solve the above technical problems, the present application discloses a real-time multi-instance segmentation method based on a Gaussian splash radiation field model, the method comprising:
[0009] Based on a pre-trained 2D Gaussian splatter radiation field model, we render the scene with spatially continuously changing viewpoints to generate an image sequence. We then use a pre-set video segmentation model to obtain a 2D instance segmentation mask that is consistent across multiple viewpoints.
[0010] The 2D instance segmentation masks from multiple views are mapped to 3D space, and a preset voting strategy is used to assign a unique instance label to each Gaussian basis element in the 2D Gaussian splatter radiation field model.
[0011] For the two-dimensional Gaussian splash radiation field model with instance labels, the Gaussian splash algorithm is used to obtain the coarse instance segmentation mask and corresponding color image with multi-view consistency under any viewing angle;
[0012] The coarse instance segmentation mask and color image are input into a lightweight post-processing network for edge detection and region connectivity restoration to obtain the target two-dimensional instance segmentation mask.
[0013] The present application also discloses a real-time multi-instance segmentation device based on a Gaussian splash radiation field model, the device comprising:
[0014] The image segmentation module is used to render spatially continuously changing viewpoints based on a pre-trained 2D Gaussian splatter radiation field model to generate image sequences, and to obtain multi-view consistent 2D instance segmentation masks for the image sequences using a preset video segmentation model.
[0015] A label generation module is used to map the 2D instance segmentation masks from multiple views into 3D space and assign a unique instance label to each Gaussian basis element in the 2D Gaussian splatter radiation field model using a preset voting strategy;
[0016] An image rendering module is used to obtain a coarse instance segmentation mask and a corresponding color image with multi-view consistency under any viewing angle using a Gaussian splashing algorithm for a two-dimensional Gaussian splashing radiation field model with instance labels;
[0017] The image fine segmentation module is used to input the coarse instance segmentation mask and color image into a lightweight post-processing network for edge detection and region connectivity repair to obtain the target two-dimensional instance segmentation mask.
[0018] The present application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, one or more methods described in the present application are implemented.
[0019] The present application also discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, one or more methods described in the present application are implemented.
[0020] In this application, based on a pre-trained two-dimensional Gaussian splash radiation field model, a perspective with continuous spatial variation is rendered to obtain an image sequence, and a preset video segmentation model is used to obtain a two-dimensional instance segmentation mask with multi-perspective consistency of the image sequence; based on the multi-perspective two-dimensional instance segmentation mask, a unique instance label is assigned to each Gaussian primitive in the two-dimensional Gaussian splash radiation field model to achieve three-dimensional instance segmentation with a granularity of Gaussian primitive level. Afterwards, for the two-dimensional Gaussian splash radiation field model with instance labels, a Gaussian splash algorithm is used to obtain a coarse instance segmentation mask with multi-perspective consistency and a corresponding color image at any perspective. The segmentation speed is higher and is not limited to the requirement that the perspective must meet the spatial continuous variation requirement. Finally, the coarse instance segmentation mask and the color image are input into a lightweight post-processing network for edge detection and regional connectivity repair, and a target two-dimensional instance segmentation mask with a complete structure and labels consistent with the three-dimensional scene is output. This application does not rely on any training or distillation process and directly acts on the Gaussian splash radiation field model. It has the advantages of fast inference speed, strong semantic consistency, and support for multi-target continuous tracking. It is suitable for real-time three-dimensional understanding task scenarios such as augmented reality, robotic perception, and three-dimensional modeling. It can meet the dual requirements of actual scenarios for instance accuracy and system response speed.
[0021] Additional aspects and advantages of the present application will be given in the following description, which will become apparent from the following description, or will be understood through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0023] Figure 1 A flowchart of a real-time multi-instance segmentation method based on a Gaussian splash radiation field model provided in an embodiment of the present application;
[0024] Figure 2 A schematic diagram of the process of generating a two-dimensional mask for multi-view rendering and video segmentation provided in an embodiment of the present application;
[0025] Figure 3 A diagram of the lightweight post-processing network structure provided in an embodiment of the present application;
[0026] Figure 4 A schematic diagram of a mask image generated by semantic Gaussian rendering according to an embodiment of the present application;
[0027] Figure 5 A comparison chart before and after optimization of the lightweight post-processing network provided in an embodiment of the present application;
[0028] Figure 6 A schematic structural diagram of a real-time multi-instance segmentation device based on a Gaussian splash radiation field model provided in an embodiment of the present application;
[0029] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following describes embodiments of the present application in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application.
[0031] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0032] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which the present invention pertains. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as herein, will not be interpreted in an idealized or overly formal sense.
[0033] The solution provided in the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server, wherein the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this. With respect to the technical problems existing in the prior art, the real-time multi-instance segmentation method and device based on the Gaussian splash radiation field model provided in this application is intended to solve at least one of the technical problems of the prior art.
[0034] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0035] The present application embodiment provides a possible implementation method, such as Figure 1 As shown, a flowchart of a real-time multi-instance segmentation method based on a Gaussian splash radiation field model is provided. The solution can be executed by any electronic device, and optionally, can be executed on a server or terminal device.
[0036] like Figure 1 As shown in , the method may include the following steps:
[0037] Step 101 : Based on a pre-trained two-dimensional Gaussian splatter radiation field model, rendering is performed on a view with continuous spatial variation to generate an image sequence, and a preset video segmentation model is used to obtain a two-dimensional instance segmentation mask with multi-view consistency of the image sequence.
[0038] The pre-trained two-dimensional Gaussian splatter radiation field model is obtained by iteratively training the input color image using the existing 2DGS algorithm. The 2DGS algorithm is a variant of the 3DGS algorithm that can extract accurate geometric structure from the input color image.
[0039] In step 102 , the two-dimensional instance segmentation mask under multiple views is mapped to a three-dimensional space, and a preset voting strategy is used to assign a unique instance label to each Gaussian basis element in the two-dimensional Gaussian splash radiation field model.
[0040] Step 103 : For the two-dimensional Gaussian splash radiation field model with instance labels, a Gaussian splash algorithm is used to obtain a coarse instance segmentation mask and a corresponding color image with multi-view consistency at any viewing angle.
[0041] In step 104 , the coarse instance segmentation mask and the color image are input into a lightweight post-processing network for edge detection and region connectivity restoration to obtain a target two-dimensional instance segmentation mask.
[0042] like Figure 2 As shown in , the embodiment of the present application includes two stages. Stage 1 is a warm-up initialization stage. In this stage, a two-dimensional instance segmentation mask with multi-perspective consistency is obtained by performing spatially continuously changing multi-perspective rendering and video segmentation model on the two-dimensional Gaussian splash radiation field model. The above-mentioned two-dimensional instance segmentation mask with multi-perspective consistency is converted into instance labels through an inverse grating projection algorithm, so that the Gaussian basis elements in the two-dimensional Gaussian splash radiation field model are assigned unique instance labels, and then a two-dimensional Gaussian splash radiation field model containing instance labels is constructed. Stage 2 is an acceleration stage. In this stage, the two-dimensional Gaussian splash radiation field model with instance labels is rendered at any angle to obtain a color image at any perspective and a coarse instance segmentation mask with multi-perspective consistency. The color image and the coarse instance segmentation mask with multi-perspective consistency are input into the post-processing network, the segmentation mask is further refined, and the multi-perspective consistent instance segmentation mask is output as the target two-dimensional instance segmentation mask to achieve fast and accurate segmentation output.
[0043] In an embodiment of the present application, based on a pre-trained two-dimensional Gaussian splash radiation field model, a perspective with spatially continuous changes is rendered to obtain an image sequence, and a preset video segmentation model is used to obtain a two-dimensional instance segmentation mask with multi-perspective consistency of the image sequence; the two-dimensional instance segmentation mask under multiple perspectives is mapped to a three-dimensional space, and a preset voting strategy is used to assign a unique instance label to each Gaussian primitive in the two-dimensional Gaussian splash radiation field model, and a two-dimensional Gaussian splash radiation field model with instance labels is constructed to achieve three-dimensional instance segmentation at the Gaussian primitive level. For the two-dimensional Gaussian splash radiation field model with instance labels, a Gaussian splash algorithm is used to obtain a coarse instance segmentation mask and a corresponding color image with multi-perspective consistency under any perspective, which has a higher segmentation speed and is not limited to the requirement that the perspective must meet the spatial continuous change requirement. Finally, the coarse instance segmentation mask and the color image are input into a lightweight post-processing network for edge detection and regional connectivity repair, and the output is a target two-dimensional instance segmentation mask with a complete structure and labels consistent with the three-dimensional scene. The embodiment of the present application does not rely on any training or distillation process, and directly acts on the Gaussian splash radiation field model. It has the advantages of fast reasoning speed, strong semantic consistency, and support for continuous tracking of multiple targets. It is suitable for real-time three-dimensional understanding task scenarios such as augmented reality, robot perception, and three-dimensional modeling, and can meet the dual requirements of actual scenarios for instance accuracy and system response speed.
[0044] In an optional embodiment, if Figure 2 As shown in , based on the pre-trained 2D Gaussian splatter radiation field model, the spatially continuously changing perspective is rendered to generate an image sequence, and a preset video segmentation model is used to obtain a multi-view consistent 2D instance segmentation mask for the image sequence, including:
[0045] Based on a pre-trained two-dimensional Gaussian splatter radiation field model, multiple perspectives with continuous spatial variation are rendered to obtain color images with continuous spatial variation. These color images have the characteristics of continuous spatial variation similar to video frames.
[0046] The spatially continuously changing color images are formed into an image sequence in time order;
[0047] Determine the initial coordinates of a target instance based on the color image where the target instance first appears in the image sequence. Specifically, perform a mouse click operation (denoted as clicks) on the first color image frame where the target instance appears to obtain the initial click coordinates of each target instance.
[0048] The image sequence and the initial click coordinates corresponding to each target instance are input into a general video segmentation model (such as Segment Anything Model 2, SAM2) to generate a two-dimensional instance segmentation mask with multi-view consistency. The output is:
[0049]
[0050] in It represents the two-dimensional instance segmentation mask of the color image of the frame, the pixel value represents the instance ID, and I1 represents the selected color image.
[0051] In an optional embodiment, the two-dimensional instance segmentation mask under multiple views is mapped to a three-dimensional space, and a preset voting strategy is used to assign a unique instance label to each Gaussian basis element in the two-dimensional Gaussian splatter radiation field model, including:
[0052] Projecting the 3D center coordinates of each Gaussian basis element in the 2D Gaussian splash radiation field model onto the 2D instance segmentation masks at different viewpoints to obtain the projected pixel positions of the Gaussian basis element in the 2D instance segmentation masks at different viewpoints;
[0053] The pixel value at the projected position is used as the semantic labeling candidate value of the Gaussian primitive at the corresponding viewing angle;
[0054] The occurrence frequencies of the semantic annotation candidate values of the Gaussian primitives under all viewpoints are counted, and the semantic annotation candidate value with the highest occurrence frequency is selected as the final instance label of the Gaussian primitive using the majority voting strategy.
[0055] In the embodiment of the present application, for each Gaussian basis element G in the two-dimensional Gaussian splash radiation field model, i Extract its 3D center coordinates Using the camera projection function π corresponding to each perspective t (·) Project it to the 2D instance mask M of different perspectives in turn t 2D The projected pixel position π of each Gaussian primitive in the t-th frame 2D instance segmentation mask is obtained. t (c i );
[0056] In the 2D instance segmentation mask, read the projected pixel position π in the t-th frame t (c i ), i.e., the corresponding instance ID, as the Gaussian basis element G i Semantic annotation candidate values under this perspective;
[0057] Count the occurrence frequencies of different instance IDs assigned to the Gaussian primitive in all viewpoints, and use the majority voting strategy to select the instance ID with the highest frequency. i As the final instance label of the Gaussian primitive, its assignment process can be expressed as:
[0058]
[0059] Among them, I[·] is the indicator function. When the condition M t 2D (π t (c i )) takes 1 when k holds, otherwise it takes 0, where k represents the set of all possible instance ID labels;
[0060] The final instance label l i Write the corresponding Gaussian basis elements to form a Gaussian set with semantic labels:
[0061]
[0062] In an optional embodiment, after mapping the two-dimensional instance segmentation masks under multiple views to a three-dimensional space and assigning a unique instance label to each Gaussian basis element in the two-dimensional Gaussian splatter radiation field model using a preset voting strategy, the method further includes:
[0063] The original spatial attributes of each Gaussian primitive are retained, and instance labels are attached to generate a two-dimensional Gaussian splash radiation field model with instance labels.
[0064] In the embodiment of the present application, the Gaussian primitives set to which instance labels are assigned in the above steps is constructed as a 2D Gaussian splatter radiation field model with instance labels, thereby achieving 3D instance segmentation at the Gaussian primitive level. The original spatial properties of the Gaussian primitives (which may include 3D center position, orientation, scale, color, and transparency) are retained, while a unique instance label is attached for subsequent rendering and segmentation tasks. Ultimately, each Gaussian primitive can be represented as a five-tuple:
[0065] G i =(c i ,∑ i ,α i ,f i ,l i )
[0066] in, represents the three-dimensional center coordinates, is the covariance matrix, α i ∈[0,1] is opacity, f i is the spherical harmonic function parameter representing color information, l i A unique tag assigned to the instance.
[0067] In an optional embodiment, for a two-dimensional Gaussian splash radiation field model with instance labels, a Gaussian splash algorithm is used to obtain a coarse instance segmentation mask and a corresponding color image with multi-view consistency at any viewing angle, including:
[0068] The two-dimensional Gaussian splash radiation field model with instance labels is rendered at any viewing angle using the Gaussian splash algorithm to obtain a color image at the corresponding viewing angle.
[0069] For each pixel in the color image, the corresponding label is determined according to the associated Gaussian primitive to form a coarse instance segmentation mask.
[0070] In an optional embodiment, for each pixel in the color image, a corresponding label is determined according to the associated Gaussian primitive to form a coarse instance segmentation mask, including:
[0071] Traverse each pixel in the color image, find the projection set of all visible Gaussian primitives in its viewing cone, assign the Gaussian primitive that meets the rendering conditions to the label corresponding to the pixel, and obtain the coarse instance segmentation mask.
[0072] In the embodiment of the present application, given the viewing angle parameter θ t Next, we use the differentiable Gaussian forward rendering function Two-dimensional Gaussian splash radiation field model G with instance labels seg Extract the color information of each Gaussian primitive and perform rendering on the image plane to obtain a color image at the corresponding viewing angle
[0073] Traverse each pixel (u, v) in the color image and find the projection set of all visible Gaussian primitives in its viewing cone Assign the Gaussian primitive that meets the rendering conditions to the label c corresponding to the pixel position k , and obtain the coarse instance segmentation mask, which is calculated as:
[0074]
[0075] in, It represents the Euclidean distance between the pixel and the center of the Gaussian projection. t_th represents the transmittance of the rendering when it passes through the Gaussian basis element. τ is an empirical parameter that can be set to 1, and α is an empirical parameter that can be set to 0.1.
[0076] In order to simultaneously acquire color images and coarse instance segmentation mask A Gaussian rasterizer that supports instance-ID rendering is used to render the Gaussian splatter radiation field model in real time at any viewpoint. This approach avoids the limitations of spatially continuous viewpoint changes and supports fast semantic segmentation at any viewpoint.
[0077] In an optional embodiment, the coarse instance segmentation mask and the color image are input into a lightweight post-processing network for edge detection and region connectivity restoration to obtain a final two-dimensional instance segmentation mask, including:
[0078] Segment the coarse instance segmentation mask and the corresponding color image according to the bounding box of each instance area to obtain a local binary mask containing a single instance and the corresponding local color image;
[0079] The local binary mask and the local color image are input into a lightweight post-processing network for edge detection and regional connectivity repair, and a binary mask is output;
[0080] The binary masks of each instance are reconstructed into a complete target 2D instance segmentation mask.
[0081] In the embodiment of the present application, the generated coarse semantic mask and the corresponding color image One-to-one correspondence, based on the approximate bounding box of each instance area in the mask, the local image area of each instance object is cropped out respectively, and a local binary mask containing a single instance and the corresponding local color image are obtained. The pixels in the target instance area in the local binary mask are assigned a value of 1, and the background area is assigned a value of 0.
[0082] Each set of local binary masks and local color images obtained above are input into the lightweight image post-processing network. Figure 2 and Figure 3As shown in , the network uses a shallow encoding-decoding structure, including an image encoder, a hint encoder, and a mask decoder. The hint encoder includes three convolutional modules, each followed by a pooling operation. In the hint encoder, a local binary mask (referred to as a coarse mask) of size (1, H, W) passes through convolution module 1 and pooling to obtain features of size (32, H / 2, W / 2). This feature passes through convolution module 2 and pooling to obtain features of size (64, H / 4, W / 4). This feature passes through convolution module 3 and pooling to obtain features of size (128, H / 8, W / 8). Finally, it is downsampled to obtain mask features of size (128, H / 32, W / 32). The image encoder includes one convolutional module and four convolutional layers. In the image encoder, a local color image of size (3, H, W) is processed by convolution module 1, normalization, and the ReLU function to obtain features of size (64, H / 2, W / 2). This feature is processed by pooling and convolution layer 1 to obtain features of size (256, H / 4, W / 4). Convolution layer 2 outputs features of size (512, H / 8, W / 8). Convolution layer 3 outputs features of size (1024, H / 16, W / 16). Convolution layer 4 outputs image features of size (2048, H / 32, W / 32). The mask features output by the hint encoder and the image features output by the image encoder are then concatenated for feature fusion. After 1×1 convolution, normalization, the ReLU function, and channel attention, the fused features of size (512, H / 32, W / 32) are obtained. The mask decoder includes 4 decoders. In the mask decoder: the fused feature is upsampled by decoder 4 (1024+512) to obtain a feature of size (256, H / 16, W / 16), which is upsampled by decoder 3 (512+256) to obtain a feature of size (128, H / 8, W / 8), which is upsampled by decoder 2 (256+128) to obtain a feature of size (64, H / 4, W / 4), which is upsampled by decoder 1 and finally outputs a new binary mask with clear boundaries and complete structure, whose size is (2, H, W).
[0083] This lightweight image post-processing network has edge perception and region repair capabilities, which can effectively refine the edge contours and morphologically correct the mask, and output a new binary mask with clear boundaries and complete structure.
[0084] The new binary mask corresponding to each instance is reassembled back to the original image size based on the original image position information and instance ID information recorded before cropping, generating a multi-view consistent instance segmentation mask with complete structure and correct labels as the target two-dimensional instance segmentation mask.
[0085] Figure 4 The diagram shows the result of semantic Gaussian rendering to generate mask images. The first column shows representative scenes from different data sets. Columns 2-4 show the segmentation outputs at view angles 0, 1, and 2 using the method in the embodiment of the present application. Column 5 shows the true value at view angle 0. Figure 3 It can be seen that in the fork scene (first row), the method in the embodiment of the present application can recover finer geometric details compared with the true value annotation; it can also produce finer output in other scenes.
[0086] Figure 5 The following figure shows a comparison before and after optimization of the lightweight post-processing network. The left side shows the segmentation result obtained directly from the Gaussian scene rendering mask, without post-processing. The right side shows the target 2D instance segmentation mask and segmentation effect after the lightweight post-processing network. It can be clearly seen that the output details of the lightweight post-processing network are much finer.
[0087] Based on the above method, the embodiment of the present application can accurately and quickly realize multi-instance segmentation and tracking in a two-dimensional Gaussian splash radiation field model. Compared with existing solutions that rely on contrastive learning training or distillation, the embodiment of the present application adopts a combination of explicit geometric voting and differentiable rendering to achieve high-quality mask generation without training. Through the Gaussian center back projection and label voting mechanism, problems such as multi-view label drift and boundary blur are effectively overcome; a lightweight post-processing network is introduced to further optimize structural connectivity and boundary clarity. This method has the advantages of easy deployment and rapid response, and is suitable for real-time three-dimensional scene understanding tasks such as augmented reality, three-dimensional mapping and intelligent perception.
[0088] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application also provides a real-time multi-instance segmentation device based on the Gaussian splash radiation field model, such as Figure 6 As shown, the device includes:
[0089] An image segmentation module 601 is configured to render a spatially continuously varying viewpoint based on a pre-trained two-dimensional Gaussian splatter radiation field model to generate an image sequence, and to obtain a multi-view consistent two-dimensional instance segmentation mask for the image sequence using a preset video segmentation model;
[0090] a label generation module 602 for mapping the two-dimensional instance segmentation mask under multiple views into a three-dimensional space and assigning a unique instance label to each Gaussian basis element in the two-dimensional Gaussian splatter radiation field model using a preset voting strategy;
[0091] An image rendering module 603 is configured to use a Gaussian splatting algorithm to obtain a coarse instance segmentation mask and a corresponding color image with multi-view consistency at any viewing angle for a two-dimensional Gaussian splatting radiation field model with instance labels;
[0092] The image fine segmentation module 604 is used to input the coarse instance segmentation mask and the color image into a lightweight post-processing network for edge detection and region connectivity restoration to obtain a target two-dimensional instance segmentation mask.
[0093] In an embodiment of the present application, based on a pre-trained two-dimensional Gaussian splash radiation field model, a perspective with spatially continuous changes is rendered to obtain an image sequence, and a preset video segmentation model is used to obtain a two-dimensional instance segmentation mask with multi-perspective consistency of the image sequence; the two-dimensional instance segmentation mask under multiple perspectives is mapped to a three-dimensional space, and a preset voting strategy is used to assign a unique instance label to each Gaussian primitive in the two-dimensional Gaussian splash radiation field model, and a two-dimensional Gaussian splash radiation field model with instance labels is constructed to achieve three-dimensional instance segmentation at the Gaussian primitive level. For the two-dimensional Gaussian splash radiation field model with instance labels, a Gaussian splash algorithm is used to obtain a coarse instance segmentation mask and a corresponding color image with multi-perspective consistency under any perspective, which has a higher segmentation speed and is not limited to the requirement that the perspective must meet the spatial continuous change requirement. Finally, the coarse instance segmentation mask and the color image are input into a lightweight post-processing network for edge detection and regional connectivity repair, and the output is a target two-dimensional instance segmentation mask with a complete structure and labels consistent with the three-dimensional scene. The embodiment of the present application does not rely on any training or distillation process, and directly acts on the Gaussian splash radiation field model. It has the advantages of fast reasoning speed, strong semantic consistency, and support for continuous tracking of multiple targets. It is suitable for real-time three-dimensional understanding task scenarios such as augmented reality, robot perception, and three-dimensional modeling, and can meet the dual requirements of actual scenarios for instance accuracy and system response speed.
[0094] The real-time multi-instance segmentation device based on the Gaussian splash radiation field model provided in the embodiment of the present application can achieve Figures 1 to 5 To avoid repetition, the various processes implemented in the method embodiment will not be described again here.
[0095] The real-time multi-instance segmentation device based on the Gaussian splash radiation field model of the embodiment of the present application can execute the real-time multi-instance segmentation method based on the Gaussian splash radiation field model provided by the embodiment of the present application. The implementation principle is similar. The actions performed by each module and unit in the real-time multi-instance segmentation device based on the Gaussian splash radiation field model in each embodiment of the present application correspond to the steps in the real-time multi-instance segmentation method based on the Gaussian splash radiation field model in each embodiment of the present application. For the detailed functional description of each module of the real-time multi-instance segmentation device based on the Gaussian splash radiation field model, please refer to the description of the corresponding real-time multi-instance segmentation method based on the Gaussian splash radiation field model shown in the previous text, which will not be repeated here.
[0096] Based on the same principle as the method shown in the embodiment of the present application, the embodiment of the present application also provides an electronic device, which may include but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the real-time multi-instance segmentation method based on the Gaussian splash radiation field model shown in any optional embodiment of the present application by calling the computer program. Compared with the existing technology, the real-time multi-instance segmentation method based on the Gaussian splash radiation field model provided by the present application does not rely on any training or distillation process, and directly acts on the Gaussian splash radiation field model. It has the advantages of fast inference speed, strong semantic consistency, and support for continuous tracking of multiple targets. It is suitable for real-time three-dimensional understanding task scenarios such as augmented reality, robot perception, and three-dimensional modeling, and can meet the dual requirements of actual scenarios for instance accuracy and system response speed.
[0097] In an optional embodiment, an electronic device is also provided, such as Figure 7 As shown, Figure 7 The electronic device 700 shown may be a server, including a processor 701 and a memory 703. The processor 701 and the memory 703 are connected, for example, via a bus 702. Optionally, the electronic device 700 may further include a transceiver 704. It should be noted that in actual applications, the number of transceivers 704 is not limited to one, and the structure of the electronic device 700 does not constitute a limitation on the embodiments of the present application.
[0098] The processor 701 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor 701 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0099] The bus 702 may include a path for transmitting information between the above components. The bus 702 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 702 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0100] The memory 703 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0101] The memory 703 is used to store application code for executing the solution of the present application, and the execution is controlled by the processor 701. The processor 701 is used to execute the application code stored in the memory 703 to implement the content shown in the above method embodiment.
[0102] Among them, electronic devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0103] The server provided in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to these. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and this application does not limit this.
[0104] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding contents of the aforementioned method embodiment.
[0105] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0106] It should be noted that the computer-readable storage medium mentioned above in this application may also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0107] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0108] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0109] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the real-time multi-instance segmentation method and apparatus based on a Gaussian splatter radiation field model provided in the various optional implementations described above.
[0110] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0111] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0112] The modules described in the embodiments of the present application may be implemented in software or in hardware. The name of the module does not, in some cases, limit the module itself. For example, the image segmentation module may also be described as "an image segmentation module for rendering a spatially continuously varying perspective based on a pre-trained two-dimensional Gaussian splatter radiation field model, generating an image sequence, and obtaining a two-dimensional instance segmentation mask with multi-perspective consistency of the image sequence using a preset video segmentation model."
[0113] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A real-time multi-instance segmentation method based on Gaussian splash radiation field model, characterized in that: The method comprises: Based on a pre-trained 2D Gaussian splatter radiation field model, the image sequence with spatially continuously changing viewpoints is rendered. A preset video segmentation model is then used to obtain a 2D instance segmentation mask with multi-view consistency for the image sequence. Mapping the two-dimensional instance segmentation mask under multiple views to a three-dimensional space, and assigning a unique instance label to each Gaussian basis element in the two-dimensional Gaussian splash radiation field model using a preset voting strategy; For the two-dimensional Gaussian splash radiation field model with instance labels, the Gaussian splash algorithm is used to obtain the coarse instance segmentation mask and corresponding color image with multi-view consistency under any viewing angle; The coarse instance segmentation mask and the color image are input into a lightweight post-processing network for edge detection and region connectivity restoration to obtain a target two-dimensional instance segmentation mask.
2. The real-time multi-instance segmentation method based on the Gaussian splash radiation field model according to claim 1, characterized in that: The method of rendering a spatially continuously changing perspective based on a pre-trained two-dimensional Gaussian splatter radiation field model to generate an image sequence and obtaining a multi-perspective consistent two-dimensional instance segmentation mask of the image sequence using a preset video segmentation model includes: Based on the pre-trained two-dimensional Gaussian splash radiation field model, multiple perspectives with continuous spatial changes are rendered to obtain spatially continuous color images; The spatially continuously changing color images are used to form the image sequence in a temporal order; determining the initial click coordinates of the target instance according to the color image where the target instance first appears in the image sequence; The image sequence and the initial click coordinates corresponding to each target instance are input into a video segmentation model to generate a two-dimensional instance segmentation mask with multi-view consistency.
3. The real-time multi-instance segmentation method based on the Gaussian splash radiation field model according to claim 1, characterized in that: Mapping the two-dimensional instance segmentation mask under multiple views to a three-dimensional space, and using a preset voting strategy to assign a unique instance label to each Gaussian basis element in the two-dimensional Gaussian splash radiation field model, includes: Projecting the 3D center coordinates of each Gaussian basis element in the 2D Gaussian splash radiation field model onto the 2D instance segmentation masks at different viewpoints to obtain the projected pixel positions of the Gaussian basis element in the 2D instance segmentation masks at different viewpoints; The pixel value at the projected position is used as a semantic annotation candidate value of the Gaussian primitive at the corresponding viewing angle; The occurrence frequencies of the semantic annotation candidate values of the Gaussian primitives under all viewpoints are counted, and the semantic annotation candidate value with the highest occurrence frequency is selected as the final instance label of the Gaussian primitive using the majority voting strategy.
4. The real-time multi-instance segmentation method based on the Gaussian splash radiation field model according to claim 1, characterized in that: After mapping the two-dimensional instance segmentation mask under multiple views to a three-dimensional space and assigning a unique instance label to each Gaussian basis element in the two-dimensional Gaussian splatter radiation field model using a preset voting strategy, the method further includes: The original spatial attributes of each Gaussian basis element are retained, and the instance label is attached, so as to generate a two-dimensional Gaussian splash radiation field model with the instance label.
5. The real-time multi-instance segmentation method based on the Gaussian splash radiation field model according to claim 1, characterized in that: The method uses a Gaussian splashing algorithm to obtain a coarse instance segmentation mask and a corresponding color image with multi-view consistency at any viewing angle for a two-dimensional Gaussian splashing radiation field model with instance labels, including: The two-dimensional Gaussian splash radiation field model with instance labels is rendered at any viewing angle using the Gaussian splash algorithm to obtain a color image at the corresponding viewing angle. For each pixel in the color image, a corresponding label is determined according to the associated Gaussian basis to form the coarse instance segmentation mask.
6. The real-time multi-instance segmentation method based on the Gaussian splash radiation field model according to claim 5, characterized in that: For each pixel in the color image, determining a corresponding label according to an associated Gaussian basis to form the coarse instance segmentation mask includes: Traversing each pixel in the color image, searching for the projection set of all visible Gaussian primitives in its viewing cone, assigning the Gaussian primitives that meet the rendering conditions as the label corresponding to the pixel, and obtaining the coarse instance segmentation mask.
7. The real-time multi-instance segmentation method based on the Gaussian splash radiation field model according to claim 1, characterized in that: The step of inputting the coarse instance segmentation mask and the color image into a lightweight post-processing network for edge detection and region connectivity restoration to obtain a final two-dimensional instance segmentation mask comprises: Segmenting the coarse instance segmentation mask and the corresponding color image according to the bounding box of each instance region to obtain a local binary mask containing a single instance and a corresponding local color image; Inputting the local binary mask and the local color image into a lightweight post-processing network for edge detection and regional connectivity repair, and outputting a binary mask; The binary masks of each instance are reconstructed into a complete target 2D instance segmentation mask.
8. A real-time multi-instance segmentation device based on Gaussian splash radiation field model, characterized in that: The device comprises: The image segmentation module is used to render spatially continuously changing viewpoints based on a pre-trained 2D Gaussian splatter radiation field model to generate image sequences, and to obtain multi-view consistent 2D instance segmentation masks for the image sequences using a preset video segmentation model. a label generation module, configured to map the two-dimensional instance segmentation masks under multiple views into a three-dimensional space, and assign a unique instance label to each Gaussian basis element in the two-dimensional Gaussian splatter radiation field model using a preset voting strategy; An image rendering module is used to obtain a coarse instance segmentation mask and a corresponding color image with multi-view consistency under any viewing angle using a Gaussian splashing algorithm for a two-dimensional Gaussian splashing radiation field model with instance labels; The image fine segmentation module is used to input the coarse instance segmentation mask and the color image into a lightweight post-processing network for edge detection and regional connectivity restoration to obtain a target two-dimensional instance segmentation mask.
9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method according to any one of claims 1 to 7 is implemented when the processor executes the program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Three-dimensional point cloud segmentation method and device, equipment and storage medium
CN121259011A
Three-dimensional point cloud segmentation method, device and equipment and storage medium
CN121259011B
Scene segmentation and editing method and system based on three-dimensional Gaussian sputtering
CN121482352A
Ancient building three-dimensional component semantic decomposition method and system
CN122023746A
A method and system for semantic decomposition of three-dimensional components of ancient buildings
CN122023746B