A 3D gaussian monomerization and semanticization method based on a multi-modal visual large model

By combining multimodal visual large model technology, the challenges of semanticization and individualization in 3D Gaussian splashing were solved, achieving efficient 3D scene semantic understanding and individual segmentation, and improving computational efficiency and interactivity.

CN120164125BActive Publication Date: 2025-11-18ZHONGKE XINGTU DIGITAL EARTH HEFEI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510148084.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-11-18
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

3D Gaussian splashing suffers from problems such as insufficient accuracy and completeness of semantic tags, poor robustness of single-unit segmentation, and low computational efficiency in terms of semanticization and single-unit segmentation.

Method used

By combining multimodal visual large model technology, through data acquisition, preprocessing, 3D Gaussian model construction and training, semanticization and individualization, and by using deep learning and neural networks to optimize Gaussian parameters, semantic feature extraction and individual segmentation are achieved, supporting real-time rendering and interaction.

Benefits of technology

It improves the semantic and individualization accuracy of 3D Gaussian models, enhances computational efficiency, supports multiple input formats and interaction methods, and achieves efficient 3D scene management and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164125B_ABST
    Figure CN120164125B_ABST
Patent Text Reader

Abstract

The application discloses a 3D Gaussian monomerization and semanticization method based on a multimodal visual large model, and comprises the following steps: S1, data acquisition of a scene is performed, and the collected data is preprocessed, including semantic labeling of images and obtaining point cloud information of the scene; S2, a 3D Gaussian model is constructed, and training and optimization of Gaussian parameters are performed; S3, the 3D Gaussian model after the training and optimization of the parameters is subjected to semanticization; S4, the 3D Gaussian model is subjected to monomerization; and S5, real-time rendering and interaction are performed. The application realizes multi-granularity segmentation and adapts to various prompts, including text prompts, point selection, scribbling and 2D masks. The method can complete 3D segmentation within a few milliseconds, and provides a new tool for understanding and interaction of a 3D scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to a 3D Gaussian singletonization and semanticization method based on a multimodal visual large model. Background Technology

[0002] 3D Gaussian splashing, as an emerging technique in computer vision and graphics, has made significant progress in recent years. This technique utilizes Gaussian ellipsoid sets to model scenes and achieves efficient rendering by rasterizing the Gaussian ellipsoids to the image. Compared to traditional neural implicit representations (such as NeRF neural radiation fields), 3DGS avoids complex point sampling and neural network queries, thus achieving faster rendering speeds and better quality. However, despite the significant technological advancements in 3D Gaussian splashing, some challenges remain regarding semanticization and individualization.

[0003] (1) Accuracy and completeness of semantic labels: Achieving semanticization in 3D Gaussian splashes requires assigning accurate semantic labels to each Gaussian point. However, due to the complexity and diversity of 3D scenes, ensuring the accuracy and completeness of semantic labels is a challenging task.

[0004] (2) Robustness of single-object segmentation: In the process of single-object segmentation, 3D Gaussian splashing needs to accurately identify and segment each object. However, due to factors such as occlusion, overlap and lighting changes between objects, the robustness of single-object segmentation is still a problem that needs to be solved.

[0005] (3) Optimization of computational efficiency: In order to improve the rendering efficiency of 3D Gaussian splashing and reduce the consumption of computing resources, researchers need to continuously explore and optimize relevant algorithms and technologies. This includes optimizing the representation of Gaussian points, reducing unnecessary calculations, and using parallel computing techniques to improve computational efficiency.

[0006] To address the aforementioned problems and challenges, this invention combines multimodal visual large model technology with 3D Gaussian splashing to achieve semantic understanding of three-dimensional scenes. Summary of the Invention

[0007] To address the existing problems, this invention provides a 3D Gaussian singletonization and semanticization method based on a multimodal visual large model, the specific solution of which is as follows:

[0008] A 3D Gaussian singletonization and semanticization method based on a multimodal visual large model includes the following steps:

[0009] S1, to collect scene data and preprocess the collected data, including semantic annotation of images and obtaining point cloud information of the scene;

[0010] S2, construct a 3D Gaussian model, and perform training and optimization of Gaussian parameters;

[0011] S3, semanticize the 3D Gaussian model after training and parameter optimization;

[0012] S4, simplifies the 3D Gaussian model into individual units;

[0013] S5 enables real-time rendering and interaction.

[0014] Preferably, step S1 specifically includes the following steps:

[0015] S11 uses photographic equipment, including drones and mobile phones, to collect data about the scene;

[0016] S12, use a pre-trained semantic segmentation model to perform semantic annotation on the image;

[0017] S13 performs empty triangulation on the image to obtain the point cloud information of the scene, which is used as the input for 3D Gaussian splashing.

[0018] Preferably, step S2 specifically includes the following steps:

[0019] S21, Construct a 3D Gaussian model and introduce a semantic feature extraction module;

[0020] S22 uses a neural network training method based on stochastic gradient descent to optimize a 3D Gaussian model, and uses multi-channel supervision—color, depth, and semantics—to jointly optimize the Gaussian parameters.

[0021] Preferably, the semanticization in step S3 includes the following steps:

[0022] S31, Semantic Feature Extraction: Using deep learning techniques to extract semantic features from models or scenes, including the category, location, and shape of objects;

[0023] S32, Gaussian distribution modeling: The extracted semantic features are mapped onto a Gaussian distribution, and the distribution characteristics of semantic information are represented by the parameters of the Gaussian distribution—mean and variance.

[0024] S33, Semantic Information Fusion: Fusion of semantic information represented by Gaussian distribution with 3D models or scenes to achieve 3D Gaussian semanticization.

[0025] Preferably, step S4, which realizes the 3D Gaussian model singletonization, includes logical singletonization and physical singletonization;

[0026] Logical unitization: By vectorizing existing two-dimensional data, including road surfaces and building surfaces, and overlaying, highlighting, and attaching attribute information to the three-dimensional model, unitization is achieved.

[0027] Physical monolithization: Through 3D reconstruction, urban components including the ground, buildings, and roads are formed into a single entity that can be selected and separated.

[0028] Preferably, in step S5, a visibility-aware rendering algorithm is used to accelerate the training and rendering process, support anisotropic Gaussian, and provide a user interaction interface, allowing users to process elements in the scene in real time through control signals.

[0029] Preferably, the system based on any of the methods described above includes a data management module, a 3D Gaussian reconstruction module, a 3D Gaussian semantic unitization module, a task management module, a real-time rendering and display module, and a cloud service module.

[0030] The data management module is used for data input and data storage;

[0031] The 3D Gaussian reconstruction module is used for scene initialization, training of the 3D Gaussian model, and parameter optimization.

[0032] The 3D Gaussian semanticization and individualization module is used to semanticize and individualize the 3D Gaussian model;

[0033] The task management module is used to control the midpoint, start, pause, and continuation of tasks;

[0034] The real-time rendering and display module is used for real-time rendering, display, and interactive editing;

[0035] The cloud service module is used for online viewing and sharing of 3D reconstruction tasks and results.

[0036] The present invention also discloses a computer-readable storage medium storing a computer program, which, when executed, performs the method described in any of the above-mentioned embodiments.

[0037] The present invention also discloses a computer system including a processor, a storage medium storing a computer program, and the processor reading from the storage medium and running the computer program to perform the method described in any of the preceding claims.

[0038] The beneficial effects of this invention are as follows:

[0039] This invention is based on 3D Gaussian splashing and multimodal visual large model technology to achieve the individualization and semanticization of 3D Gaussian models. Specific technical features include:

[0040] (1) 3D Gaussian model generation technology supports multiple input formats such as images and videos, and 3D fully automatic reconstruction without manual intervention.

[0041] (2) Full-element semantic generation technology: Design a specific deep learning network structure to extract semantic features, and design a suitable Gaussian distribution modeling method to represent these features. Finally, effectively integrate the semantic information with the 3D model or scene to achieve semantic understanding capability.

[0042] (3) Individual object generation technology, which uses deep learning algorithms to extract and reconstruct each object in a 3D scene. At the same time, appropriate data structures and algorithms are designed to manage and manipulate these individual objects to achieve more refined 3D scene management and analysis.

[0043] (4) Supports smooth rendering of 3D Gaussian models and semantic models, including selection, rotation, movement, scaling, etc.

[0044] (5) It supports text prompts for interaction, which can realize the semantics of a certain type of single element in the 3D Gaussian model through text prompts.

[0045] (6) Supports semanticization of point selection in 3D Gaussian model. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 A diagram illustrating the singlet and semantic architecture of the Gaussian model;

[0048] Figure 2 This is a system principle framework diagram of the present invention;

[0049] Figure 3 This is a diagram illustrating the 3D Gaussian model reconstruction task in an embodiment of the present invention.

[0050] Figure 4 This is a smooth rendering diagram from an embodiment of the present invention;

[0051] Figure 5 This is a semantically encoded graph of all elements in an embodiment of the present invention;

[0052] Figure 6 This is a semantic diagram of the text prompt interaction method in an embodiment of the present invention;

[0053] Figure 7 In an embodiment of the present invention, a single-element diagram of a certain type of element or a single element is realized based on text prompts or click methods. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] 3D Gaussian splatter technology is an innovative 3D reconstruction and viewpoint synthesis method that uses a large number of 3D Gaussian functions (Splats) to represent 3D scenes and generate new views. To achieve higher semantic and individualization effects, this invention combines the latest research results on multimodal visual large models and proposes a complete solution.

[0056] The architecture diagram of the Gaussian model's individualization and semanticization in this invention is as follows: Figure 1 As shown, it includes a hardware layer, a scheduling layer, an algorithm layer, and a rendering layer.

[0057] The computing center hardware layer typically includes high-performance computing servers, accelerator cards (such as GPUs and FPGAs), and high-speed network devices. These hardware devices work together to provide powerful computing capabilities, meeting the demands of large-scale data computation and complex algorithms in 3D model production. For example, the rendering and simulation of 3D models requires processing large amounts of geometric, texture, and lighting data, all of which require high-performance computing equipment.

[0058] Scheduling Layer: This layer dynamically allocates computing resources through intelligent algorithms to ensure the efficient operation of 3D model production tasks. It rationally distributes computing resources based on task priority, resource requirements, and real-time load, avoiding resource idleness or overload. This dynamic allocation mechanism significantly improves production efficiency and shortens model production cycles, thereby meeting market demands for rapid 3D model delivery.

[0059] Algorithm layer:

[0060] a) First, a 3D Gaussian model is constructed, and a semantic feature extraction module is introduced. Then, neural network training methods such as stochastic gradient descent are used to optimize the model parameters, and the GS parameters are jointly optimized through multi-channel supervision (color, depth, semantics).

[0061] b) Semanticization of 3D Gaussian models, including semantic feature extraction, Gaussian distribution modeling, and semantic information fusion. Semantic feature extraction: Using techniques such as deep learning, semantic features are extracted from the model or scene, such as object category, location, and shape. Gaussian distribution modeling: The extracted semantic features are mapped onto a Gaussian distribution, and the distribution characteristics of the semantic information are represented by the parameters of the Gaussian distribution (such as mean and variance). Semantic information fusion: The semantic information represented by the Gaussian distribution is fused with the 3D model or scene to achieve 3D Gaussian semanticization.

[0062] c) 3D unitization is commonly used in fields such as 3D city modeling and Building Information Modeling (BIM) to achieve more refined 3D scene management and analysis. The main methods for achieving 3D unitization include logical unitization and physical unitization. Logical unitization: This involves vectorizing existing 2D data (such as road surfaces, building facades, etc.) and overlaying, highlighting, and attaching attribute information to the 3D model to achieve unitization. This method relies heavily on the accuracy of the 2D data and the precision of the 3D model. Physical unitization: This involves reconstructing urban components such as the ground, buildings, and roads into selectable and separable entities through 3D reconstruction. This method requires high-precision 3D reconstruction techniques and algorithms to ensure the accuracy and integrity of each unit object.

[0063] Rendering Layer: Implements efficient visibility-aware rendering algorithms, accelerating the training and rendering process. Supports anisotropic Gaussian rendering to improve rendering quality and real-time performance. Provides a user interface, allowing users to manipulate elements in the scene in real time through simple control signals.

[0064] This invention proposes a 3D interactive segmentation method that combines a 2D multimodal visual large model with 3D Gaussian splashing (3DGS). Utilizing carefully designed contrastive training, the 2D segmentation results are efficiently embedded into 3D Gaussian point features, thereby achieving multi-granularity segmentation and adapting to various prompts, including text prompts, clicks, doodles, and 2D masks. This method can complete 3D segmentation within milliseconds, providing a new tool for understanding and interacting with 3D scenes. Specifically, a 3D Gaussian singleton and semanticization method based on a multimodal visual large model includes the following steps:

[0065] S1 involves collecting scene data and preprocessing the collected data, including semantic annotation of images and obtaining point cloud information of the scene. Specifically, this includes the following steps:

[0066] S11 uses photographic equipment, including drones and mobile phones, to collect data about the scene;

[0067] S12, use a pre-trained semantic segmentation model to perform semantic annotation on the image;

[0068] S13 performs empty triangulation on the image to obtain the point cloud information of the scene, which is used as the input for 3D Gaussian splashing.

[0069] S2, construct a 3D Gaussian model and perform training and Gaussian parameter optimization. Specifically, this includes the following steps:

[0070] S21, Construct a 3D Gaussian model and introduce a semantic feature extraction module;

[0071] S22 uses a neural network training method based on stochastic gradient descent to optimize a 3D Gaussian model, and uses multi-channel supervision—color, depth, and semantics—to jointly optimize the Gaussian parameters.

[0072] S3. Semantize the 3D Gaussian model after training and parameter optimization. Semanticization includes the following steps:

[0073] S31, Semantic Feature Extraction: Using deep learning techniques to extract semantic features from models or scenes, including the category, location, and shape of objects;

[0074] S32, Gaussian distribution modeling: The extracted semantic features are mapped onto a Gaussian distribution, and the distribution characteristics of semantic information are represented by the parameters of the Gaussian distribution—mean and variance.

[0075] S33, Semantic Information Fusion: Fusion of semantic information represented by Gaussian distribution with 3D models or scenes to achieve 3D Gaussian semanticization.

[0076] S4, which performs individualization of the 3D Gaussian model.

[0077] Specifically, the realization of 3D Gaussian model singletonization includes logical singletonization and physical singletonization;

[0078] Logical unitization: By vectorizing existing two-dimensional data, including road surfaces and building surfaces, and overlaying, highlighting, and attaching attribute information to the three-dimensional model, unitization is achieved.

[0079] Physical monolithization: Through 3D reconstruction, urban components including the ground, buildings, and roads are formed into a single entity that can be selected and separated.

[0080] S5 enables real-time rendering and interaction.

[0081] The specific method for rendering interaction is as follows: use a visibility-aware rendering algorithm to accelerate the training and rendering process, support anisotropic Gaussian, provide a user interaction interface, and allow users to process elements in the scene in real time through control signals.

[0082] like Figure 2As shown, the present invention discloses a system based on the above method, including a data management module, a 3D Gaussian reconstruction module, a 3D Gaussian semantic unitization module, a task management module, a real-time rendering and display module, and a cloud service module.

[0083] The data management module is used for data input and storage. It ensures that needs related to data storage, organization, visualization, evaluation, task management, synchronous updates, and data security are met. Powered by AI technology, the data management module achieves greater intelligence and automation, providing more efficient and convenient data support for 3D model production.

[0084] The 3D Gaussian reconstruction module is used for scene initialization, training of the 3D Gaussian model, and parameter optimization. Specifically, the 3D Gaussian reconstruction module first recovers the 3D point cloud from the 2D image using Structure of Motion (SfM) technology. Then, each point is converted into a 3D Gaussian distribution containing position, color information, and covariance matrix. Next, a neural network is used to train the parameters of the Gaussian distribution, optimizing its position, color, and transparency attributes. Finally, the Gaussian distribution is converted into pixels that can be displayed on the screen through a rasterization process, thereby generating a realistic 3D rendered image or model.

[0085] The 3D Gaussian semantic characterization and individualization module is used to semanticize and individualize 3D Gaussian models. Specifically, firstly, each object or entity in the 3D scene is extracted and separated as a separate 3D Gaussian model to achieve individualization; then, deep learning or machine learning algorithms are used to semantically annotate these individual models, giving them specific semantic information, such as object category, attributes, etc., thereby achieving semanticization.

[0086] The task management module is used to control the midpoint, start, pause, and continuation of tasks.

[0087] The real-time rendering and display module is used for real-time rendering, display, and interactive editing. Specifically, this module is a component capable of high-quality, high-efficiency 3D scene rendering. It uses a 3D Gaussian distribution to represent the geometric and color information in the scene, and converts these Gaussian models into realistic images or animations through real-time rendering technology. It also supports user interaction and various display effects, providing visual support for fields such as virtual reality, augmented reality, and games.

[0088] The cloud service module is used for online viewing and sharing of 3D reconstruction tasks and results.

[0089] This invention supports the following functions:

[0090] 1. Supports image (JPG, PNG, and other common formats) and video input. Based on the StarMap Earth Brain 3D reconstruction algorithm and automatic computing power, the results are saved to the cloud or transferred to the local machine. For example... Figure 3 Reconstructing a task graph for a 3D Gaussian model.

[0091] 2. A platform that supports smooth rendering of 3D Gaussian and semantic models, including functions such as selection, rotation, movement, and scaling, providing high interactivity and flexibility. For example... Figure 4 This is a smooth rendering of the invention.

[0092] 3. 3D Gaussian splashing technology can be combined with multimodal visual large model technology to extract rich semantic information from input RGB images using a pre-trained semantic feature extractor. This semantic information can be fused with appearance features to generate spatially consistent high-dimensional semantic features, providing a foundation for subsequent 3D semantic mapping. By embedding semantic features into a 3D Gaussian distribution, 3D Gaussian splashing technology can achieve 3D semantic mapping. This mapping not only preserves the geometric structure of the scene but also contains rich semantic information, enabling objects in the scene to be accurately identified and classified. A 3D Gaussian full-element semantic representation is shown below. Figure 5 As shown. Extract a specific type of single element based on the text prompts, such as... Figure 6 As shown. Based on text prompts or click methods, you can achieve the individualization of a certain type of element or a single element, such as... Figure 7 As shown.

[0093] The present invention also discloses a computer-readable storage medium and a computer system, wherein the computer-readable storage medium stores a computer program, and the computer program, upon execution, performs the method described in any of the preceding claims. A computer system includes a processor and a storage medium, the storage medium storing a computer program, and the processor reading from the storage medium and running the computer program to perform the method described in any of the preceding claims.

[0094] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in a generalized manner in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.

[0095] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein can be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in alternatives, it may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.

[0096] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read and write information to / from the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and storage medium may reside as discrete components in the user terminal.

[0097] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functionality may be stored or transmitted as one or more instructions or code on or through a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium accessible to a computer. By way of example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Any connection is also legitimately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. As used in this article, disk and disc include compact discs (CDs), laser discs, optical discs, digital multi-purpose discs (DVDs), floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of these should also be included within the scope of computer-readable media.

[0098] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0099] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A 3D Gaussian singletonization and semanticization method based on a multimodal visual large model, characterized in that, Includes the following steps: S1, to collect scene data and preprocess the collected data, including semantic annotation of images and obtaining point cloud information of the scene; S2, construct a 3D Gaussian model, and perform training and optimization of Gaussian parameters; S3, semanticize the 3D Gaussian model after training and parameter optimization; S4, simplifies the 3D Gaussian model into individual units; S5 enables real-time rendering and interaction; Semanticization in step S3 includes the following steps: S31, Semantic Feature Extraction: Using deep learning techniques to extract semantic features from models or scenes, including the category, location, and shape of objects; S32, Gaussian distribution modeling: The extracted semantic features are mapped onto a Gaussian distribution, and the distribution characteristics of semantic information are represented by the parameters of the Gaussian distribution—mean and variance. S33, Semantic Information Fusion: Fusion of semantic information represented by Gaussian distribution with 3D models or scenes to achieve 3D Gaussian semanticization.

2. The method according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11 uses photographic equipment, including drones and mobile phones, to collect data about the scene; S12, use a pre-trained semantic segmentation model to perform semantic annotation on the image; S13 performs empty triangulation on the image to obtain the point cloud information of the scene, which is used as the input for 3D Gaussian splashing.

3. The method according to claim 1, characterized in that, Step S2 specifically includes the following steps: S21, Construct a 3D Gaussian model and introduce a semantic feature extraction module; S22 uses a neural network training method based on stochastic gradient descent to optimize a 3D Gaussian model, and uses multi-channel supervision—color, depth, and semantics—to jointly optimize the Gaussian parameters.

4. The method according to claim 1, characterized in that: Step S4, which implements the 3D Gaussian model singletonization, includes logical singletonization and physical singletonization. Logical unitization: By vectorizing existing two-dimensional data, including road surfaces and building surfaces, and overlaying, highlighting, and attaching attribute information to the three-dimensional model, unitization is achieved. Physical monolithization: Through 3D reconstruction, urban components including the ground, buildings, and roads are formed into a single entity that can be selected and separated.

5. The method according to claim 1, characterized in that, Step S5 uses a visibility-aware rendering algorithm to accelerate the training and rendering process, supports anisotropic Gaussian, and provides a user interaction interface, allowing users to process elements in the scene in real time through control signals.

6. A system based on the method described in any one of claims 1-5, characterized in that: It includes a data management module, a 3D Gaussian reconstruction module, a 3D Gaussian semantic unitization module, a task management module, a real-time rendering and display module, and a cloud service module; The data management module is used for data input and data storage; The 3D Gaussian reconstruction module is used for scene initialization, training of the 3D Gaussian model, and parameter optimization. The 3D Gaussian semanticization and individualization module is used to semanticize and individualize the 3D Gaussian model; The task management module is used to control the termination, start, pause, and continuation of tasks; The real-time rendering and display module is used for real-time rendering, display, and interactive editing; The cloud service module is used for online viewing and sharing of 3D reconstruction tasks and results.

7. A computer-readable storage medium, characterized in that: The medium contains a computer program, which, when run, performs the method as described in any one of claims 1 to 5.

8. A computer system, characterized in that: It includes a processor and a storage medium, on which a computer program is stored, and the processor reads from the storage medium and runs the computer program to perform the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for generating monomeric building from Mesh model of city scene

    CN115600307A

  • Live-action three-dimensional logic monomerization method and device and electronic equipment

    CN118015197A