Three-dimensional scene generation method and electronic device

By using an explicit-neural bimorphic Gaussian representation method and combining viewpoint direction features to optimize 3D scene generation, the problems of inconsistency in multiple viewpoints and slow optimization in existing technologies are solved, and efficient and stable 3D scene editing is achieved.

CN122454071APending Publication Date: 2026-07-24INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2026-06-25
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies suffer from inconsistencies in geometry and appearance from multiple perspectives during the process of editing two-dimensional images to three-dimensional enhancement. They also suffer from low efficiency in iterative optimization, and the editing results are prone to over-editing or under-editing. Furthermore, the optimization process is unstable.

Method used

An explicit-neural bimorphic Gaussian representation method is adopted. By inserting the target 3D model corresponding to the reference image into the specified position of the source 3D model, explicit Gaussian parameters and neural Gaussian parameters are extracted. Combined with the viewpoint direction features, rendering and optimization are performed to generate a 3D scene with consistency across multiple viewpoints.

Benefits of technology

It significantly improves the realism, stability, and convergence speed of 3D scene editing results, solves the problems of inconsistency between multiple perspectives and slow optimization, and reduces resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454071A_ABST
    Figure CN122454071A_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional scene generation method and an electronic device, relates to the technical field of computer vision, and comprises the following steps: inserting a target three-dimensional model corresponding to a reference image into a source three-dimensional model to construct a first three-dimensional scene, and extracting parameters of explicit Gaussian points from the first three-dimensional scene; combining three-dimensional coordinates of the Gaussian points with preset visual angle direction features to generate neural Gaussian parameters capable of modeling visual angle dependent geometry and appearance changes; jointly rendering the explicit Gaussian parameters and the neural Gaussian parameters to obtain a first image, rendering a second image by using the source three-dimensional model, and jointly optimizing the two types of Gaussian parameters by comparing differences between the two images; and finally reconstructing a second three-dimensional scene that fuses target objects, maintains structural consistency and has multi-view reality based on the updated explicit and neural Gaussian parameters, so that efficient, accurate and visually coherent three-dimensional editing is realized, and the problems of multi-view geometry and appearance inconsistency in a three-dimensional scene in the related art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to methods and electronic devices for generating 3D scenes. Background Technology

[0002] In related technologies, based on two-dimensional image editing, the editing is first completed on the reference view according to the text or image prompts, and then the two-dimensional editing result is improved to three-dimensional space through iterative optimization, so as to realize the three-dimensional editing that conforms to the text or reference image guidance in the specified area. However, this type of method faces significant challenges in practical applications: (1) Improving the image editing result to three dimensions will lead to inconsistencies in the geometry and appearance (including lighting, style, etc.) of the three-dimensional scene after editing; (2) The three-dimensional editing is completed through iterative optimization, which is time-consuming and not conducive to practical applications; and the optimization process is unstable, and the editing result is prone to problems such as over-editing or under-editing. Summary of the Invention

[0003] This application provides a method and electronic device for generating three-dimensional scenes, which at least solves the problem in the related art where the iterative optimization mechanism from image editing to 3D enhancement leads to inconsistencies in the geometry and appearance of the edited 3D scene from multiple perspectives, low efficiency of iterative optimization, and problems such as over-editing or under-editing of the editing results.

[0004] This application provides a method for generating a three-dimensional scene, comprising: acquiring a target three-dimensional model corresponding to a reference image and a source three-dimensional model to be edited; inserting the target three-dimensional model into the target editing position of the source three-dimensional model to generate a first three-dimensional scene; extracting explicit Gaussian parameters of explicit three-dimensional Gaussian points in the first three-dimensional scene, wherein the explicit Gaussian parameters include at least one of the three-dimensional coordinates, scale, color, and opacity of the explicit three-dimensional Gaussian points; determining neural Gaussian parameters based on the three-dimensional coordinates in the explicit Gaussian parameters and pre-set viewpoint direction features, wherein the neural Gaussian parameters include at least one of the geometric parameters and appearance parameters of the neural three-dimensional Gaussian points; rendering the explicit Gaussian parameters and neural Gaussian parameters to generate a first image; rendering the source three-dimensional model to generate a second image; updating the explicit Gaussian parameters and neural Gaussian parameters based on the first image and the second image; and generating a second three-dimensional scene using the updated explicit Gaussian parameters and neural Gaussian parameters.

[0005] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described three-dimensional scene generation methods.

[0006] This application also provides a non-volatile computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described three-dimensional scene generation methods.

[0007] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described three-dimensional scene generation methods.

[0008] This application embodiment constructs an initial first 3D scene by inserting the target 3D model corresponding to the reference image into a specified position of the source 3D model to be edited, and extracting parameters of explicit 3D Gaussian points from it, including at least one of 3D coordinates, scale, color, and opacity; then, combining the 3D coordinates of these Gaussian points with preset viewpoint direction features, generating neural Gaussian parameters capable of modeling viewpoint-dependent geometric and appearance changes; using the explicit Gaussian parameters and neural Gaussian parameters to jointly render a first image, and simultaneously rendering the source 3D model to generate a second image, and jointly optimizing the two types of Gaussian parameters by comparing the differences between the two images; finally, based on the updated explicit and neural Gaussian parameters, a second 3D scene is reconstructed that integrates the target object, maintains structural consistency, and has a multi-view realism, significantly improving the realism, stability, and convergence speed of the editing results, overcoming the shortcomings of existing technologies such as multi-view inconsistency, slow optimization, and high resource consumption. Attached Figure Description

[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart illustrating a three-dimensional scene generation method provided in this application embodiment; Figure 2 A flowchart of consistent 3D editing based on explicit-neural bimorphic Gaussian provided for embodiments of this application; Figure 3 This is a schematic diagram of an explicit-neural bistate Gaussian 3D scene representation provided in an embodiment of this application; Figure 4 This is a block diagram of a three-dimensional scene generation device provided in an embodiment of this application. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0012] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0013] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0014] Existing applications in embodied intelligence, autonomous driving, and world modeling rely heavily on the generation and editing of 3D scene models. Text- or image-based 3D scene editing, in particular, involves altering the geometry and appearance of existing 3D assets based on textual or reference image prompts to create diverse and high-quality 3D assets. Due to its ease of use, it has wide-ranging application value. Current methods, based on 2D image editing, iteratively optimize the results to 3D, but this often leads to inconsistencies between multi-view geometry and appearance (such as lighting and style). Furthermore, the iterative optimization process is time-consuming and unstable, frequently resulting in under-editing or over-editing. While some work has attempted to mitigate inconsistencies using attention mechanisms and joint color and geometric losses, their effectiveness and efficiency remain limited by scene complexity and the capabilities of the image editing models themselves, making it difficult to meet practical application needs.

[0015] This application proposes a consistent 3D scene editing and generation method based on dual-state Gaussian in the field of 3D vision, aiming to solve the problems of poor realism, low accuracy and slow editing speed in existing 3D scene editing technologies based on text or reference images.

[0016] The embodiments of this application provide a method for generating a three-dimensional scene. The method is described in detail below in conjunction with the execution flow of the three-dimensional scene generation method.

[0017] Specifically, Figure 1 This application provides a flowchart illustrating a method for generating a three-dimensional scene.

[0018] like Figure 1 As shown, the 3D scene generation method includes the following steps: In step S101, the target 3D model corresponding to the reference image and the source 3D model to be edited are obtained.

[0019] It is understood that the embodiments of this application can obtain the target 3D model corresponding to the reference image and the source 3D model to be edited, so as to quickly construct an initial 3D scene containing new content.

[0020] It should be noted that the target 3D model corresponding to the reference image refers to a 3D model that matches the semantic and geometric structure of the image content, obtained through 3D generation or reconstruction techniques based on a given reference image (such as a photograph or rendering containing a specific object, character, or scene).

[0021] In step S102, the target 3D model is inserted into the target editing position of the source 3D model to generate a first 3D scene. The explicit Gaussian parameters of the explicit 3D Gaussian points in the first 3D scene are extracted, wherein the explicit Gaussian parameters include at least one of the 3D coordinates, scale, color and opacity of the explicit 3D Gaussian points.

[0022] It is understood that the embodiments of this application can quickly construct a first 3D scene containing new content by precisely inserting the target 3D model into a specified editing position of the source 3D model, and extracting the parameters of the explicit 3D Gaussian points, including at least one of 3D coordinates, scale, color, and opacity, thereby preserving the geometric and appearance features of the inserted object in an efficient and structurally clear explicit representation, which is beneficial for subsequent fusion of view-related details and multi-view... Figure 1 Consistency optimization provides a high-quality initial representation.

[0023] It should be noted that the target editing location refers to the specific spatial region or semantic location in the source 3D model where the user wants to insert or modify content, and is used to guide the embedding operation of the target 3D model.

[0024] Specifically, such as Figure 2 As shown, this application, given a reference image, utilizes a single-image-based 3D generation method to obtain a target 3D model corresponding to the reference image. Given a source 3D model to be edited and specifying the insertion or replacement position (represented by a 3D editing box), the obtained 3D model corresponding to the reference image is inserted into the specified position in the source 3D model, thus obtaining a preliminary result of the edited 3D scene. This method avoids the iterative image editing process required in previous 3D editing methods, allowing for a rapid acquisition of preliminary editing results. However, this initial editing result is too rough, with inconsistencies in geometry and appearance from multiple perspectives.

[0025] Preliminary results of the edited 3D scene The results were then reconstructed and optimized. An explicit-neural bimorphic Gaussian representation of the 3D scene was used to improve the 3D scene representation capability, optimize editing results, and increase video memory utilization.

[0026] Among them, explicit three-dimensional Gaussian Each explicit 3D Gaussian point contains 8 parameters. , , Indicates the location of the center of the three-dimensional Gaussian. Representing isotropic scales, the viewpoint remains unchanged at each point, and the color is determined by... This indicates transparency Compared to the 59-dimensional parameters of traditional 3D Gaussian, explicit Gaussian has a significantly reduced number of parameters, thus reducing the amount of video memory required. However, explicit Gaussian inherently lacks view-dependent color and fine geometric representation (lacking anisotropic scale and rotation parameters), which can easily lead to oversmoothing representations of 3D scenes. Neural 3D Gaussian can address this problem.

[0027] In step S103, neural Gaussian parameters are determined based on the three-dimensional coordinates in the explicit Gaussian parameters and the preset viewpoint direction features. The neural Gaussian parameters include at least one of the geometric parameters and appearance parameters of the neural three-dimensional Gaussian points.

[0028] It is understood that the embodiments of this application can construct neural Gaussian parameters based on the three-dimensional coordinates in the explicit Gaussian parameters and combined with the pre-set viewpoint direction features to model the viewpoint-related geometric details and appearance changes. The neural Gaussian parameters include the geometric parameters, appearance parameters or a combination of both of the neural three-dimensional Gaussian points. This enhances the ability to express lighting, material and viewpoint-dependent effects while retaining the efficiency of explicit representation, and effectively improves the multi-view realism and visual consistency of the edited three-dimensional scene.

[0029] It should be noted that geometric parameters include opacity, scale, and rotation, while appearance parameters include color, without specific limitations.

[0030] In this embodiment, determining neural Gaussian parameters based on the three-dimensional coordinates in the explicit Gaussian parameters and a pre-set viewing direction includes: inputting the three-dimensional coordinates in the explicit Gaussian parameters and the pre-set viewing direction features into a neural Gaussian network; generating neural Gaussian parameters through the neural Gaussian network, wherein the neural Gaussian network includes a geometric prediction network and an appearance prediction network; generating a first feature through the geometric prediction network and the three-dimensional coordinates; predicting the geometric parameters of the neural three-dimensional Gaussian point based on the decoding result of the first feature; generating a second feature through the appearance prediction network, the three-dimensional coordinates, and the viewing direction features; and predicting the appearance parameters of the neural three-dimensional Gaussian point based on the decoding result of the second feature.

[0031] It is understood that the embodiments of this application can input the three-dimensional coordinates in the explicit Gaussian parameters and the pre-set view direction features into the geometric prediction network and appearance prediction network in the neural Gaussian network, respectively. The geometric prediction network generates a first feature based on the three-dimensional coordinates and decodes the geometric parameters of the neural three-dimensional Gaussian points. At the same time, the appearance prediction network fuses the three-dimensional coordinates and view direction features to generate a second feature and decodes the corresponding appearance parameters. This achieves decoupled modeling of geometric structure and view-related appearance, improving the consistency of three-dimensional scenes across multiple views while enhancing the ability to express details and the realism of rendering.

[0032] Specifically, such as Figure 2 As shown, the neural 3D Gaussian employs a grid-based neural radiation field structure. Optionally, multi-resolution hash-encoded neural radiation field technology can be used to simultaneously decompose the neural 3D Gaussian into a geometric neural 3D Gaussian and an appearance neural 3D Gaussian, thus balancing efficiency and performance. Specifically, the geometric neural 3D Gaussian and the appearance neural 3D Gaussian use two independent neural radiation fields to predict the geometric and appearance parameters of the neural 3D Gaussian, respectively.

[0033] Given the center coordinates of a 3D Gaussian Geometrically relevant Gaussian parameters: Opacity ,scale and rotation Since it is independent of the viewpoint, high-dimensional features can be directly extracted from the center point coordinates. It is stored using multi-resolution hash encoding, and then the high-dimensional features are directly decoded. The predicted value is obtained using the following formula:

[0034]

[0035] in, and Geometric feature encoding network and decoding network The parameters all use a multilayer perceptron as the network structure. Let these represent the center coordinates, opacity, scale, and rotation of the i-th Gaussian. This is the high-dimensional feature extracted from the coordinates of the i-th three-dimensional Gaussian center point.

[0036] Appearance-related features are viewpoint-dependent. To model the variation of 3D Gaussian color with viewpoint, the viewpoint direction is introduced. As additional input, the viewpoint orientation feature is encoded as follows: The specific calculations are as follows:

[0037]

[0038]

[0039] in, , The camera origin, Indicates appearance characteristics, PE represents the direction feature of the viewpoint, and PE represents the direction encoding function. , This indicates that the view direction is obtained by dire encoding the view direction. , Let i be the center coordinates of the i-th Gaussian. Appearance feature encoding network The parameters, d represents the viewpoint direction. This is a feature encoder.

[0040] The 3D Gaussian color decoding calculation is as follows:

[0041] in, This represents the concatenation of eigenvectors. For appearance feature decoding network The parameters all use a multilayer perceptron as the network structure. Indicates appearance characteristics, Indicates the characteristics of the viewpoint direction.

[0042] Using only explicit Gaussian representation of a 3D scene degenerates into a primitive 3D Gaussian representation. However, the primitive 3D Gaussian representation requires a large number of Gaussian spheres to accurately represent the 3D scene, resulting in high storage consumption. Furthermore, the large number of Gaussian spheres written and written during training causes insufficient bandwidth, further reducing editing speed. An explicit-neural bimorphic Gaussian representation is adopted, where explicit Gaussian handles high-frequency local details, and neural Gaussian fills in / corrects low-frequency or holed regions. In addition, the hyperparameters of neural Gaussian, such as mesh resolution and number of channels, are fixed and do not change with the number of explicit Gaussian spheres. It can also significantly reduce the amount of explicit Gaussian data, while using neural Gaussian to compensate for the shortcomings of explicit Gaussian in terms of appearance and geometric representation, thus achieving a balance between multiple requirements for representation accuracy, storage consumption, and editing speed.

[0043] It should be noted that by using this explicit-neural bimorphic Gaussian representation, after training, we can obtain: the explicit Gaussian is responsible for representing the high-frequency part of the entire 3D scene, and the neural Gaussian is responsible for representing the low-frequency or hollow areas of the 3D scene (here, hollow areas refer to 3D scene areas that are not represented by explicit Gaussian).

[0044] In step S104, the explicit Gaussian parameters and neural Gaussian parameters are rendered to generate a first image, the source 3D model is rendered to generate a second image, the explicit Gaussian parameters and neural Gaussian parameters are updated based on the first image and the second image, and a second 3D scene is generated using the updated explicit Gaussian parameters and neural Gaussian parameters.

[0045] It is understood that the embodiments of this application can update these parameters by rendering the difference between a first image generated using explicit Gaussian parameters and neural Gaussian parameters and a second image directly rendered from the source 3D model. This makes the 3D scene constructed based on the updated parameters more accurately approximate the appearance of the original 3D model, thereby improving the quality and visual effect of 3D scene reconstruction while maintaining efficient computation. It fully utilizes the advantages of statistical parameterization methods and modern neural network technology to achieve high-quality and fast 3D content generation.

[0046] Specifically, it integrates explicit Gaussian and neural Gaussian bistate representations for rendering. Specifically, given an explicit Gaussian representation... Based on explicit Gaussian position The opacity is predicted by processing the neural Gaussian using formulas (1)-(6). ,scale Rotation and color The explicit Gaussian and neural Gaussian parameter information is fused and input into the original 3D Gaussian splash rasterization to obtain the rendered free-viewpoint image. The specific calculation is as follows:

[0047] in, This represents a multilayer perceptron. , , These are the opacity, color, and scale of explicit Gaussian. , , These are the opacity, color, and scale predicted by the neural Gaussian method. , s and s represent the opacity, appearance, and scale after fusing explicit Gaussian parameters and neural Gaussian parameters, respectively.

[0048] The color of any pixel in the rendered free-view image is then:

[0049]

[0050] in, Let represent the opacity and color of the i-th Gaussian, respectively. and These are the three-dimensional Gaussian mean and covariance, respectively. Let be the intersection point of the ray and the i-th three-dimensional Gaussian ray. Let G represent the cumulative transparency of the i-th Gaussian function, and let G represent the three-dimensional Gaussian function. This represents the result obtained by substituting the intersection point x of the ray and the i-th three-dimensional Gaussian ray into the function G. This represents the opacity weight of the i-th Gaussian.

[0051] In this embodiment of the application, updating explicit Gaussian parameters and neural Gaussian parameters based on a first image and a second image includes: performing a Fourier transform on the first image and the second image to the frequency domain; calculating a first frequency loss and a second frequency loss based on the first image and the second image in the frequency domain, and calculating an image loss between the first image and the second image in the frequency domain; updating explicit Gaussian parameters based on the first frequency loss and the image loss; and updating neural Gaussian parameters based on the second frequency loss and the image loss.

[0052] It is understood that the embodiments of this application can transform the first image and the second image to the frequency domain, calculate the first frequency loss and the second frequency loss corresponding to the low frequency and high frequency components respectively, and construct a multi-scale optimization target by combining the overall image loss. The low frequency information mainly guides the update of the neural Gaussian parameters to stabilize the geometric structure and overall appearance, while the high frequency information drives the adjustment of the explicit Gaussian parameters to restore details and edge sharpness. This achieves decoupling and collaborative optimization of explicit and neural Gaussian parameters, which enhances the geometric accuracy, texture details and multi-view consistency of the 3D scene while improving the fidelity of the rendered image.

[0053] It should be noted that the first frequency loss refers to the measurement of reconstruction error of low-frequency components (such as large-scale structures, smooth regions, and overall illumination) in the frequency domain, mainly focusing on spectral components with radii smaller than the target cutoff frequency; the second frequency loss refers to the measurement of reconstruction error of high-frequency components (such as edges, textures, and detailed structures) in the frequency domain, corresponding to spectral components with radii greater than or equal to the target cutoff frequency; the image loss can refer to the difference between the first image (generated by explicit and neural Gaussian joint rendering) and the second image (generated by rendering the source 3D model) calculated in the spatial domain. This loss reflects the overall reconstruction error and is used to constrain the edited scene to maintain visual consistency with the original scene.

[0054] Specifically, the rendering image loss is split into low-frequency and high-frequency components, and explicit Gaussian and neural Gaussian are optimized respectively. Specifically, the source 3D model and the initial 3D editing results are... Multi-view rendering from the same perspective is performed. An edit mask (obtained by projecting a manually specified edit box onto the same viewpoint) is used to blend the rendered images, resulting in a pseudo-ground image. .

[0055] Render the image With pseudo-true value image Perform FFT transformation to the frequency domain:

[0056] in, and These represent the images to be rendered. With pseudo-true value image The value after transformation to the frequency domain This is a function that transforms an image to the frequency domain.

[0057] In this embodiment of the application, calculating the first frequency loss and the second frequency loss based on the first image and the second image in the frequency domain includes: the number of updates based on explicit Gaussian parameters and neural Gaussian parameters; determining the target cutoff frequency based on the number of updates; identifying the spectral coordinates of the first image and the second image in the frequency domain; calculating the radius based on the spectral coordinates; and determining the first frequency loss and the second frequency loss based on the target cutoff frequency and the radius.

[0058] It is understood that the embodiments of this application can dynamically determine the target cutoff frequency based on the number of updates of explicit Gaussian parameters and neural Gaussian parameters, and calculate the radius of each frequency component based on the spectral coordinates of the first image and the second image in the frequency domain. Then, the low-frequency and high-frequency regions are divided with the target cutoff frequency as the boundary, and the first frequency loss for optimizing neural Gaussian parameters and the second frequency loss for optimizing explicit Gaussian parameters are calculated respectively. In this way, the collaborative recovery of low-frequency structure and high-frequency details is adaptively guided during the training process, effectively improving the stability, convergence speed and multi-scale visual fidelity of 3D scene reconstruction.

[0059] It should be noted that the target cutoff frequency refers to the frequency threshold used to distinguish between low-frequency and high-frequency components during frequency domain optimization. It is usually represented by the distance from the center (zero frequency) to a certain radius in the Fourier spectrum. This threshold is not fixed, but is dynamically adjusted according to the number of updates of the explicit Gaussian parameters and the neural Gaussian parameters.

[0060] Specifically, setting the dynamic cutoff frequency :

[0061] in, This is the current iteration number; the total number of iterations is set to [number]. Base cutoff frequency and γ is a hyperparameter, and γ is an exponent.

[0062] In the frequency domain, the low-frequency and high-frequency regions of the image are respectively labeled as binary masks. , For spectral coordinates Define radius ,but:

[0063] in, Let r be the current iteration number and r be the radius. The cutoff frequency, For low-frequency masking, For high-frequency mask, This is an indicator function that takes the value 1 when the condition is met, and 0 otherwise.

[0064] Calculate frequency loss, low-frequency loss High-frequency loss Update the neural Gaussian parameters and the explicit Gaussian parameters separately:

[0065]

[0066] in, and These represent the images to be rendered. With pseudo-true value image The value after transformation to the frequency domain Indicates bitwise AND. express Select the red (R), green (G), and blue (B) color channels in sequence. For low-frequency masking, It is a high-frequency mask.

[0067] In this embodiment of the application, before updating the explicit Gaussian parameters and the neural Gaussian parameters based on the first image and the second image, the process includes: calculating the rendering loss based on the first image and the second image; calculating the total loss based on the first frequency loss, the second frequency loss, and the rendering loss; if the total loss is greater than a preset loss threshold, then updating the explicit Gaussian parameters and the neural Gaussian parameters based on the first image and the second image; otherwise, stopping the updating of the explicit Gaussian parameters and the neural Gaussian parameters.

[0068] The pre-set loss threshold can be set according to actual needs without specific limitations.

[0069] It is understood that, in the embodiments of this application, before updating the explicit Gaussian parameters and the neural Gaussian parameters, the rendering loss between the first image and the second image can be calculated first, and the total loss can be constructed by combining the first frequency loss and the second frequency loss. By comparing the total loss with a preset loss threshold, parameter updates are only performed when the total loss exceeds the threshold, otherwise the optimization process is terminated, thereby avoiding overfitting or redundant iterations, and significantly improving training efficiency and convergence stability while ensuring editing quality.

[0070] Specifically, the total loss function is:

[0071]

[0072] in, , and For hyperparameters, and These represent low-frequency loss and high-frequency loss, respectively. Represents the loss function in the image domain. To render the image, This is a pseudo-true value image. To render the image With pseudo-true value image The squared error.

[0073] The optimization is performed in the frequency domain space using frequency domain masks, without the need for additional networks. Frequency-guided decoupling optimization allows the neural Gaussian to focus on smooth areas of the image, while the explicit Gaussian focuses on details, significantly improving training convergence speed and making editing and optimization more efficient.

[0074] In this embodiment of the application, generating a second three-dimensional scene using updated explicit Gaussian parameters and neural Gaussian parameters includes: generating bi-state Gaussian parameters based on explicit Gaussian parameters and neural Gaussian parameters; generating a rendered image of at least one target viewpoint based on the bi-state Gaussian parameters; and fusing the rendered image of at least one target viewpoint to generate a second three-dimensional scene.

[0075] It is understood that when the embodiments of this application generate a second three-dimensional scene using updated explicit Gaussian parameters and neural Gaussian parameters, they first combine the two to generate bi-state Gaussian parameters, then create rendered images for at least one target viewpoint based on these parameters, and finally fuse these rendered images to construct a complete second three-dimensional scene. This process effectively improves the accuracy and visual effect of three-dimensional scene reconstruction, while enhancing the consistency and realism of observing the scene from different viewpoints.

[0076] In this embodiment of the application, after fusing rendered images from at least one target viewpoint to generate a second three-dimensional scene, the method further includes: obtaining a third image of the second three-dimensional scene rendered from a reference viewpoint; obtaining a fourth image of the second three-dimensional scene rendered from a target viewpoint; calculating the amount of color change of the second three-dimensional scene between the target viewpoint and the reference viewpoint based on the third image and the fourth image; and optimizing the second three-dimensional scene based on the amount of color change.

[0077] It is understood that, after generating a second three-dimensional scene by fusing multi-view rendered images, this application embodiment can obtain the third and fourth images of the scene under the reference view and the target view, calculate the amount of color change between the two, and optimize the second three-dimensional scene accordingly. This effectively suppresses the problem of inconsistent lighting, color tone or material caused by view switching, and further improves the appearance continuity and overall realism of the scene under multiple views.

[0078] Specifically, this application can be edited based on the multi-view consistency of residual prediction: The geometry of a 3D scene remains unchanged across different viewpoints, but colors can vary due to lighting and other factors. Traditional 3D editing methods optimize both geometry and appearance simultaneously, making it impossible to maintain geometric consistency across different viewpoints. The dual-state Gaussian proposed in this application provides a foundation for solving this problem. Specifically, it sets the explicit Gaussian to remain unchanged, while simultaneously keeping the neural Gaussian geometry prediction network unchanged, optimizing only the color prediction network, and employing a residual prediction method. This involves selecting the front view as the reference viewpoint and retaining the front view's neural Gaussian color prediction. Unchanged, using a neural Gaussian color prediction network To predict the color residuals from other viewpoints relative to the positive viewpoint, the rendering calculation is as follows:

[0079]

[0080] in, This represents a multilayer perceptron. This represents the predicted color residual. This represents a color prediction network, which consists of a multilayer perceptron. This represents the parameters of the color prediction network. For explicit Gaussian color, The color is a neuro-Gaussian color. For appearance characteristics, To extract high-dimensional features from the coordinates of the i-th 3D Gaussian center point, is the neural Gaussian appearance residual term, and c is the output rendering color.

[0081] In this embodiment of the application, before generating the second three-dimensional scene using the updated explicit Gaussian parameters and neural Gaussian parameters, the method includes: performing Gaussian sampling on the first image and the second image; obtaining the rendering contribution value and the number of intersecting tiles of the candidate Gaussian points obtained from the Gaussian sampling; calculating the activity of the candidate Gaussian points based on the rendering contribution value and the number of intersecting tiles; and replacing the candidate Gaussian points with one of the explicit three-dimensional Gaussian points and the neural three-dimensional Gaussian points based on the activity of the candidate Gaussian points.

[0082] It is understood that, before generating the second three-dimensional scene, this application embodiment can obtain the rendering contribution value of candidate Gaussian points and the number of tiles they cover by performing Gaussian sampling on the first and second images, and calculate the activity of each candidate point accordingly. Then, based on the activity, the point is dynamically selected to be represented as an explicit three-dimensional Gaussian point or a neural three-dimensional Gaussian point, thereby optimizing the allocation of computing resources and improving rendering efficiency and adaptability of scene representation while ensuring the expression of details in important areas.

[0083] It should be noted that this application can add a rendering pruning technique based on activity and semantic guidance on the basis of explicit-neural bimorphic Gaussian coupling rendering, thereby reducing video memory usage and improving video memory utilization and editing efficiency.

[0084] Specifically, the activity level is calculated as follows:

[0085] in, This is an indicator function that takes the value 1 when a certain condition is met, and 0 otherwise. This indicates the number of pixels that the Gaussian point is responsible for, measuring the degree of contribution of the Gaussian point to the pixel value. All Gaussian points contributed middle, Greater than the threshold When, then the pixel is considered From Gauss point Responsible, Indicates the point with Gaussian The number of tiles where the ellipses intersect is obtained from the intersection stage of the Gaussian rasterizer preprocessing, which measures the computational cost of calculating this Gaussian. Let be the cumulative transparency factor of the i-th Gaussian point at pixel j. Let be the opacity of the i-th Gaussian point.

[0086] In this embodiment of the application, replacing a candidate Gaussian point with one of an explicit three-dimensional Gaussian point and a neural three-dimensional Gaussian point based on the activity level of the candidate Gaussian point includes: obtaining a preset activity threshold; if the activity level of the candidate Gaussian point is greater than or equal to the activity threshold, then replacing the candidate Gaussian point with an explicit three-dimensional Gaussian point; if the activity level of the candidate Gaussian point is less than the activity threshold, then replacing the candidate Gaussian point with a neural three-dimensional Gaussian point.

[0087] The activity threshold can be set according to actual needs, without specific limitations.

[0088] It is understood that the embodiments of this application can distinguish candidate Gaussian points by setting an activity threshold, replace points with high activity that contribute significantly to rendering with explicit 3D Gaussian points to preserve high-frequency details and geometric accuracy, and replace points with low activity with neural 3D Gaussian points to utilize neural networks to model low-frequency structures and view-related appearances, thereby achieving efficient and high-quality hybrid 3D representation while controlling video memory overhead.

[0089] Specifically, the activity-based rendering pruning technique works as follows: During training, Gaussians with activity levels above a threshold are used as explicit 3D Gaussians and rendered using a Gaussian rasterizer and a neural Gaussian. Gaussians with activity levels below the threshold are automatically collapsed into individual neural Gaussians for rendering. Specifically, a Gaussian Activity (GA) is maintained for each Gaussian, measuring the ratio of a Gaussian point's contribution to a pixel value to its computational cost. Higher activity levels indicate that the explicit Gaussian is more important and should be retained.

[0090] When activity When the activity level exceeds a threshold, the explicit Gaussian resides in memory and participates in the coupled rendering of the Gaussian rasterizer and the neural Gaussian. When the activity level is below the threshold, the explicit Gaussian is replaced in situ with the neural Gaussian, i.e., a single entry in the hash table, and the explicit Gaussian parameters are released. By dynamically using explicit-neural bimorphic Gaussian during training and dynamically switching storage modes, memory utilization is improved and editing is accelerated. At the same time, the use of explicit-neural bimorphic Gaussian can significantly reduce the number of explicit Gaussian parameters and the amount of data required, thereby significantly reducing memory usage.

[0091] In this embodiment of the application, before generating a three-dimensional scene using the updated explicit Gaussian parameters and neural Gaussian parameters, the method further includes: segmenting the first three-dimensional scene into at least one region to be edited; determining the corresponding semantic priority based on the at least one region to be edited; removing explicit Gaussians with semantic priorities lower than a preset semantic priority level, and retaining explicit Gaussians with semantic priorities greater than or equal to the preset semantic priority level.

[0092] The preset semantic priority level can be set according to actual needs without specific limitations.

[0093] It is understood that, before generating a 3D scene using the updated explicit Gaussian parameters and neural Gaussian parameters, the embodiments of this application can segment the first 3D scene to obtain multiple regions to be edited, assign semantic priorities to each region, and then remove explicit Gaussian points with semantic priorities lower than a preset level, retaining only explicit Gaussian points with high semantic importance. This effectively focuses computing resources on key semantic regions, reduces redundant representations, and improves editing accuracy and rendering efficiency.

[0094] It should be noted that, based on the semantic priority tags, we can set which Gaussian sigmas can reside in video memory. During rendering, Gaussian sigmas that reside in video memory can be directly used for rendering, but this will increase the video memory load. Therefore, based on different semantic priorities, we can set whether to reside in video memory, prioritize putting them in video memory, or not put them in video memory at all.

[0095] Specifically, the semantically guided Gaussian rendering technique utilizes a semantic segmentation method to assign semantic priority labels to explicit Gaussian expressions, assigning different semantic priority labels to foreground objects and backgrounds based on whether they are within the editing area. Specifically, foreground objects within the editing area are given the highest semantic priority, forcing their explicit Gaussian expressions to reside permanently in video memory to ensure a high interactive frame rate; the background within the editing area is set to the second highest priority, with its explicit Gaussian expressions written to video memory preferentially; and foreground and background objects outside the editing area are set to the lowest priority, not residing permanently in video memory, to reduce video memory burden and improve video memory utilization. Simultaneously, this approach allows the editing optimization process to focus on optimizing the editing area, accelerating editing optimization and improving editing efficiency.

[0096] During training, this application first uses formulas (1)-(4) to render and optimize the loss of the 3D scene. When the number of iterations meets a certain threshold, formula (5) is used to correct the rendering, while other steps remain unchanged. During inference, the rendering method corrected by formula (5) is also used.

[0097] Initial results of 3D editing Issues such as inconsistencies in geometry / appearance from multiple perspectives and inaccurate editing exist. These are addressed and optimized using bi-state Gaussian reconstruction. During training, the initial 3D editing results are used... Initial explicit Gaussians are obtained using the farthest point sampling method, typically 5k; during iteration, the number is adaptively increased or decreased, with a cloning or splitting operation performed every 300 iterations to split Gaussians with large gradients; Gaussian pruning is also performed simultaneously: opacity... Gaussians below a threshold or with scales that are too large or too small are directly deleted. The total number of explicit Gaussians varies with scene complexity, generally not exceeding 40k. During rendering, a rendering pruning technique based on activity and semantic guidance is used, adaptively employing explicit-neural Gaussian coupled rendering, while simultaneously utilizing a frequency-domain guided decoupling optimization strategy for gradient backpropagation optimization. After a certain number of training iterations, a 3D scene editing strategy based on residual prediction is adopted. The explicit Gaussian color and geometry-related parameters are kept unchanged in the editing area, and only the neural Gaussian network is trained. Geometry and appearance are optimized through residual prediction, achieving consistent appearance and geometry editing across multiple viewpoints.

[0098] In summary, this application proposes a consistent 3D scene editing method based on dual-state Gaussian, achieving consistent representation and optimization of the edited 3D scene from multiple perspectives through coupled representation and decoupled optimization. Simultaneously, by utilizing dynamic storage, residual prediction, and hybrid pruning techniques, it improves editing efficiency while achieving high-quality 3D editing based on reference images. This method can be used for efficient editing of existing 3D scenes, obtaining more diverse 3D scenes, and significantly saving modeling time and economic costs. It can support the technological development of server products and expand applications in the field of 3D content generation and editing. Furthermore, it is applicable to cutting-edge fields such as embodied intelligence, autonomous driving, and world model generation and editing, possessing broad market prospects.

[0099] This application proposes a 3D scene representation method based on explicit-neural bistate Gaussian. It uses explicit-neural bistate Gaussian coupling to represent the 3D scene and dynamically switches according to the Gaussian activity. At the same time, the video memory usage adaptively scales with the Gaussian switching, which improves the 3D scene representation capability after editing, while reducing video memory usage and improving editing efficiency.

[0100] This application proposes a multi-view consistency editing method based on residual prediction. By keeping the explicit Gaussian invariant, the residual predicts the appearance changes of the neural Gaussian at different views, and utilizes geometric and appearance decoupling prediction to reduce the scale of scene optimization after editing, while increasing the stability of the editing optimization process and accelerating editing.

[0101] This application proposes a frequency-domain guided coupled training strategy, which decomposes the training loss into low-frequency and high-frequency losses through frequency-domain masking, realizes explicit and neural Gaussian decoupling optimization, ensures consistent transition of geometry and appearance of the edited 3D scene, does not lose details, and significantly improves the training convergence speed.

[0102] This application proposes a method for rapidly obtaining pre-editing results of 3D scenes. By utilizing image-based 3D generation, it can quickly complete the insertion or replacement editing of 3D scenes based on reference images, avoid iterative image editing, and quickly obtain preliminary results of 3D scene editing.

[0103] The three-dimensional scene generation method according to the embodiments of this application constructs an initial first three-dimensional scene by inserting the target three-dimensional model corresponding to the reference image into a specified position of the source three-dimensional model to be edited, and extracting parameters of explicit three-dimensional Gaussian points from it, including at least one of three-dimensional coordinates, scale, color, and opacity; then, combining the three-dimensional coordinates of these Gaussian points with preset viewpoint direction features, generating neural Gaussian parameters that can model viewpoint-dependent geometric and appearance changes; using the explicit Gaussian parameters and neural Gaussian parameters to jointly render a first image, and simultaneously rendering the source three-dimensional model to generate a second image, and jointly optimizing the two types of Gaussian parameters by comparing the differences between the two images; finally, based on the updated explicit and neural Gaussian parameters, a second three-dimensional scene that integrates the target object, maintains structural consistency, and has multi-view realism is reconstructed, significantly improving the realism, stability, and convergence speed of the editing results, and overcoming the defects of inconsistency in multiple views, slow optimization, and high resource consumption in the prior art.

[0104] The embodiments of this application provide a three-dimensional scene generation apparatus, and the apparatus is described in detail in conjunction with the execution flow of the three-dimensional scene generation method.

[0105] like Figure 4 As shown, the 3D scene generation device 10 includes: an acquisition module 100, an insertion module 200, a processing module 300, and a rendering module 400.

[0106] The module 100 is used to acquire the target 3D model corresponding to the reference image and the source 3D model to be edited; the insertion module 200 is used to insert the target 3D model into the target editing position of the source 3D model to generate a first 3D scene, and extract the explicit Gaussian parameters of the explicit 3D Gaussian points in the first 3D scene, wherein the explicit Gaussian parameters include at least one of the 3D coordinates, scale, color and opacity of the explicit 3D Gaussian points; the processing module 300 is used to determine the neural Gaussian parameters according to the 3D coordinates in the explicit Gaussian parameters and the pre-set viewpoint direction features, wherein the neural Gaussian parameters include at least one of the geometric parameters and appearance parameters of the neural 3D Gaussian points; the rendering module 400 is used to render the explicit Gaussian parameters and the neural Gaussian parameters to generate a first image, render the source 3D model to generate a second image, update the explicit Gaussian parameters and the neural Gaussian parameters according to the first image and the second image, and generate a second 3D scene using the updated explicit Gaussian parameters and the neural Gaussian parameters.

[0107] In this embodiment, the processing module 300 is further configured to: input the three-dimensional coordinates and pre-set viewpoint direction features from the explicit Gaussian parameters into a neural Gaussian network, generate neural Gaussian parameters through the neural Gaussian network, wherein the neural Gaussian network includes a geometric prediction network and an appearance prediction network, generate a first feature through the geometric prediction network and the three-dimensional coordinates, and predict the geometric parameters of the neural three-dimensional Gaussian point based on the decoding result of the first feature; generate a second feature through the appearance prediction network, the three-dimensional coordinates and the viewpoint direction features, and predict the appearance parameters of the neural three-dimensional Gaussian point based on the decoding result of the second feature.

[0108] In this embodiment, the rendering module 400 is further configured to: perform Fourier transform on the first image and the second image to the frequency domain; calculate the first frequency loss and the second frequency loss based on the first image and the second image in the frequency domain, and calculate the image loss between the first image and the second image in the frequency domain; update the explicit Gaussian parameters based on the first frequency loss and the image loss; and update the neural Gaussian parameters based on the second frequency loss and the image loss.

[0109] In this embodiment, the rendering module 400 is further configured to: determine the target cutoff frequency based on the number of updates of explicit Gaussian parameters and neural Gaussian parameters; identify the spectral coordinates of the first image and the second image in the frequency domain; calculate the radius based on the spectral coordinates; and determine the first frequency loss and the second frequency loss based on the target cutoff frequency and the radius.

[0110] In this embodiment, the rendering module 400 is further configured to: calculate rendering loss based on the first image and the second image; calculate total loss based on the first frequency loss, the second frequency loss and the rendering loss; if the total loss is greater than a preset loss threshold, update the explicit Gaussian parameters and the neural Gaussian parameters based on the first image and the second image; otherwise, stop updating the explicit Gaussian parameters and the neural Gaussian parameters.

[0111] In this embodiment, the rendering module 400 is further configured to: generate dual-state Gaussian parameters based on explicit Gaussian parameters and neural Gaussian parameters; generate a rendered image of at least one target viewpoint based on the dual-state Gaussian parameters; and fuse the rendered images of at least one target viewpoint to generate a second three-dimensional scene.

[0112] In this embodiment of the application, it further includes: an optimization module, configured to, after fusing rendered images from at least one target viewpoint to generate a second three-dimensional scene, obtain a third image of the second three-dimensional scene rendered from a reference viewpoint; obtain a fourth image of the second three-dimensional scene rendered from a target viewpoint; calculate the amount of color change of the second three-dimensional scene between the target viewpoint and the reference viewpoint based on the third image and the fourth image; and optimize the second three-dimensional scene based on the amount of color change.

[0113] In this embodiment, the system further includes: a replacement module, configured to perform Gaussian sampling on the first image and the second image before generating the second three-dimensional scene using the updated explicit Gaussian parameters and neural Gaussian parameters; obtain the rendering contribution value and the number of intersecting tiles of the candidate Gaussian points obtained by Gaussian sampling; calculate the activity of the candidate Gaussian points based on the rendering contribution value and the number of intersecting tiles; and replace the candidate Gaussian points with one of the explicit three-dimensional Gaussian points and the neural three-dimensional Gaussian points based on the activity of the candidate Gaussian points.

[0114] In this embodiment of the application, the replacement module is further configured to: obtain a preset activity threshold; if the activity of a candidate Gaussian point is greater than or equal to the activity threshold, then replace the candidate Gaussian point with an explicit three-dimensional Gaussian point; if the activity of a candidate Gaussian point is less than the activity threshold, then replace the candidate Gaussian point with a neural three-dimensional Gaussian point.

[0115] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0116] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the three-dimensional scene generation method.

[0117] Embodiments of this application also provide a non-volatile computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the three-dimensional scene generation method at runtime.

[0118] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0119] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the three-dimensional scene generation method.

[0120] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the three-dimensional scene generation method.

[0121] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0122] The above provides a detailed description of a three-dimensional scene generation method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for generating a three-dimensional scene, characterized in that, include: Obtain the target 3D model corresponding to the reference image and the source 3D model to be edited; The target 3D model is inserted into the target editing position of the source 3D model to generate a first 3D scene. The explicit Gaussian parameters of the explicit 3D Gaussian points in the first 3D scene are extracted, wherein the explicit Gaussian parameters include at least one of the 3D coordinates, scale, color and opacity of the explicit 3D Gaussian points. The neural Gaussian parameters are determined based on the three-dimensional coordinates in the explicit Gaussian parameters and the pre-set viewpoint direction features, wherein the neural Gaussian parameters include at least one of the geometric parameters and appearance parameters of the neural three-dimensional Gaussian points; The explicit Gaussian parameters and the neural Gaussian parameters are rendered to generate a first image, the source 3D model is rendered to generate a second image, the explicit Gaussian parameters and the neural Gaussian parameters are updated based on the first image and the second image, and a second 3D scene is generated using the updated explicit Gaussian parameters and the neural Gaussian parameters.

2. The three-dimensional scene generation method according to claim 1, characterized in that, The step of determining the neural Gaussian parameters based on the three-dimensional coordinates in the explicit Gaussian parameters and the pre-set viewpoint direction features includes: The three-dimensional coordinates and pre-set viewpoint direction features in the explicit Gaussian parameters are input into a neural Gaussian network. The neural Gaussian parameters are generated through the neural Gaussian network, which includes a geometric prediction network and an appearance prediction network. A first feature is generated through the geometric prediction network and the three-dimensional coordinates. The geometric parameters of the neural three-dimensional Gaussian point are predicted based on the decoding result of the first feature. A second feature is generated through the appearance prediction network, the three-dimensional coordinates, and the viewpoint direction features. The appearance parameters of the neural three-dimensional Gaussian point are predicted based on the decoding result of the second feature.

3. The three-dimensional scene generation method according to claim 1, characterized in that, The step of updating the explicit Gaussian parameters and the neural Gaussian parameters based on the first image and the second image includes: Perform Fourier transform on the first image and the second image to the frequency domain; Calculate the first frequency loss and the second frequency loss based on the first image and the second image in the frequency domain, and calculate the image loss between the first image and the second image in the frequency domain; The explicit Gaussian parameters are updated based on the first frequency loss and the image loss; The neural Gaussian parameters are updated based on the second frequency loss and the image loss.

4. The three-dimensional scene generation method according to claim 3, characterized in that, The step of calculating the first frequency loss and the second frequency loss based on the first image and the second image in the frequency domain includes: Based on the number of updates of the explicit Gaussian parameters and the neural Gaussian parameters; The target cutoff frequency is determined based on the number of updates, the spectral coordinates of the first image and the second image in the frequency domain are identified, and the radius is calculated based on the spectral coordinates. The first frequency loss and the second frequency loss are determined based on the target cutoff frequency and the radius.

5. The three-dimensional scene generation method according to claim 3, characterized in that, Before updating the explicit Gaussian parameters and the neural Gaussian parameters based on the first image and the second image, the process includes: Calculate the rendering loss based on the first image and the second image; Calculate the total loss based on the first frequency loss, the second frequency loss, and the rendering loss; If the total loss is greater than a preset loss threshold, the explicit Gaussian parameters and the neural Gaussian parameters are updated based on the first image and the second image; otherwise, the updating of the explicit Gaussian parameters and the neural Gaussian parameters is stopped.

6. The three-dimensional scene generation method according to claim 1, characterized in that, The process of generating a second 3D scene using the updated explicit Gaussian parameters and the neural Gaussian parameters includes: Generate bi-state Gaussian parameters based on the explicit Gaussian parameters and the neural Gaussian parameters; Generate at least one rendered image from the target viewpoint based on the dual-state Gaussian parameters; The rendered images from at least one of the target viewpoints are fused to generate the second three-dimensional scene.

7. The three-dimensional scene generation method according to claim 6, characterized in that, After fusing rendered images from at least one of the target viewpoints to generate the second 3D scene, the method further includes: Obtain the third image of the second 3D scene rendered from the reference viewpoint; Obtain the fourth image of the second 3D scene rendered from the target viewpoint; Based on the third image and the fourth image, calculate the amount of color change of the second three-dimensional scene between the target viewpoint and the reference viewpoint; The second 3D scene is optimized based on the amount of color change.

8. The three-dimensional scene generation method according to claim 1, characterized in that, Before generating the second 3D scene using the updated explicit Gaussian parameters and the neural Gaussian parameters, the process includes: Gaussian sampling is performed on the first image and the second image; Obtain the rendering contribution value and the number of intersecting tiles of the candidate Gaussian points obtained by Gaussian sampling, and calculate the activity of the candidate Gaussian points based on the rendering contribution value and the number of intersecting tiles; The candidate Gaussian point is replaced with one of an explicit 3D Gaussian point and a neural 3D Gaussian point based on the activity of the candidate Gaussian point.

9. The three-dimensional scene generation method according to claim 8, characterized in that, The step of replacing the candidate Gaussian point with one of an explicit 3D Gaussian point and a neural 3D Gaussian point based on the activity of the candidate Gaussian point includes: Get the pre-set activity threshold; If the activity level of the candidate Gaussian point is greater than or equal to the activity level threshold, then the candidate Gaussian point is replaced with the explicit three-dimensional Gaussian point. If the activity level of the candidate Gaussian point is less than the activity threshold, then the candidate Gaussian point is replaced with the neural 3D Gaussian point.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the three-dimensional scene generation method as described in any one of claims 1 to 9 when executing the computer program.