Method for constructing three-dimensional scene model, apparatus, electronic device, program product and medium

WO2026200644A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/084186
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-18
Publication Date
2026-10-01

Smart Images

  • Figure CN2026084186_01102026_PF_FP_ABST
    Figure CN2026084186_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of computer vision. Disclosed in the embodiments of the present application are a method for constructing a three-dimensional scene model, an apparatus, an electronic device, a program product and a medium, which are used for controlling the optimization process of Gaussian primitives, so as to avoid redundant Gaussian primitives, thereby avoiding the waste of computing power resources. The method comprises: acquiring a first scene model corresponding to a target scene, the first scene model comprising a plurality of Gaussian primitives, the Gaussian primitives being three-dimensional Gaussian distributions of data points in a point cloud of the target scene, and the first scene model being used for an electronic device to generate a two-dimensional image of the target scene; determining a target attribute value of each Gaussian primitive in the first scene model, the target attribute value representing the degree of influence of the Gaussian primitive on the quality of the two-dimensional image; and, on the basis of the target attribute value of each Gaussian primitive in the first scene model, optimizing the first scene model to obtain a second scene model corresponding to the target scene, the optimization comprising densification and / or pruning.
Need to check novelty before this filing date? Find Prior Art

Description

A method, apparatus, electronic device, program product, and medium for constructing a three-dimensional scene model.

[0001] This application claims priority to Chinese Patent Application No. 202510391078.9, filed on March 28, 2025, entitled "A method, apparatus, electronic device, program product and medium for constructing a three-dimensional scene model", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer vision, and more particularly to a method, apparatus, electronic device, program product, and medium for constructing a three-dimensional scene model. Background Technology

[0003] 3D rendering refers to the process of converting a 3D physical scene or model into a 2D image, and it is widely used in various industries and fields. Among these, novel perspective synthesis technology refers to rendering an image of the scene from a new perspective, given multiple known scene images. This technology is widely used in downstream tasks of physical world modeling and simulation, such as Virtual Reality (VR), Augmented Reality (AR), and dynamic scene modeling. Traditional 3D rendering, which requires a large amount of physical information as input, has limited applications and is now being replaced by neural rendering based on generative artificial intelligence (AI).

[0004] Three-Dimensional Gaussian Splatting (3DGS) is one of the most widely used neural rendering methods. 3DGS is primarily used for visualizing 3D data. Its core idea is to represent data points in a 3D physical scene as a Gaussian distribution and "project" these distributions onto a 2D screen, thereby generating a smooth 2D image.

[0005] However, in the 3DGS process, there are redundant and uncontrollable Gaussian distributions, which not only lead to insufficient memory, but also waste computing resources during the rendering process. Summary of the Invention

[0006] This application provides a method, apparatus, electronic device, program product, and medium for constructing a three-dimensional scene model, which can effectively control the Gaussian element optimization process, avoid the occurrence of redundant Gaussian elements, and thus avoid the waste of computing resources.

[0007] Firstly, a method for constructing a three-dimensional scene model is provided. This method is applied to an electronic device. The method includes: obtaining a first scene model corresponding to a target scene; the first scene model includes multiple Gaussian elements, where each Gaussian element is a three-dimensional Gaussian distribution of data points in the point cloud of the target scene; the first scene model is used by the electronic device to generate a two-dimensional image of the target scene; determining the target attribute value of each Gaussian element in the first scene model, where the target attribute value represents the degree of influence of the Gaussian element on the quality of the two-dimensional image; and optimizing the first scene model according to the target attribute value of each Gaussian element in the first scene model to obtain a second scene model corresponding to the target scene, wherein the optimization includes densification and / or pruning.

[0008] As shown above, optimizing the first scene model by assigning target attribute values ​​to each Gaussian gramm avoids uncontrollable densification during the optimization process, which leads to Gaussian metadata redundancy and directly affects the utilization of GPU memory resources. The above method effectively controls the Gaussian gramm optimization process and avoids wasting computing resources.

[0009] In one possible implementation, the target attribute value includes a first attribute value and a second attribute value; the first attribute value is used to represent the strength of the influence of Gaussian elements on the quality of the two-dimensional image; and the second attribute value is used to represent the weakness of the influence of Gaussian elements on the quality of the two-dimensional image.

[0010] In one possible implementation, a first target Gaussian element is determined, and the first scene model is densified based on the first target Gaussian element; the first attribute value of the first target Gaussian element satisfies a first preset condition; densification means adding the first target Gaussian element; a second target Gaussian element is determined, and the first scene model is pruned based on the second target Gaussian element; the second attribute value of the second target Gaussian element satisfies a second preset condition; pruning means removing the second target Gaussian element.

[0011] As can be seen from the above, densifying the first target Gaussian unit whose first attribute value meets the first preset condition and pruning the second target Gaussian unit whose second attribute value meets the second preset condition can not only improve the accuracy of the 3D scene model construction, but also avoid uncontrollable densification.

[0012] In one possible implementation, for each Gaussian element in the first scene model, a first attribute value of the Gaussian element is determined based on its loss function, gradient, and / or projected area; the first attribute value is proportional to the loss function, gradient, and / or projected area; wherein, the loss function represents the loss of the first scene model relative to the target scene; the gradient represents the sum of the gradients of the attribute parameters of the Gaussian element; the attribute parameters include mean, covariance matrix, color, and / or opacity; and the projected area represents the two-dimensional projected area of ​​the Gaussian element.

[0013] As shown above, a larger Gaussian loss function and gradient indicate that the Gaussian has not yet converged and the corresponding model is not accurate enough; Gaussian elements with excessively large projected areas will cause blur artifacts. Therefore, the first attribute value is proportional to the loss function, gradient, and / or projected area. Determining the first attribute value using the above method can improve the accuracy of Gaussian densification.

[0014] In one possible implementation, for each Gaussian element in the first scene model, a second attribute value of the Gaussian element is determined based on the blending weight and / or opacity of the Gaussian element; the second attribute value is inversely proportional to the blending weight and / or opacity; wherein the blending weight represents the weight of a single Gaussian element among the multiple Gaussian elements when the multiple Gaussian elements are projected onto the target pixel.

[0015] As can be seen from the above, the lower the blending weight or the lower the opacity of a Gaussian primitive, the less influence that primitive has on the rendered pixel color, meaning its impact on the quality of the 2D image is weaker. Using the above method to determine the second attribute value can improve the accuracy of Gaussian primitive pruning.

[0016] In one possible implementation, if the scale of the first target Gaussian unit is smaller than a first preset threshold, the first target Gaussian unit is copied; if the scale of the first target Gaussian unit is larger than the first preset threshold, the target Gaussian unit is split.

[0017] As can be seen from the above, different densification methods for splitting or copying the first target Gaussian units of different sizes can improve the accuracy of the densified Gaussian units, thereby further improving the realism of the model rendering.

[0018] In one possible implementation, if the second attribute value includes a pruning score and the second preset condition includes exceeding a second preset threshold, then each Goski element with a pruning score greater than the second preset threshold is identified as a second target Goski element, and the second target Goski element is removed.

[0019] As can be seen from the above, removing the second target Gaussian primitives that have a lower impact on the rendering of pixel color can effectively reduce the number of invalid primitives, solve the problem of excessive memory occupation, and avoid wasting computing resources.

[0020] In one possible implementation, a second attribute value is determined for each Gaussian element; the second attribute value includes a pruning probability; the pruning probability represents the probability that a Gaussian element will be removed; based on the pruning quantity and the pruning probability, a second target Gaussian element is determined, and the second target Gaussian element is removed; the number of second target Gaussian elements is the pruning quantity.

[0021] As shown above, since the weight ratios and different opacities of Gaussian primitives at local geometric structures in a scene model are often quite similar, normalizing the pruning scores corresponding to each Gaussian primitive to a pruning probability, and then randomly selecting the second target Gaussian primitive based on the pruning quantity, can avoid large-scale pruning of the same part of the scene model, which would cause geometric structure loss. This further improves the accuracy of model rendering.

[0022] In one possible implementation, an original image set of the target scene is obtained; the original image set includes images of the target scene from at least one viewpoint; based on the original image set, a point cloud of the target scene is determined; based on the point cloud of the target scene, an initial scene model of the target scene is generated; wherein, the first scene model is the initial scene model or the initial scene model is obtained after optimization of the initial scene model for a preset number of times.

[0023] In one possible implementation, an image of the target scene from the target perspective is generated based on the target viewpoint and the second scene model; the target viewpoint is different from the viewpoint corresponding to the original image.

[0024] As can be seen from the above, electronic devices acquire images of the target scene from the target perspective based on the target viewpoint and the second scene model. This expands the applicability of the 3D scene model construction method, enabling it to be applied to new perspective synthesis techniques and improving the accuracy of the 2D images acquired in these techniques.

[0025] Secondly, a device for constructing a three-dimensional scene model is provided. In embodiments of this application, the device can be divided into functional modules according to the method provided in the first aspect. For example, different functional modules can be defined for each function, or two or more functions can be inherited into a single processing module. For instance, embodiments of this application can divide the three-dimensional scene model construction device into an acquisition module, a calculation module, and an optimization module according to their functions. The descriptions of the possible technical solutions and beneficial effects of each of the above-described functional modules can refer to the technical solutions provided in the first aspect or its corresponding possible implementations, and will not be repeated here.

[0026] Thirdly, embodiments of this application provide an electronic device, which includes a processor and a memory for storing processor-executable instructions; the processor is configured to execute instructions, causing the electronic device to perform the above-described method for constructing a three-dimensional scene model.

[0027] Fourthly, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the method for constructing a three-dimensional scene model provided in the various optional implementations of the first aspect described above.

[0028] Fifthly, embodiments of this application provide a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor of an electronic device to implement the method for constructing a three-dimensional scene model as described in the first aspect above.

[0029] For a detailed description of the second to fifth aspects and their various implementations in the embodiments of this application, please refer to the detailed description in the first aspect and its various implementations; and for a detailed description of the beneficial effects of the second to fifth aspects and their various implementations, please refer to the beneficial effect analysis in the various implementations of the first aspect, which will not be repeated here.

[0030] These or other aspects of the embodiments of this application will become more apparent in the following description. Attached Figure Description

[0031] Figure 1 shows a hardware schematic diagram of an electronic device 100 provided in an embodiment of this application;

[0032] Figure 2 shows a flowchart illustrating a method for constructing a three-dimensional scene model according to an embodiment of this application;

[0033] Figure 3 shows a schematic diagram of obtaining the original image set of a target scene 200 according to an embodiment of this application;

[0034] Figure 4 shows a schematic diagram of a Gaussian element projected onto a two-dimensional plane according to an embodiment of this application;

[0035] Figure 5 shows a schematic diagram of the software architecture of a three-dimensional scene model construction system 400 provided in an embodiment of this application;

[0036] Figure 6 shows a schematic diagram of the structure of a three-dimensional scene model construction device 800 provided in an embodiment of this application. Detailed Implementation

[0037] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish the same or similar items with basically the same function and effect.

[0038] Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or order of execution, and that "first," "second," etc., are not necessarily different. Furthermore, in some embodiments of this application, words such as "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.

[0039] Furthermore, the device architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of device architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0040] First, the application scenarios of the embodiments of this application will be introduced by way of example.

[0041] 3DGS is a technique for 3D rendering, primarily used to visualize 3D data. 3DGS can represent a 3D physical scene as multiple Gaussian distributions, each with its own position, size, color, and other attributes. The shape of the Gaussian distribution can be spherical, ellipsoidal, or other complex shapes. Typically, 3DGS uses 3D Gaussian ellipsoids as primitives to model the 3D scene as a series of ellipsoid sets, and finally performs rasterization rendering through "projection" to obtain a 2D image from a new perspective.

[0042] In the 3DGS optimization process, densification is often used to increase the number of primitives, thereby improving rendering accuracy. However, the number of high-order primitives directly affects the use of video memory resources. Therefore, redundant and uncontrollable high-order primitives in the 3DGS process will lead to a waste of computing resources.

[0043] In related technologies, compressing Gaussian parameters such as color, position, and covariance by clustering or codebooking Gaussian units can save some computational resources. However, this method degrades the quality of the rendered 2D image. Alternatively, Gaussian units can be directly output using a pre-trained neural network model, avoiding the densification process. While this method avoids redundancy in the number of Gaussian units, the quality of the rendered 2D image is affected by the model, resulting in unstable rendering effects.

[0044] In view of this, embodiments of this application provide a method for constructing a three-dimensional scene model. This method is applied to an electronic device, which acquires a first scene model corresponding to a target scene. The first scene model includes multiple Gaussian elements, which represent the three-dimensional Gaussian distribution of data points in the point cloud of the target scene. The first scene model is used by the electronic device to generate a two-dimensional image of the target scene. The electronic device determines the target attribute values ​​of each Gaussian element in the first scene model. The target attribute values ​​represent the degree of influence of the Gaussian elements on the quality of the two-dimensional image. Based on the target attribute values ​​of each Gaussian element in the first scene model, the electronic device optimizes the first scene model to obtain a second scene model corresponding to the target scene. The optimization includes densification and / or pruning. As can be seen from the above method, optimizing the first scene model using the target attribute values ​​of each Gaussian element not only improves the construction accuracy of the three-dimensional scene model but also avoids the waste of computing resources caused by uncontrollable densification during the optimization process.

[0045] Secondly, the system architecture of the embodiments of this application will be described by way of example.

[0046] The electronic devices provided in this application embodiment may include mobile phones, smartwatches, tablet computers, foldable electronic devices, desktop computers, laptop computers, handheld computers, laptops, Ultra-Mobile Personal Computers (UMPCs), netbooks, cellular phones, Personal Digital Assistants (PDAs), Augmented Reality (AR) devices, Virtual Reality (VR) devices, Artificial Intelligence (AI) devices, wearable devices, in-vehicle devices, and other industrial terminals with communication capabilities, as well as servers. Specifically, the server may be a blade server, a high-density server, a rack server, or a high-performance server. This application embodiment does not impose any special limitations on the specific type of electronic device.

[0047] Furthermore, the operating system installed on electronic devices includes, but is not limited to, Alternatively, other operating systems may be used. This application does not limit the specific type of electronic device or the type of operating system installed. Those skilled in the art can determine the type of electronic device and the installed operating system according to actual needs, and these designs do not exceed the protection scope of the embodiments of this application.

[0048] For example, Figure 1 shows a hardware schematic diagram of an electronic device 100 provided in an embodiment of this application. The electronic device 100 is used to implement a method for constructing a three-dimensional scene model. As shown in Figure 1, the electronic device 100 includes a processor 110, a memory 120, a communication interface 130, and a bus 140. The processor 110, the memory 120, and the communication interface 130 are connected via the bus 140. It should be understood that this application does not limit the number of processors 110 and memory 120 in the electronic device 200.

[0049] Specifically, processor 110 can be a central processing unit (CPU) or a graphics processing unit (GPU). Processor 110 is used to acquire a first scene model corresponding to the target scene. The first scene model includes multiple Gaussian elements, which represent the three-dimensional Gaussian distribution of data points in the point cloud of the target scene. The first scene model is used by the electronic device to generate a two-dimensional image of the target scene.

[0050] The processor 110 is also used to determine the target attribute values ​​of each Gaussian unit in the first scene model. The target attribute values ​​represent the degree of influence of the Gaussian unit on the quality of the two-dimensional image. Based on the target attribute values ​​of each Gaussian unit in the first scene model, the processor 110 optimizes the first scene model to obtain a second scene model corresponding to the target scene. The optimization includes densification and / or pruning.

[0051] Memory 120 is used to store processor-executable instructions. Memory 120 can be a data-preserving storage space that does not lose data when power is off, specifically non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), etc. Communication interface 130 is used to realize data transmission and communication between the electronic device and other devices.

[0052] It should be noted that the system architecture and application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0053] For ease of understanding, the following description, in conjunction with the accompanying drawings, illustrates the method for constructing a three-dimensional scene model provided in this application, which is performed by the electronic device 100 shown in Figure 1.

[0054] Figure 2 shows a flowchart illustrating a method for constructing a three-dimensional scene model according to an embodiment of this application. This method for constructing the three-dimensional scene model is executed by the electronic device 100 shown in Figure 1 and includes the following steps:

[0055] S101, the electronic device acquires the first scene model corresponding to the target scene.

[0056] The first scene model includes multiple Gaussian elements, which are the three-dimensional Gaussian distributions of data points in the point cloud of the target scene; the first scene model is used by electronic devices to generate two-dimensional images of the target scene.

[0057] In one possible implementation, the electronic device acquires a raw image set of the target scene; based on the raw image set, the electronic device determines the point cloud of the target scene. Based on the point cloud of the target scene, the electronic device generates an initial scene model of the target scene.

[0058] The original image set includes images of the target scene from at least one viewpoint; the first scene model is the initial scene model or the initial scene model after a preset number of optimizations.

[0059] In some examples, the electronic device acquires a raw set of images transmitted by a camera, which captures images of the target scene from at least one viewpoint. Based on the images in the raw set, the electronic device determines a point cloud of the target scene.

[0060] For example, Figure 3 shows a schematic diagram of acquiring a raw image set of a target scene 200 according to an embodiment of this application. As shown in Figure 3, the target scene 200 includes a first camera 210 and a second camera 220. The first camera 210 and the second camera 220 are located at different positions. The first camera 210 and the second camera 220 are used to capture images. Because the first camera 210 and the second camera 220 are located at different positions, the images of the target scene 200 captured by the first camera 210 and the second camera 220 are images of the target scene 200 from different perspectives. The electronic device acquires the images transmitted by the first camera 210 and the second camera 220.

[0061] Electronic devices extract feature points from 2D images using the Scale Invariant Feature Transform (SIFT) algorithm. These feature points are located at positions with unique properties within the 2D image. Feature points from other viewpoints of 2D images are matched using the K-Nearest Neighbors (KNN) algorithm to find the correspondence between feature points in 2D images from different viewpoints. Specifically, the electronic device calculates the distance (e.g., Euclidean distance) between feature points in one 2D image and feature points in other 2D images from different viewpoints, selecting the K nearest points as candidate matching points. Matching feature points in 2D images from different viewpoints determines the corresponding positions of the same feature point in 2D images from different viewpoints. The SIFT algorithm is used to extract representative and unique key data points from the image and generate a feature vector for each key data point. The core idea of ​​the KNN algorithm is that if a sample's K nearest neighbors in the feature space mostly belong to a certain class, then the sample also belongs to that class and possesses the characteristics of samples in that class. The KNN algorithm can be used to match feature points in images.

[0062] After obtaining the pixel coordinates of corresponding feature points in multiple two-dimensional images through feature matching, electronic devices perform three-dimensional reconstruction using the principle of triangulation. Specifically, by combining the coordinates of corresponding data points in two-dimensional images from different perspectives with the parameters of at least one imaging device, the coordinates of the matched feature points in the three-dimensional coordinate system are calculated, thereby obtaining a three-dimensional point cloud.

[0063] In other examples, the electronic device constructs a Gaussian distribution centered on each data point in the point cloud to determine Gaussian primitives. Based on these Gaussian primitives, the electronic device renders an initial model of the target scene.

[0064] For example, an electronic device uses data points in a 3D point cloud as centers and employs scientific computing and point cloud processing libraries in Python to establish a Gaussian distribution. Preset parameters are assigned to each attribute of the Gaussian distribution to initialize it, resulting in multiple Gaussian primitives. These attributes include center position, shape, color, and / or opacity. The center position of a Gaussian primitive is determined by the mean of the Gaussian normal distribution function; the shape is determined by the covariance of the Gaussian normal distribution function; and the color is determined by the fourth-order spherical harmonic function. The shapes of Gaussian primitives can include complex shapes such as spheres or ellipsoids. Based on the preset parameters of each Gaussian primitive, the electronic device draws the surface of the Gaussian distribution function corresponding to each primitive and connects them to form a 3D mesh or surface, thereby constructing an initial scene model corresponding to the target scene.

[0065] In one possible implementation, the electronic device optimizes the initial scene model to obtain a first scene model.

[0066] For example, the electronic device acquires a 2D image of the initial scene model from a first-view perspective, calculates the loss function between the 2D image of the initial scene from the first-view perspective and the 2D image of the target scene from the first-view perspective, and updates the attribute parameters of each Gaussian unit through backpropagation. After a preset number of optimizations, the first scene model is obtained.

[0067] The loss function, also known as the mean absolute error loss function, is used to measure the average absolute error between the pre-test and the true value.

[0068] For example, the electronic device projects each Gaussian unit from the initial scene model onto a first plane and obtains a 2D image of the initial scene model through differentiable rasterization rendering. The electronic device compares the 2D image of the initial scene model with a 2D image of the target scene from the same viewpoint to obtain a loss function. If the loss function is large, it indicates that the initial scene model differs significantly from the real target scene model. In this case, backpropagation is performed to update the attribute parameters of each Gaussian unit. After a preset number of backpropagations, the first scene model is obtained.

[0069] S102, the electronic device determines the target attribute values ​​of each Gaussian element in the first scene model.

[0070] The target attribute value represents the degree of influence of Gaussian elements on the quality of the two-dimensional image.

[0071] In this embodiment of the application, the target attribute value includes a first attribute value and a second attribute value; the first attribute value is used to represent the strength of the influence of Gaussian elements on the quality of the two-dimensional image; the second attribute value is used to represent the weakness of the influence of Gaussian elements on the quality of the two-dimensional image.

[0072] In one possible implementation, the electronic device determines a first target Gaussian element. The electronic device then densifies a first scene model based on the first target Gaussian element. The first attribute value of the first target Gaussian element satisfies a first preset condition. Densification represents the addition of the first target Gaussian element.

[0073] The electronic device determines a second target Gaussian element. Based on this second target Gaussian element, the electronic device prunes the first scene model. Specifically, the second attribute value of the second target Gaussian element satisfies a second preset condition; pruning means removing the second target Gaussian element.

[0074] In some examples, the electronic device determines the first attribute value of each Gaussian element in the first scene model based on the Gaussian element's loss function, gradient, and / or projected area.

[0075] The first attribute value is proportional to the loss function, gradient, and / or projected area. The loss function represents the loss of the first scene model relative to the target scene; the gradient represents the sum of the gradients of the attribute parameters of the Gaussian unit; the attribute parameters include mean, covariance matrix, color, and / or opacity; and the projected area represents the two-dimensional projected area of ​​the Gaussian unit.

[0076] For example, the first attribute value includes a density score, and the first preset condition includes exceeding a third preset threshold. The density score formula is as follows: D=L(X1*G+X2*S)

[0077] Where D represents the density score; L represents the loss function of the Gaussian primitive; G represents the gradient of the Gaussian primitive; S represents the projected area of ​​the Gaussian primitive; and X1 and X2 represent custom coefficients. The electronic device determines the density score of each Gaussian primitive based on its loss function, gradient, projected area, and the density score formula described above. A higher density score indicates a stronger influence of the Gaussian primitive on the quality of the generated 2D image. The electronic device identifies Gaussian primitives with density scores exceeding a third preset threshold as first target Gaussian primitives. Furthermore, it densifies these first target Gaussian primitives, increasing their number.

[0078] In other examples, the electronic device determines the second attribute value of each Gaussian in the first scene model based on the Gaussian's blending weights and / or opacity.

[0079] The second attribute value is inversely proportional to the blending weight and / or opacity. The blending weight represents the weight of an individual Gaussian cell among multiple Gaussian cells when projected onto a target pixel.

[0080] For example, if the second attribute value includes a pruning score, and the second preset condition includes exceeding a second preset threshold, the electronic device will identify each Goxsky unit with a pruning score greater than the second preset threshold as the second target Goxsky unit and discard the second target Goxsky units. The pruning score formula is as follows: P = 1 / (Y1*W + Y2*O)

[0081] Where P represents the pruning score; W represents the blending weight; O represents the opacity; and Y1 and Y2 represent custom coefficients. The electronic device determines the pruning score of each Gaussian unit based on its blending weight, opacity, and the aforementioned pruning score formula. A higher pruning score indicates a weaker impact of the Gaussian unit on the quality of the generated 2D image. The electronic device identifies Gaussian units with pruning scores exceeding a second preset threshold as second target Gaussian units. Furthermore, the electronic device prunes and removes these second target Gaussian units.

[0082] Optionally, the electronic device determines a second attribute value for each Gaussian unit. The second attribute value includes the pruning probability. Based on the pruning quantity and pruning probability, the electronic device determines a second target Gaussian unit and removes it.

[0083] Here, the pruning probability represents the probability of a Gaussian element being removed; the number of second-objective Gaussian elements is the number of pruning elements.

[0084] For example, the electronic device obtains the pruning score of each Gaussian unit according to the pruning scoring formula. The electronic device normalizes the pruning score of each Gaussian unit to obtain the pruning probability of each Gaussian unit. The electronic device determines the second target Gaussian unit based on the pruning quantity and the pruning probability of each Gaussian unit.

[0085] For example, the pruning probability formula for any Gaussian element is as follows: Qi = Pi / (P1 + P2 + ... + Pj) * 100%

[0086] Where i is a positive integer less than j. The electronic device can obtain scores of 1, 2, 3, and 4 for each Gaussian unit based on the pruning scoring formula. The pruning probability formula for each Gaussian unit determines the pruning probabilities as follows: 1 / (1+2+3+4)*100% = 10%, 2 / (1+2+3+4)*100% = 20%, 3 / (1+2+3+4)*100% = 30%, and 4 / (1+2+3+4)*100% = 40%. If the pruning quantity is 2, the electronic device randomly selects 2 second target Gaussian units for pruning based on the pruning quantity and the pruning probability of each Gaussian unit.

[0087] As shown above, since the weight ratios and different opacities of Gaussian primitives at local geometric structures in a scene model are often quite similar, normalizing the pruning scores corresponding to each Gaussian primitive to a pruning probability, and then randomly selecting the second target Gaussian primitive based on the pruning quantity, can avoid large-scale pruning of the same part of the scene model, which would cause geometric structure loss. This further improves the accuracy of model rendering.

[0088] In one possible implementation, during the rendering of each Gaussian pixel, the electronic device projects each Gaussian pixel onto different pixels. For a target pixel, at least one Gaussian pixel is projected onto it. If the sum of the density scores of at least one Gaussian pixel on the target pixel satisfies a preset condition, the electronic device identifies at least one Gaussian pixel on the target pixel as a first target Gaussian pixel. If the sum of the pruning scores of at least one Gaussian pixel on the target pixel satisfies a preset condition, the electronic device identifies at least one Gaussian pixel on the target pixel as a second target Gaussian pixel.

[0089] Figure 4 illustrates a schematic diagram of Gaussian primitives projected onto a two-dimensional plane according to an embodiment of this application. As shown in Figure 4, the first Gaussian primitive 310 and the second Gaussian primitive 320 are projected onto the target pixel region 300 during the rendering process, and are projected as the first Gaussian primitive 310' and the second Gaussian primitive 320' in the target pixel region 300. The electronic device determines the density score and / or pruning score of the first Gaussian primitive 310 and the second Gaussian primitive 320 respectively according to the above-mentioned Gaussian primitive density score formula and / or pruning score formula. The electronic device determines the density score and / or pruning score of the target pixel based on the sum of the density score and / or pruning score of the first Gaussian primitive 310 and the second Gaussian primitive 320. Thus, the first Gaussian primitive 310 and the second Gaussian primitive 320 corresponding to the target pixel are determined as the first target Gaussian primitive or the second target Gaussian primitive.

[0090] As can be seen from the above, by determining the target attribute value of Gaussian elements based on their various attribute parameters and quantifying the influence of Gaussian elements on the quality of two-dimensional images, we can more accurately identify the Gaussian elements that need to be densified or pruned. This improves the accuracy of optimizing the target scene model, avoids redundant Gaussian elements, and further avoids wasting computing resources during scene model rendering.

[0091] S103, the electronic device optimizes the first scene model based on the target attribute values ​​of each Gaussian element in the first scene model to obtain the second scene model corresponding to the target scene.

[0092] Optimization includes densification and / or pruning.

[0093] For example, the electronic device determines a first target Gaussian element and a second target Gaussian element based on the target attribute values ​​of each Gaussian element. The electronic device densifies the first target Gaussian element and prunes the second Gaussian element. Densification means adding the first target Gaussian element; pruning means removing the second target Gaussian element. The first attribute value of the first target Gaussian element satisfies a first preset condition; the second attribute value of the second target Gaussian element satisfies a second preset condition.

[0094] In one possible implementation, if the scale of the first target Gaussian unit is smaller than a first preset threshold, the first target Gaussian unit is copied; if the scale of the first target Gaussian unit is larger than the first preset threshold, the target Gaussian unit is split.

[0095] Among the attribute parameters of a Gausky element, the covariance matrix describes the covariance relationship between multiple variables, and variance reflects the degree of dispersion of the variables themselves. A larger variance in a Gausky element indicates a greater degree of data dispersion along the corresponding coordinate axis, meaning the Gausky element has a larger scale along that axis. A larger-scale Gausky element will be more "flat" or "stretched" in the corresponding dimension. Therefore, a larger-scale Gausky element can be considered as having a larger volume.

[0096] For example, if the scale of the first target Gaussian unit is smaller than a first preset threshold, it indicates that the volume of the first target Gaussian unit is small. The electronic device determines the center position of the newly copied Gaussian unit based on the mean and gradient direction of the first target Gaussian unit, and copies the first target Gaussian unit to generate a Gaussian unit identical to the first target Gaussian unit. If the scale of the first target Gaussian unit is larger than the first preset threshold, it indicates that the volume of the first target Gaussian unit is large. The electronic device uses a splitting method to replace the first target Gaussian unit with multiple smaller Gaussian units.

[0097] As can be seen from the above, different densification methods for splitting or copying the first target Gaussian units of different sizes can improve the accuracy of the densified Gaussian units, thereby further improving the realism of the model rendering.

[0098] Optionally, the electronic device obtains a second scene model corresponding to the target scene based on the optimized first target Gaussian unit and / or second target Gaussian unit.

[0099] It should be noted that during the backpropagation optimization of the first scenario model, the target attribute values ​​of each Gaussian unit can be determined in each optimization step, and densification and / or pruning can be performed. Alternatively, the target attribute values ​​of each Gaussian unit can be determined and densification and / or pruning can be performed after the number of optimization iterations meets a preset limit. The specific optimization method used depends on the specific application scenario and does not constitute a specific limitation.

[0100] In one possible implementation, the electronic device generates an image of the target scene from the target viewpoint, based on the target viewpoint and a second scene model.

[0101] The target viewpoint differs from the viewpoint corresponding to the original image. The second scene model is obtained by optimizing the first scene model.

[0102] For example, as shown in Figure 3 above, the first camera 210 and the second camera 220 are used to capture raw images and send them to an electronic device. The viewpoints corresponding to the first camera 210 and the second camera 220 are different from the target viewpoint. Based on the target viewpoint, the electronic device projects each Gaussian element in the second scene model onto the two-dimensional plane where the target viewpoint is located, and performs operations such as segmentation, depth sorting, and blending rendering through differentiable raster rendering to obtain a two-dimensional image of the target scene under the target viewpoint.

[0103] As can be seen from the above, the aforementioned electronic device acquires an image of the target scene from the target perspective based on the target viewpoint and the second scene model, thereby improving the applicability of the three-dimensional scene model construction method and enabling the three-dimensional scene model construction method to be applied to the new perspective synthesis technology, thus improving the accuracy of the two-dimensional images acquired in the new perspective synthesis technology.

[0104] The above describes the construction method of a 3D scene model from the perspective of method implementation. The following describes the software architecture for implementing the construction method of a 3D scene model from the perspective of software architecture.

[0105] Figure 5 shows a schematic diagram of the software architecture of a three-dimensional scene model construction system 400 provided in an embodiment of this application. As shown in Figure 5, the three-dimensional scene model construction system 400 includes a motion recovery module 410, an initialization module 420, a primitive projection module 430, a differentiable rasterization rendering module 440, and a sparse densification module 450.

[0106] The motion recovery module 410 acquires the original image set of the target scene and determines the point cloud of the target scene based on the original image set. The initialization module 420 constructs a Gaussian distribution centered on each data point in the point cloud and assigns preset parameters to each attribute of the Gaussian distribution to obtain multiple Gaussian primitives. The primitive projection module 430 projects each Gaussian primitive in the scene model onto a two-dimensional plane to determine the two-dimensional Gaussian. The differentiable rasterization rendering module 440 performs operations such as block segmentation, depth sorting, and blending on the two-dimensional Gaussian projected onto the two-dimensional plane to render the corresponding two-dimensional image. The sparse densification module 450 determines the target attribute values ​​of each Gaussian primitive in the first scene model, determines the first target Gaussian primitive and the second target Gaussian primitive based on the target attribute values ​​of each Gaussian primitive, densifies the first target Gaussian primitive, and prunes the second Gaussian primitive.

[0107] Optionally, the software architecture of the above-mentioned three-dimensional scene model construction system 400 can be applied to the electronic device 100 shown in Figure 1.

[0108] In summary, this application provides a method for constructing a 3D scene model, used to control the Gaussian primitive optimization process, avoiding the generation of redundant Gaussian primitives, thereby preventing the waste of computing resources. This method is applied to an electronic device, which acquires a first scene model corresponding to a target scene. The first scene model includes multiple Gaussian primitives, which represent the 3D Gaussian distribution of data points in the point cloud of the target scene. The first scene model is used by the electronic device to generate a 2D image of the target scene. The electronic device determines the target attribute values ​​of each Gaussian primitive in the first scene model. The target attribute values ​​represent the degree of influence of the Gaussian primitives on the quality of the 2D image. Based on the target attribute values ​​of each Gaussian primitive in the first scene model, the electronic device optimizes the first scene model to obtain a second scene model corresponding to the target scene. The optimization includes densification and / or pruning. As can be seen from the above method, optimizing the first scene model using the target attribute values ​​of each Gaussian primitive in the first scene model can effectively avoid uncontrollable densification during the optimization process, avoid the generation of redundant Gaussian primitives, and further avoid the waste of computing resources during the rendering process.

[0109] The foregoing mainly describes the solutions of the embodiments of this application from a methodological perspective. It is understood that the 3D scene model construction apparatus, in order to realize the functions in the above-described 3D scene model construction method, includes at least one of the hardware structures and software modules corresponding to each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0110] This application embodiment can divide the 3D scene model construction device into functional units according to the above method example. For example, each function can be divided into separate functional units, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0111] For example, Figure 6 shows a schematic diagram of the structure of a three-dimensional scene model building device 800 provided in an embodiment of this application. As shown in Figure 6, the three-dimensional scene model building device can be applied to the electronic device 100 shown in Figure 1. The three-dimensional scene model building device 800 includes:

[0112] The acquisition module 810 is used to acquire a first scene model corresponding to the target scene; the first scene model includes multiple Gaussian elements, which are the three-dimensional Gaussian distributions of each data point in the point cloud of the target scene; the first scene model is used by the electronic device to generate a two-dimensional image of the target scene.

[0113] The calculation module 820 is used to determine the target attribute value of each Gaussian element in the first scene model. The target attribute value represents the degree of influence of the Gaussian element on the quality of the two-dimensional image.

[0114] The optimization module 830 is used to optimize the first scene model based on the target attribute values ​​of each Gaussian element in the first scene model to obtain the second scene model corresponding to the target scene. The optimization includes densification and / or pruning.

[0115] In one possible implementation, the acquisition module 810 is further configured to acquire an original image set of the target scene; the original image set includes images of the target scene from at least one viewpoint;

[0116] Based on the original image set, determine the point cloud of the target scene;

[0117] Based on the point cloud of the target scene, an initial scene model of the target scene is generated; wherein, the first scene model is the initial scene model or the initial scene model after being optimized a preset number of times.

[0118] In one possible implementation, the computation module 820 is further configured to determine a first target Gaussian element, and to densify the first scene model based on the first target Gaussian element; the first attribute value of the first target Gaussian element satisfies a first preset condition; densification means adding the first target Gaussian element; determine a second target Gaussian element, and to prune the first scene model based on the second target Gaussian element; the second attribute value of the second target Gaussian element satisfies a second preset condition; pruning means removing the second target Gaussian element.

[0119] In one possible implementation, the computation module 820 is further configured to determine a first attribute value for each Gaussian element in the first scene model based on the loss function, gradient, and / or projected area of ​​the Gaussian element; the first attribute value is proportional to the loss function, gradient, and / or projected area.

[0120] Here, the loss function represents the loss of the first scene model relative to the target scene; the gradient represents the sum of the gradients of the attribute parameters of the Gaussian unit; the attribute parameters include mean, covariance matrix, color and / or opacity; and the projected area represents the two-dimensional projected area of ​​the Gaussian unit.

[0121] In one possible implementation, the computation module 820 is further configured to determine a second attribute value for each Gaussian element in the first scene model, based on the blending weights and / or opacity of the Gaussian elements; the second attribute value is inversely proportional to the blending weights and / or opacity.

[0122] Here, the hybrid weight represents the weight of a single Gaussian cell among multiple Gaussian cells when multiple Gaussian cells are projected onto a target pixel.

[0123] In one possible implementation, the calculation module 820 is further configured to, when the second attribute value includes a pruning score and the second preset condition includes exceeding a second preset threshold, identify each Goski element with a pruning score greater than the second preset threshold as a second target Goski element and remove the second target Goski element.

[0124] In one possible implementation, the computation module 820 is further configured to determine a second attribute value for each Gaussian element; the second attribute value includes a pruning probability; the pruning probability represents the probability that a Gaussian element will be removed.

[0125] Based on the number of pruning elements and the pruning probability, determine the second objective Gaussian elements and remove them; the number of second objective Gaussian elements is the number of pruning elements.

[0126] In one possible implementation, the optimization module 830 is further configured to: copy the first target Gaussian Gaussian if the scale of the first target Gaussian is less than a first preset threshold; and split the target Gaussian if the scale of the first target Gaussian is greater than the first preset threshold.

[0127] In one possible implementation, the 3D scene model building device 800 further includes a rendering module for generating an image of the target scene from the target viewpoint, based on the target viewpoint and the second scene model; the target viewpoint is different from the viewpoint corresponding to the original image.

[0128] The acquisition module 810, the calculation module 820, and the optimization module 830 can all be implemented in software or in hardware.

[0129] For example, the implementation of the acquisition module 810 will be described below. Similarly, the implementation of the calculation module 820 and the optimization module 830 can refer to the implementation of the acquisition module 810.

[0130] As an example of a software functional unit, the acquisition module 810 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the acquisition module 810 may include code running on multiple hosts / virtual machines / containers. It should be understood that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0131] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, one VPC is set up within one region, and two VPCs within the same region...

[0132] Cross-region communication between VPCs, and between VPCs in different regions, requires setting up a communication gateway within each VPC to achieve interconnection between VPCs.

[0133] As an example of a hardware functional unit, the acquisition module 810 may include at least one computing device, such as a server. Alternatively, the acquisition module 810 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0134] The multiple computing devices included in the acquisition module 810 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 810 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 810 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0135] It should be understood that in other embodiments, the acquisition module 810 can be used to execute any step in the method for constructing a 3D scene model, and the calculation module 820 and the optimization module 830 can be used to execute any step in the method for constructing a 3D scene model. The steps implemented by the acquisition module 810, the calculation module 820 and the optimization module 830 can be specified as needed. By implementing different steps in the method for constructing a 3D scene model through the acquisition module 810, the calculation module 820 and the optimization module 830 respectively, all functions of the 3D scene model construction device 800 can be realized.

[0136] This application also provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform operations corresponding to any implementation scheme and various feasible implementation schemes of the method for constructing a three-dimensional scene model.

[0137] This application also provides a computer program product containing instructions that, when run on an electronic device, causes the electronic device to execute any implementation scheme and various feasible implementation schemes corresponding to the method for constructing a three-dimensional scene model.

[0138] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0139] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0140] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions, which, when loaded and executed on a computer, generate all or part of the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one network site, computer, server, or data center to another network site, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or it can be a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape, etc.), an optical medium (e.g., DVD, etc.), or a semiconductor medium (e.g., solid-state drive), etc.

[0141] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0142] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for constructing a three-dimensional scene model, characterized in that, Applied to electronic devices, the method includes: A first scene model corresponding to the target scene is obtained; the first scene model includes multiple Gaussian elements, and the Gaussian elements are the three-dimensional Gaussian distribution of each data point in the point cloud of the target scene; the first scene model is used by an electronic device to generate a two-dimensional image of the target scene; Determine the target attribute value of each Gaussian element in the first scene model, wherein the target attribute value represents the degree of influence of the Gaussian element on the quality of the two-dimensional image; Based on the target attribute values ​​of each Gaussian element in the first scene model, the first scene model is optimized to obtain a second scene model corresponding to the target scene. The optimization includes densification and / or pruning.

2. The construction method according to claim 1, characterized in that, The target attribute value includes a first attribute value and a second attribute value; the first attribute value is used to represent the strength of the influence of the Gaussian elements on the quality of the two-dimensional image; the second attribute value is used to represent the weakness of the influence of the Gaussian elements on the quality of the two-dimensional image.

3. The construction method according to claim 2, characterized in that, The step of optimizing the first scene model based on the target attribute values ​​of each Gaussian element in the first scene model to obtain a second scene model corresponding to the target scene includes: A first target Gaussian element is determined, and the first scene model is densified based on the first target Gaussian element; the first attribute value of the first target Gaussian element satisfies a first preset condition; the densification means increasing the first target Gaussian element. A second target Gaussian element is determined, and the first scene model is pruned based on the second target Gaussian element; the second attribute value of the second target Gaussian element satisfies a second preset condition; the pruning means removing the second target Gaussian element.

4. The construction method according to claim 2, characterized in that, Determining the target attribute values ​​of each Gaussian element in the first scene model includes: For each Gaussian element in the first scene model, the first attribute value of the Gaussian element is determined based on its loss function, gradient, and / or projected area; the first attribute value is proportional to the loss function, the gradient, and / or the projected area. Wherein, the loss function is used to represent the loss of the first scene model relative to the target scene; the gradient represents the sum of the gradients of the attribute parameters of the Gaussian element; the attribute parameters include mean, covariance matrix, color and / or opacity; and the projected area represents the two-dimensional projected area of ​​the Gaussian element.

5. The construction method according to claim 2, characterized in that, Determining the target attribute values ​​of each Gaussian element in the first scene model includes: For each Gaussian element in the first scene model, a second attribute value is determined based on the blending weight and / or opacity of the Gaussian element; the second attribute value is inversely proportional to the blending weight and / or the opacity. The hybrid weight represents the weight of a single Gaussian cell among the multiple Gaussian cells when the multiple Gaussian cells are projected onto the target pixel.

6. The construction method according to claim 3, characterized in that, The step of determining the first target Gaussian element and performing the densification on the first scene model based on the first target Gaussian element includes: If the scale of the first target Gaussian element is smaller than the first preset threshold, copy the first target Gaussian element; If the scale of the first target Gaussian unit is greater than the first preset threshold, the target Gaussian unit is split.

7. The construction method according to claim 3, characterized in that, The step of determining the second target Gaussian elements and pruning the first scene model based on the second target Gaussian elements includes: If the second attribute value includes a pruning score and the second preset condition includes exceeding a second preset threshold, then each of the Gaussian elements whose pruning score is greater than the second preset threshold is identified as the second target Gaussian element, and the second target Gaussian element is removed.

8. The construction method according to claim 1 or 2, characterized in that, The first scene model is optimized based on the target attribute values ​​of each Gaussian element in the first scene model to obtain a second scene model corresponding to the target scene, including... Determine a second attribute value for each Gaussian element; the second attribute value includes a pruning probability; the pruning probability represents the probability that the Gaussian element is removed. Based on the number of pruning operations and the pruning probability, a second target Gaussian element is determined, and then the second target Gaussian element is removed; the number of the second target Gaussian elements is the number of pruning operations.

9. The construction method according to any one of claims 1-8, characterized in that, Before obtaining the first scene model corresponding to the target scene, the method further includes: Obtain the original image set of the target scene; the original image set includes images of the target scene from at least one viewpoint; Based on the original image set, determine the point cloud of the target scene; Based on the point cloud of the target scene, an initial scene model of the target scene is generated; wherein, the first scene model is the initial scene model or the initial scene model after being optimized a preset number of times.

10. The construction method according to claim 9, characterized in that, The method further includes: Based on the target viewpoint and the second scene model, an image of the target scene is generated from the target viewpoint; the target viewpoint is different from the viewpoint corresponding to the original image.

11. A device for constructing a three-dimensional scene model, characterized in that, The device includes: The acquisition module is used to acquire a first scene model corresponding to the target scene; the first scene model includes multiple Gaussian elements, and the Gaussian elements are the three-dimensional Gaussian distribution of each data point in the point cloud of the target scene; the first scene model is used by the electronic device to generate a two-dimensional image of the target scene; The calculation module is used to determine the target attribute value of each Gaussian element in the first scene model, wherein the target attribute value represents the degree of influence of the Gaussian element on the quality of the two-dimensional image; An optimization module is used to optimize the first scene model based on the target attribute values ​​of each Gaussian element in the first scene model to obtain a second scene model corresponding to the target scene. The optimization includes densification and / or pruning.

12. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a computer program / instructions stored in the memory; the processor executes the computer program / instructions to enable the electronic device to implement the method for constructing a three-dimensional scene model as described in any one of claims 1-10.

13. A computer program product, characterized in that, The computer program product includes instructions that, when executed by an electronic device, cause the electronic device to perform the method for constructing a three-dimensional scene model as described in any one of claims 1-10.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer program instructions that, when executed by an electronic device, enable the electronic device to perform a method for constructing a three-dimensional scene model as described in any one of claims 1-10.