Gaussian model generation method for scene reconstruction and scene reconstruction method

By constructing anchor Gaussian sets and partitioned images to generate training sets and determining constraint expressions based on three-dimensional Gaussian parameters, the problem of poor scene reconstruction quality in the prior art is solved, and efficient and accurate scene reconstruction is achieved.

CN118982611BActive Publication Date: 2025-06-10SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410987435.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-06-10
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

The existing three-dimensional Gaussian scattering technology is difficult to capture complex geometric surfaces during scene rendering, resulting in poor quality of scene reconstruction.

Method used

By obtaining the sequence of images to be trained, the training data is determined, including building anchor Gaussian sets and partitioned images to generate training sets, and determining constraint expressions based on three-dimensional Gaussian parameters, the Gaussian model is trained to generate scene reconstruction Gaussian model.

Benefits of technology

The calculation efficiency and reconstruction accuracy of scene reconstruction are improved, the reconstruction quality of the model is enhanced, and it is suitable for large-scale scenario reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118982611B_ABST
    Figure CN118982611B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a Gaussian model for scene reconstruction and a scene reconstruction method, including: determining training data according to a sequence of images to be trained, training a Gaussian model based on the training data to generate a Gaussian model for scene reconstruction; determining training data according to a sequence of images to be trained, including: determining point cloud data according to a sequence of images to be trained and constructing an anchor Gaussian set at different levels; partitioning the sequence of images to be trained and determining the image acquisition device corresponding to each image partition, projecting the anchor points in the anchor Gaussian sets at different levels to determine all observable anchor points of each image acquisition device, and generating a training set for each image partition; determining the constraint expression of each image to be trained in the training image sequence according to the sequence of images to be trained and three-dimensional Gaussian parameters; using the training set of the image partition and the constraint expression as training data, solving the problem of poor scene reconstruction quality, ensuring computational efficiency and reconstruction accuracy, and improving the reconstruction quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method for generating a Gaussian model for scene reconstruction and a scene reconstruction method. Background Art

[0002] A Neural Radiance Field (NeRF) builds a 3D scene model by learning a continuous volumetric scene function that maps 3D coordinates and viewing directions to corresponding RGB colors and volume densities. This method can synthesize new views of complex scenes from a set of input images.

[0003] 3D Gaussian Splatting (3DGS) is another innovative method in the field of neural rendering that uses Gaussian functions to represent volumetric data. The technique involves placing 3D Gaussian kernels in the scene to approximate the spatial distribution of radiance and density. More importantly, 3DGS has an explicit representation method and can utilize rasterization rendering methods for real-time rendering. Subsequent work has mainly focused on improving rendering quality, further simplifying 3DGS to increase rendering speed, and extending its application to reflective surfaces.

[0004] When performing scene rendering, the 3D Gaussian scattering in the prior art either focuses on object-level reconstruction or has difficulty capturing complex geometric surfaces, resulting in poor reconstruction quality during scene reconstruction. Summary of the Invention

[0005] The present invention provides a method for generating a Gaussian model for scene reconstruction and a scene reconstruction method to solve the problem of poor scene reconstruction quality.

[0006] According to one aspect of the present invention, a method for generating a Gaussian model for scene reconstruction is provided, including:

[0007] Obtaining at least one sequence of images to be trained;

[0008] For each sequence of images to be trained, determining training data according to the sequence of images to be trained, and training a Gaussian model based on the training data to generate a Gaussian model for scene reconstruction;

[0009] Wherein, the determining of the training data according to the sequence of images to be trained includes:

[0010] Determining point cloud data according to the sequence of images to be trained, and constructing an anchor Gaussian set based on the point cloud data;

[0011] Partitioning the sequence of images to be trained to obtain each image partition, projecting the anchor points in the anchor Gaussian set, determining all observable anchor points corresponding to each image partition, and generating a training set for each image partition;

[0012] Determine the constraint expression of each to-be-trained image in the to-be-trained image sequence according to the to-be-trained image sequence and the three-dimensional Gaussian parameters;

[0013] Use the training set of each image partition and the corresponding constraint expression as training data.

[0014] According to another aspect of the present invention, a scene construction method is provided, including:

[0015] Obtain a to-be-constructed image sequence;

[0016] Input the to-be-constructed image sequence into a pre-generated scene reconstruction Gaussian model to generate reconstructed scene data;

[0017] Wherein, the scene reconstruction Gaussian model is generated by using the scene reconstruction Gaussian model generation method described in any embodiment of the present invention.

[0018] According to another aspect of the present invention, a scene reconstruction Gaussian model generation device is provided, including:

[0019] A to-be-trained sequence acquisition module, configured to acquire at least one to-be-trained image sequence;

[0020] A model generation module, configured to, for each to-be-trained image sequence, determine training data according to the to-be-trained image sequence, train a Gaussian model according to the training data, and generate a scene reconstruction Gaussian model;

[0021] Wherein, the model generation module includes:

[0022] An anchor Gaussian set construction unit, configured to determine point cloud data according to the to-be-trained image sequence, and construct anchor Gaussian sets at different levels according to the point cloud data;

[0023] A training set generation unit, configured to partition the to-be-trained image sequence to obtain image partitions, determine the image acquisition device corresponding to each image partition, project the anchor points in the anchor Gaussian sets at different levels, determine all observable anchor points of each image acquisition device, determine all observable anchor points corresponding to each image partition according to the projection result, and generate a training set for each image partition;

[0024] A constraint expression determination unit, configured to determine the constraint expression of each to-be-trained image in the to-be-trained image sequence according to the to-be-trained image sequence and the three-dimensional Gaussian parameters;

[0025] A training data determination unit, configured to use the training set of each image partition and the corresponding constraint expression as training data.

[0026] According to another aspect of the present invention, there is provided a scene construction device, comprising:

[0027] A sequence to be constructed acquisition module for acquiring an image sequence to be constructed;

[0028] A scene reconstruction module for inputting the image sequence to be constructed into a pre-generated scene reconstruction Gaussian model to generate reconstructed scene data;

[0029] Wherein, the scene reconstruction Gaussian model is generated by using the scene reconstruction Gaussian model generation method described in any embodiment of the present invention.

[0030] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0031] At least one processor, and a memory communicatively connected to the at least one processor;

[0032] Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the scene reconstruction Gaussian model generation method or the scene construction method described in any embodiment of the present invention.

[0033] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the scene reconstruction Gaussian model generation method or the scene construction method described in any embodiment of the present invention when executed.

[0034] According to another aspect of the present invention, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the scene reconstruction Gaussian model generation method or the scene construction method described in any embodiment of the present invention.

[0035] The technical solution of the embodiment of the present invention includes obtaining at least one image sequence to be trained; for each image sequence to be trained, determining training data according to the image sequence to be trained, and training a Gaussian model according to the training data to generate a scene reconstruction Gaussian model; wherein, determining training data according to the image sequence to be trained includes: determining point cloud data according to the image sequence to be trained, and constructing anchor Gaussian sets at different levels according to the point cloud data; partitioning the image sequence to be trained to obtain image partitions, determining the image acquisition device corresponding to each image partition, projecting the anchor points in the anchor Gaussian sets at different levels, determining all observable anchor points of each image acquisition device, determining all observable anchor points corresponding to each image partition according to the projection result, and generating a training set for each image partition; determining the constraint expression of each to-be-trained image in the to-be-trained image sequence according to the to-be-trained image sequence and three-dimensional Gaussian parameters; using the training set of each image partition and the corresponding constraint expression as training data, which solves the problem of poor scene reconstruction quality. By partitioning the to-be-trained image sequence to obtain image partitions, determining the image acquisition device corresponding to each image partition, then projecting the anchor points in the anchor Gaussian set respectively to determine all the anchor points that can be observed by each image acquisition device, and finally determining all the observable anchor points corresponding to the image partition according to the projection result to form a training set. By projecting in this way, it can be ensured that each partition has sufficient supervision, that is, it can be ensured that each partition can obtain the maximum degree of supervision, and each image acquisition device can present a complete image; at the same time, determining the constraint expression of each to-be-trained image, and using the training set and the constraint expression as training data to train the Gaussian model to generate a scene reconstruction Gaussian model. The scene reconstruction Gaussian model generated by the method provided in the embodiment of the present application can ensure the calculation efficiency and reconstruction accuracy during scene reconstruction, improve the model reconstruction quality, and can be applied to the reconstruction of large-scale scenes.

[0036] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 is a flowchart of a method for generating a scene reconstruction Gaussian model according to Embodiment 1 of the present invention;

[0039] Figure 2 It is a flowchart of a method for generating a scene reconstruction Gaussian model provided in Embodiment 2 of the present invention;

[0040] Figure 3 It is a visualization example diagram of different levels of effects provided in Embodiment 2 of the present invention;

[0041] Figure 4a It is an example diagram of the influence of a loss term on the final optimization process provided in Embodiment 2 of the present invention;

[0042] Figure 4b It is another example diagram of the influence of a loss term on the final optimization process provided in Embodiment 2 of the present invention;

[0043] Figure 4c It is another example diagram of the influence of a loss term on the final optimization process provided in Embodiment 2 of the present invention;

[0044] Figure 4d It is another example diagram of the influence of a loss term on the final optimization process provided in Embodiment 2 of the present invention;

[0045] Figure 4e It is another example diagram of the influence of a loss term on the final optimization process provided in Embodiment 2 of the present invention;

[0046] Figure 5 It is an example diagram of a comparison of visualization results provided in Embodiment 2 of the present invention;

[0047] Figure 6 It is a visualization example diagram of surface reconstruction provided in Embodiment 2 of the present invention;

[0048] Figure 7 It is a flowchart of a scene reconstruction method provided in Embodiment 3 of the present invention;

[0049] Figure 8 It is a schematic structural diagram of a device for generating a scene reconstruction Gaussian model provided in Embodiment 4 of the present invention;

[0050] Figure 9 It is a schematic structural diagram of a scene reconstruction device provided in Embodiment 5 of the present invention;

[0051] Figure 10 It is a schematic structural diagram of an electronic device provided in Embodiment 6 of the present invention. Detailed implementation manners

[0052] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0054] Embodiment 1

[0055] Figure 1 It is a flowchart of a method for generating a scene reconstruction Gaussian model provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of generating a high-precision scene reconstruction Gaussian model. This method can be executed by a scene reconstruction Gaussian model generation device, which can be implemented in the form of hardware and / or software, and the scene reconstruction Gaussian model generation device can be configured in an electronic device. As Figure 1 shown, the method includes:

[0056] S101. Obtain at least one sequence of images to be trained.

[0057] In this embodiment, a sequence of images to be trained can be understood as a sequence composed of one or more images to be trained for model training. An image to be trained can be understood as an image used for model training. The sequence of images to be trained includes one or more images to be trained, and each image to be trained can be sorted in a certain order. For example, it can be sorted according to the acquisition time or the acquisition position.

[0058] Images are collected in advance by an image acquisition device. The image acquisition device can move along a certain trajectory to collect images at different positions to form an image sequence. The image sequence can be saved locally or in a space such as the cloud or a database. When a model needs to be generated, the image sequence is obtained as the sequence of images to be trained.

[0059] S102. For each image sequence to be trained, determine training data according to the image sequence to be trained, and train a Gaussian model based on the training data to generate a scene reconstruction Gaussian model.

[0060] In this embodiment, the training data can be understood as the data used for model training; the scene reconstruction Gaussian model can be understood as a Gaussian model used for scene reconstruction. The scene reconstruction Gaussian model of the embodiments of the present application can be used to reconstruct large-scale scenes.

[0061] Process each image sequence to be trained to determine its corresponding training data. The training data includes a training set of image partitions and corresponding constraint expressions. The image partitions are determined by partitioning the image sequence to be trained. Train a Gaussian model based on the training data to obtain a scene reconstruction Gaussian model.

[0062] Among them, determining the training data according to the image sequence to be trained includes the following steps:

[0063] Step 1. Determine point cloud data according to the image sequence to be trained, and construct anchor Gaussian sets at different levels based on the point cloud data.

[0064] In this embodiment, the anchor Gaussian set can be understood as a set formed by anchor points. Each anchor Gaussian set includes one or more anchor points, that is, anchor Gaussians. The anchor Gaussian set level_i indicates that the anchor Gaussians in this set all correspond to the i-th level.

[0065] Among them, the anchor Gaussian is used to represent the local nearby area, and the anchors at different levels are used to describe features at different granularities. During the forward inference process, a multi-layer perceptron (MLP) is used to recover the parameters of these three-dimensional Gaussians. The parameters of the MLP are jointly trained with the features of the anchor Gaussians. The anchor Gaussians at different levels are used to represent features at different granularities.

[0066] Process the image sequence to be trained through a preset algorithm, model, etc., and convert the data of the images in the two-dimensional image sequence to be trained into three-dimensional point cloud data. For example, the image sequence to be trained can be processed by structure-from-motion (SFM) to infer three-dimensional information from the two-dimensional images to obtain point cloud data. Construct anchor Gaussian sets at different levels based on the obtained point cloud data.

[0067] In 3D reconstruction tasks, the training data usually includes information about the same object at different scales, especially in aerial images. Existing research has difficulty directly capturing features at different scales because there is no explicitly designed structure to capture the levels of detail. Therefore, the embodiments of the present application introduce a new representation method that combines the hierarchical structure of scene surface modeling with a flat form close to a plane.

[0068] Exemplarily, this embodiment provides a formula for constructing anchor Gaussian sets at different levels according to point cloud data :

[0069]

[0070] where level i is the anchor Gaussian set of the i-th level, indicating that all the anchor Gaussians in this set correspond to the i-th level and is also a numerical value; v 0 is the basic voxel size, which can be set according to the size of the scene is the voxel size of the i-th level, k is the number of forks. When k = 2, the hierarchical strategy generates an octree structure. K is the maximum number of levels, which can be set according to the scene size or manually.

[0071] Step 2: Partition the image sequence to be trained to obtain image partitions, determine the image acquisition device corresponding to each image partition, project the anchor points in the anchor Gaussian sets at different levels, determine all observable anchor points of each image acquisition device, and determine all observable anchor points corresponding to each image partition according to the projection results, so as to generate a training set for each image partition.

[0072] In this embodiment, the image partition can be understood as the partition obtained by dividing the image into different regions according to certain rules. The image acquisition device can be a camera, an infrared thermal imager, etc. The image acquisition device can move to acquire the images to be trained. Each image to be trained has an image acquisition device corresponding to it. The image acquisition devices corresponding to different images to be trained can actually be the same, but parameters such as position and angle may be different. That is, an image acquisition device can acquire different images to be trained at different positions by moving and other means.

[0073] Partition the training images in the training image sequence according to certain rules to obtain image partitions. For example, according to the filtering rule of uniform camera density, ensure that the number of cameras in each image partition c is approximately equal, that is, the number of training images in each image partition c is approximately equal. The number of image partitions can be preset, for example, set according to the scene scale, set according to the length of the training image sequence, and so on. After dividing the training images in the training image sequence into different image partitions, regard the training images corresponding to each image partition as the image acquisition devices corresponding to the image partitions, that is, each training image can be regarded as an image acquisition device. Project all the anchor points in the anchor Gaussian sets of different levels, project the anchor points in the image partition onto the image planes of all image acquisition devices, determine whether the image acquisition device can observe this anchor point, determine all the anchor points that each image acquisition device can observe, and then determine all the observable anchor points corresponding to each image partition according to the relationship between the image partition and the image acquisition device, obtain the observable anchor points corresponding to each image partition, and use the observable anchor points corresponding to the image partition as the training set of this image partition. The training set can be used for Gaussian model training.

[0074] Step 3: Determine the constraint expression of each training image in the training image sequence according to the training image sequence and the three-dimensional Gaussian parameters.

[0075] In this embodiment, the constraint expression can be understood as an expression used to constrain the relationship between data in the modeling process. For example, the constraint of geometric consistency, the constraint of multi-view Figure 1 consistency, the length constraint of the axis of the three-dimensional Gaussian kernel, etc.

[0076] Preset the types of constraints to be calculated. Based on the types of constraints, calculate according to the training images in the training image sequence and in combination with the three-dimensional Gaussian parameters to determine the constraint expression corresponding to each training image in the training image sequence. If there are multiple types of constraint expressions, after calculating each type of constraint expression, perform a comprehensive operation on all the constraint expressions to obtain the final constraint expression. For example, perform addition, weighting, etc.

[0077] Step 4: Use the training set of each image partition and the corresponding constraint expression as training data.

[0078] Since the training set of each image partition includes all the image acquisition devices within this image partition, i.e., the images to be trained, and all the observable anchor points of each image acquisition device, after determining the constraint expression of the image to be trained, the constraint expression corresponding to each image acquisition device is determined according to the constraint expression of the image to be trained, as well as all the observable anchor points of each image acquisition device. The observable anchor points and the constraint expression corresponding to each image acquisition device are used as a set of corresponding training data, and finally, the training data corresponding to the sequence of images to be trained is generated. Gaussian model training is performed using the training data to generate a scene reconstruction Gaussian model.

[0079] An embodiment of the present invention provides a method for generating a scene reconstruction Gaussian model, which solves the problem of poor scene reconstruction quality. By partitioning the sequence of images to be trained, the image acquisition devices corresponding to each image partition are determined after obtaining the image partitions. Then, the anchor points in the anchor Gaussian set are projected respectively to determine all the anchor points that can be observed by each image acquisition device. Finally, all the observable anchor points corresponding to the image partition are determined according to the projection results to form a training set. By projecting in this way, it can be ensured that each partition has sufficient supervision, that is, it can be ensured that each partition can obtain the maximum degree of supervision, and each image acquisition device can present a complete image. At the same time, the constraint expression of each image to be trained is determined, and the training set and the constraint expression are used as training data for Gaussian model training to generate a scene reconstruction Gaussian model. The scene reconstruction Gaussian model generated by the method provided in the embodiments of the present application can ensure the calculation efficiency and reconstruction accuracy during scene reconstruction, improve the model reconstruction quality, and can be applied to the reconstruction of large-scale scenes.

[0080] Embodiment 2

[0081] Figure 2 It is a flowchart of a method for generating a scene reconstruction Gaussian model provided by Embodiment 2 of the present invention. This embodiment is refined on the basis of the above embodiment. As Figure 2 shown, the method includes:

[0082] S201. Obtain at least one sequence of images to be trained.

[0083] For each sequence of images to be trained, its corresponding training data is determined through the steps of S202 - S214.

[0084] S202. For each sequence of images to be trained, determine the point cloud data according to the sequence of images to be trained, and construct anchor Gaussian sets at different levels according to the point cloud data.

[0085] S203. Partition the sequence of images to be trained to obtain image partitions, and determine the image acquisition devices corresponding to each image partition.

[0086] S204. For each image acquisition device in each image partition, determine all the anchor points within the image partition, and use all the anchor points within the image partition as the observable anchor points of the image acquisition device.

[0087] After completing the division of the image partition, determine all the anchor points within this image partition, and use all the anchor points within this image partition as the observable anchor points of each image acquisition device, that is, within an image partition, all the anchor points within this image partition can be observed by all the image acquisition devices within this image partition. This method can ensure that each image partition has sufficient supervision, project the anchor points of the image partition onto the image planes of all the image acquisition devices, and regard the image acquisition devices that can observe the anchor points in the image partition as the training objects of this image partition. Generally, an image partition includes multiple image acquisition devices and multiple anchor points.

[0088] S205. Determine the other anchor points outside the image partition as alternative anchor points.

[0089] In this embodiment, the alternative anchor points can be understood as the anchor points that may be observed by the image acquisition device. Determine the other anchor points outside this image partition and denote them as alternative anchor points.

[0090] S206. For each alternative anchor point, determine the anchor Gaussian set and the visibility of the level corresponding to the alternative anchor point. If the visibility of the level is less than the level corresponding to the anchor Gaussian set, determine whether the alternative anchor point can be projected into the viewing cone of the image acquisition device; if so, determine the alternative anchor point as the observable anchor point of the image acquisition device.

[0091] Pre-determine the visibility of the level of the alternative anchor point, determine the level corresponding to the alternative anchor point according to the anchor Gaussian set corresponding to the alternative purchase point, compare the visibility of the level and the level corresponding to the anchor Gaussian set. If the visibility of the level is less than the level corresponding to the anchor Gaussian set, this alternative anchor point can be used for projection; otherwise, this alternative anchor point cannot be used for projection, and continue to judge the next alternative anchor point. Project the alternative anchor points that can be used for projection onto the image plane of the image acquisition device, and determine whether this alternative anchor point can be projected into the viewing cone of the image acquisition device; if so, determine the alternative anchor point as the observable anchor point of the image acquisition device; otherwise, this alternative anchor point cannot be observed by this image acquisition device.

[0092] S207. Determine all the observable anchor points corresponding to the image partition according to all the observable anchor points of all the image acquisition devices corresponding to each image partition.

[0093] For each image partition, determine all the image acquisition devices within this image partition, and then determine all the observable anchor points of the image acquisition device, and use this part of the anchor points as all the observable anchor points corresponding to the image partition.

[0094] S208. Determine all observable anchor points corresponding to each image partition according to the projection result, and generate a training set for each image partition.

[0095] Exemplarily, an embodiment of the present application provides a method for determining the visibility of a hierarchy. In the rendering stage, the visibility of the hierarchy upper_level(d) is determined according to its position distance d from the viewpoint:

[0096]

[0097] where the parameter d max represents the point cloud data the maximum distance between points, which can be calculated before training; is used to represent the rounding operation. As the distance d increases, the low-level three-dimensional Gaussians involved in the rendering process will decrease. During the training process, the OctreeGS method can be used to add and remove anchor Gaussians. The visibility of the hierarchy can be used to determine whether an anchor point is used at level i to calculate color and depth; each anchor point has a corresponding visibility of the hierarchy.

[0098] In order to ensure that each image partition has sufficient supervision, an embodiment of the present application projects the anchor points of the image partition onto the image planes of all image acquisition devices, and regards the cameras that can observe the anchor points in the image partition as the training objects of this partition. This process is a greedy process and does not require any manual setting of thresholds, but can ensure that each image partition can obtain the maximum degree of supervision. Finally, the partition is extended by alternative anchor points outside the image partition, so that each image acquisition device can present a complete image. Therefore, an embodiment of the present application will project all anchor points onto the image planes of each image acquisition device, and add all visible anchor points to the training set of the corresponding partition through the anchor Gaussian set and the visibility of the hierarchy. In the prior art, the supervision of a specific area exists in other partitions, but due to the threshold-based selection strategy, it is not included in the training data of the current partition, resulting in insufficient partition supervision. The present application can well solve the problem of insufficient partition supervision and enable the image partition to obtain the maximum degree of supervision. The scalable partition strategy provided by the embodiment of the present application avoids the limitations of hardware and training time, enables the method of the present application to use more three-dimensional Gaussian kernels to represent large scenes, and even reaches the scale of gigabytes (Giga), realizing the reconstruction of large-scale scenes.

[0099] S209. For each image to be trained in the training image sequence, determine the minimum axis of the Gaussian kernel according to the three-dimensional Gaussian parameters.

[0100] Preset three-dimensional Gaussian parameters. The three-dimensional Gaussian parameters may include the axial lengths of three-dimensional Gaussian spheres, and the axial lengths of the three-dimensional Gaussian spheres may be randomly generated. The number of three-dimensional Gaussian spheres includes one or more. Determine the axial lengths of all three-dimensional Gaussian spheres according to the three-dimensional Gaussian parameters, and determine the minimum axis of the Gaussian kernel according to the axial lengths of each three-dimensional Gaussian sphere. For example, compare the axial lengths of all three-dimensional Gaussian spheres to determine the minimum axis.

[0101] As an optional embodiment, this optional embodiment further optimizes determining the minimum axis of the Gaussian kernel according to the three-dimensional Gaussian parameters as follows:

[0102] A1. Determine the shortest axes of all three-dimensional Gaussian spheres according to the three-dimensional Gaussian parameters.

[0103] A three-dimensional Gaussian sphere has three axes, so there are three corresponding axial lengths. Analyze the three-dimensional Gaussian parameters to determine the axial length of each axis of all three-dimensional Gaussian spheres. For each three-dimensional Gaussian sphere, compare the three axial lengths of this three-dimensional Gaussian sphere to determine the shortest axis.

[0104] A2. Calculate the average value of the shortest axes of all three-dimensional Gaussian spheres as the minimum axis of the Gaussian kernel.

[0105] Take the average of the shortest axes of all three-dimensional Gaussian spheres, and the obtained average value is the minimum axis of the Gaussian kernel.

[0106] Since the shortest axis of the three-dimensional Gaussian kernel itself can provide an accurate estimate of the normal vector, the embodiments of the present application strive to compress the minimum axis of each Gaussian kernel during the training phase. Exemplarily, the embodiments of the present application provide a calculation formula for the minimum axis of the Gaussian kernel:

[0107]

[0108] Among them, is a set containing three-dimensional Gaussians, represents the number of three-dimensional Gaussian spheres; M is the number of three-dimensional Gaussian spheres; s i is the length of the three axes of the i-th three-dimensional Gaussian sphere. The purpose of doing this is to limit the shortest axis of the three-dimensional Gaussian kernel to be perpendicular to the scene surface, so as to facilitate using 3DGS to fit the scene surface.

[0109] S210. Perform image rendering according to the image to be trained and the three-dimensional Gaussian parameters to determine the rendered image. Determine the simulated illumination image according to the rendered image and its corresponding embedding. Determine the appearance loss according to the rendered image, the simulated illumination image, and the image to be trained.

[0110] In this embodiment, the rendered image can be understood as the image obtained through image rendering during the scene reconstruction process. The simulated illumination image can be understood as the image formed under the illumination conditions of a specific perspective, and the simulated illumination image can accurately represent the illumination conditions of a specific perspective; the appearance loss can be understood as the loss generated by the simulated illumination image relative to the image to be trained.

[0111] The image to be trained is rendered through three-dimensional Gaussian parameters, and the three-dimensional Gaussian parameters can include position coordinates, covariance matrix describing the kernel arrangement, opacity, and spherical harmonic coefficients encoding view-related colors, etc. Image rendering is performed through the key attributes of the above-mentioned three-dimensional Gaussian. During the rendering process, the three-dimensional Gaussian is projected onto the two-dimensional Gaussian distribution of a specific perspective, and the final rendering output of this viewpoint can be generated through α-blending.

[0112] Exemplarily, an embodiment of the present application provides a rendering formula:

[0113]

[0114] All parameters are updated through a differentiable rendering process; where C is the color of a pixel point finally rendered; M is the number of three-dimensional Gaussian spheres to be rendered, set in advance according to the scale; c i is the color of the i-th three-dimensional Gaussian sphere, α i is the opacity of the i-th three-dimensional Gaussian sphere, T i is the transparency before reaching the i-th three-dimensional Gaussian sphere, α j is the opacity of the j-th three-dimensional Gaussian sphere.

[0115] Each pixel point in the image to be trained is rendered through the above formula to generate a rendered image. The pixel colors of the rendered image are adjusted according to the embedding corresponding to the rendered image to generate a simulated illumination image; the differences among the rendered image, the simulated illumination image, and the image to be trained are compared, and calculated according to different loss function calculation formulas to obtain the appearance loss.

[0116] When performing rendering, it is possible to determine whether to use level i to calculate color and depth according to the visibility of the anchor point level.

[0117] As an optional embodiment, this optional embodiment further optimizes determining the simulated illumination image according to the rendered image and its corresponding embedding to:

[0118] B1. Pixel color adjustment is performed on the rendered image and its corresponding embedding according to a pre-determined appearance model to obtain pixel color adjustment values.

[0119] In this embodiment, the appearance model can be understood as a model used to capture the appearance transformation of each image, which can be a calculation formula, a neural network model, etc. The appearance model is determined in advance, and an embedding is assigned to each rendered image, that is, the embeddings from different perspectives are generated in advance. At this time, each embedding corresponds to a training image, and the embedding corresponding to the rendered image is determined according to the corresponding relationship between the training image and the rendered image. The rendered image and the embedding are processed by the appearance model to adjust the pixel colors of the rendered image, and a pixel color adjustment value is obtained. For example, the rendered image and the embedding are used as the input or known parameters of the appearance model, and the output obtained is the pixel color adjustment value.

[0120] B2. Multiply the pixel color adjustment value by the rendered image to obtain a simulated illumination image.

[0121] Due to factors such as exposure and lighting conditions, directly applying existing representation methods to an outdoor photo collection may result in inaccurate reconstructions. These reconstructions will exhibit severe ghosting, over-smoothing, and more artifacts. Therefore, the embodiment of the present application introduces an appearance model to capture the appearance changes of each image. Similarly, this model is also co-learned with the planar representation.

[0122] The embodiment of the present application learns to model the appearance changes of each image in a low-dimensional latent space, such as exposure, lighting, weather, and post-processing effects. In this method, by assigning an embedding to each training perspective v, that is, the embedding corresponding to the rendered image; thus, the pixel color adjustment value of the image can be obtained using the appearance model φ. By multiplying these adjustment values by the rendered image I, a simulated illumination image can be obtained, which can accurately represent the lighting conditions of a specific perspective:

[0123] I a = φ(I, emb v )I,...(5)

[0124] where, I a is the simulated illumination image obtained after being processed by this appearance model, and emb v is the embedding of the rendered image v.

[0125] As an optional embodiment, this optional embodiment further optimizes the determination of the appearance loss according to the simulated illumination image and the training image as:

[0126] C1. Calculate the first loss according to the training image and the simulated illumination image.

[0127] C2. Calculate the second loss according to the rendered image and the training image.

[0128] In this embodiment, the first loss and the second loss are two losses obtained through calculation. They can be calculated through two different pre-set loss calculation methods, or can be obtained through the same method. The loss calculation method can be L1, SSIM, MSELoss, etc.

[0129] Exemplarily, through calculate the first loss, and calculate the second loss through SSIM(I, I 0 ).

[0130] C3. Weight the first loss and the second loss to determine the appearance loss.

[0131] Pre-set weights and perform a weighted operation on the first loss and the second loss to obtain the appearance loss.

[0132] Exemplarily, the embodiments of the present application provide a calculation method for the appearance loss:

[0133]

[0134] Wherein, is the appearance loss, is the first loss, and SSIM(I, I 0 ) is the second loss; I 0 is the real image to be trained, I a is the simulated illumination image, and I is the rendered image; λ is the weighting coefficient, which can be pre-set.

[0135] S211. Determine the geometric consistency of the depth map and the surface normal map according to the depth map and the surface normal map corresponding to the image to be trained.

[0136] Process and analyze the image to be trained to determine the depth map and the surface normal map of the image to be trained. Calculate the geometric consistency of the depth map and the surface normal map according to the mutual correlation of each pixel point in the common plane of the image to be trained. For each pixel point in the image to be trained, query the depths of the adjacent pixel points of the pixel point according to the depth map, calculate the differences in the depths of the adjacent pixel points of the pixel point, determine the normal of the pixel point through the surface normal map, and calculate the geometric consistency of the depth map and the surface normal map according to the difference in depth in combination with the normal.

[0137] Exemplarily, the embodiments of the present application provide a calculation method for the normal and a calculation method for the depth:

[0138] PGSR (Chen et al, 2024) introduces specific shape and distribution constraints into the original 3DGS method. This method essentially provides an accurate estimate of the normal vector, and the normal vector corresponds to the shortest axis of the Gaussian kernel. Utilizing this inherent property, PGSR was initially proposed for view-dependent normal vector rendering:

[0139] N = ∑ i∈M R c n i α i T i ,...(7)

[0140] where N is the normal; R c is the rotation matrix from camera coordinates to world coordinates; n i is the normal of the i-th 3D Gaussian sphere, α i is the transparency of the i-th 3D Gaussian sphere, and T i is the transmittance before reaching the i-th 3D Gaussian sphere.

[0141] Different from previous methods, PGSR adopts a different approach that does not directly render based on the spatial position of the 3DGS kernel. Instead, it assumes that the 3D Gaussian kernel can be tiled into a plane and fitted to the actual surface.

[0142] Then, the distance D from the origin of the image acquisition device to this Gaussian plane is rendered:

[0143] D = ∑ i∈M d i α i T i ,...(8)

[0144] where represents the distance from the camera origin to the i-th Gaussian kernel, μ i is the 3D coordinate of the center point of the i-th 3D Gaussian sphere, and T c is the 3D coordinate of the camera.

[0145] After obtaining the distance and normal of the plane, PGSR determines the corresponding depth map by the intersection of rays with these planes. This intersection operation ensures that the depth shape aligns with the plane assumed by the Gaussian kernel, thereby obtaining a depth map that accurately reflects the actual surface:

[0146]

[0147] where p is the 2D position on the image plane. represents the homogeneous coordinate of p, K is the internal parameter matrix of the image acquisition device, and N(p) is the normal vector corresponding to the pixel p.

[0148] As an alternative embodiment, this alternative embodiment further optimizes determining the geometric consistency of the depth map and the surface normal map based on the depth map and the surface normal map corresponding to the image to be trained to:

[0149] D1. For each pixel point in the image to be trained, obtain the previous depth of the previous pixel point corresponding to the pixel point, the next depth of the next pixel point corresponding to the pixel point, the left depth of the left pixel point corresponding to the pixel point, and the right depth of the right pixel point corresponding to the pixel point from the corresponding depth map.

[0150] In this embodiment, the previous pixel point can be understood as the pixel point adjacent above the pixel point in the image; the next pixel point can be understood as the pixel point adjacent below the pixel point in the image; the left pixel point can be understood as the pixel point adjacent to the left of the pixel point in the image; the right pixel point can be understood as the pixel point adjacent to the right of the pixel point in the image. The previous depth can be understood as the depth of the previous pixel point; the next depth is the depth of the next pixel point; the left depth can be understood as the depth of the left pixel point; the right depth is the depth of the right pixel point.

[0151] For each pixel point in the image to be trained, obtain the pixel point adjacent above this pixel point from the image to be trained according to the coordinates of the pixel point, denoted as the previous pixel point; obtain the pixel point adjacent below this pixel point, denoted as the next pixel point; obtain the pixel point adjacent to the left of this pixel point, denoted as the left pixel point; obtain the pixel point adjacent to the right of this pixel point, denoted as the right pixel point. Obtain the corresponding depth from the depth map corresponding to the image to be trained according to the coordinates of each pixel point, and obtain the previous depth of the previous pixel point, the next depth of the next pixel point, the left depth of the left pixel point, and the right depth of the right pixel point.

[0152] D2. Calculate the uncertainty coefficient based on the previous depth, the next depth, the left depth, and the right depth.

[0153] In this embodiment, the uncertainty coefficient is a parameter used to quantify the possibility that pixel i belongs to the surface boundary. Predetermine the calculation formula of the uncertainty coefficient, and substitute the previous depth, the next depth, the left depth, and the right depth into the corresponding formula to calculate the uncertainty coefficient. For example, calculate the difference between the previous depth and the next depth, denoted as the first difference, calculate the difference between the left depth and the right depth, denoted as the second difference. Calculate the product, ratio, etc. of the first difference and the second difference to obtain the uncertainty coefficient.

[0154] Exemplarily, an embodiment of the present application provides a calculation formula for the uncertainty coefficient:

[0155] w i =|(P i,0 -P i,1 )(P i,2 -P i,3 )|,……(10);

[0156] where, w i is the uncertainty coefficient; Pi,0 is the depth of the previous pixel point of pixel i in the coordinate system of the image acquisition device, i.e., the previous depth; P i,1 is the depth of the next pixel point of pixel i in the coordinate system of the image acquisition device, i.e., the next depth; P i,2 is the depth of the leftmost pixel point of pixel i in the coordinate system of the image acquisition device, i.e., the leftmost depth; P i,3 is the depth of the rightmost pixel point of pixel i in the coordinate system of the image acquisition device, i.e., the rightmost depth.

[0157] For pixels belonging to the depth discontinuity surface, an uncertainty coefficient w i is introduced to quantify the possibility that pixel i belongs to the surface boundary. By multiplying the hypothesis with this uncertainty factor, the inherent probability related to the surface edge of the pixel is considered. It can be seen that in regions with large depth differences, the value of the above dot product will decrease, thus identifying it as an edge region. Therefore, the influence of these regions on the depth and normal consistency constraints should be reduced.

[0158] D3. Obtain the normal value corresponding to the pixel point from the corresponding surface normal map.

[0159] Obtain the normal value corresponding to this pixel point from the surface normal map corresponding to the image to be trained according to the coordinates of the pixel point.

[0160] D4. Determine the geometric consistency based on the previous depth, next depth, leftmost depth, rightmost depth, combined with the normal value and the uncertainty coefficient.

[0161] Predetermine the calculation formula of geometric consistency, substitute the previous depth, next depth, leftmost depth, rightmost depth, combined with the normal value and the uncertainty coefficient into the formula for calculation, and obtain the geometric consistency.

[0162] D5. Determine the geometric consistency of the depth map and the surface normal map based on the geometric consistency of each pixel point.

[0163] Perform a comprehensive operation on the geometric consistency of each pixel point. For example, calculate the average value, maximum value, minimum value, etc., and determine the geometric consistency of the depth map and the surface normal map based on the result of the comprehensive operation.

[0164] For the boundary points of the image, for example, the top row of pixel points, which have no corresponding previous pixel points, and the bottom row of pixel points, which have no corresponding next pixel points, in this case, the geometric consistency of this part of pixel points can not be calculated. It is also possible to assign the previous pixel of the previous pixel point or the next pixel of the next pixel point as a preset value, etc.

[0165] To ensure that the flattened three-dimensional Gaussian is consistent with the actual surface, the embodiments of the present application require consistency between the unbiased depth map and the normal map for each perspective. At the same time, it is clear that controlling the consistency of each viewpoint has a positive impact on the final surface reconstruction quality. Ordinary 3DGS mainly relies on the image reconstruction loss and often encounters challenges in local overfitting optimization. In the absence of effective regularization techniques, the inherent ability of 3DGS to capture visual appearance leads to an inherent ambiguity in the relationship between the three-dimensional shape and brightness. Therefore, this ambiguity allows for the acceptance of degenerate solutions, resulting in a difference between the Gaussian shape estimate and the actual surface representation. To solve this problem, the embodiments of the present application introduce a direct regularization constraint aimed at strengthening the geometric consistency between the local depth map and the surface normal map:

[0166]

[0167] where P i,j is the three-dimensional position of the adjacent pixel j relative to pixel i in the up, down, left, and right directions in the camera coordinate system, that is, P i,0 is the depth of the previous pixel point of pixel i in the coordinate system of the image acquisition device, that is, the previous depth; P i,1 is the depth of the next pixel point of pixel i in the coordinate system of the image acquisition device, that is, the next depth; P i,2 is the depth of the left pixel point of pixel i in the coordinate system of the image acquisition device, that is, the left depth; P i,3 is the depth of the right pixel point of pixel i in the coordinate system of the image acquisition device, that is, the right depth. n i is the normal value of pixel point i. Assuming that the points represented by the four pixels above, below, left, and right of pixel point i are in the same plane in three-dimensional space, this assumption is of significant meaning because it proves the mutual correlation of adjacent pixels within a common plane. I is the set of image pixels, and after adding the absolute value, it represents the number of image pixel points.

[0168] S212. Convert the rendered image corresponding to the image to be trained into a grayscale image to obtain a grayscale image. Determine the multi-view photometric constraint according to the grayscale image, determine the multi-view geometric consistency constraint according to the image to be trained, and determine the multi-view consistency constraint based on the multi-view photometric constraint and the multi-view geometric consistency constraint.

[0169] Perform grayscale conversion on the rendered image corresponding to the image to be trained, convert the color image into a grayscale image, and supervise the geometric parameters through the grayscale image. Calculate the normalized cross-correlation based on the grayscale image and use it as the multi-view photometric constraint; convert the pixel points in the image to be trained to adjacent viewpoints to obtain the corresponding homogeneous coordinates, and determine the multi-view geometric consistency constraint according to the homogeneous coordinates; perform operations on the multi-view photometric constraint and the multi-view geometric consistency constraint, such as directly adding and summing, weighted summing, etc., and use the obtained result as the multi-view consistency constraint.

[0170] Exemplarily, an embodiment of the present application provides a method for determining the multi-view consistency constraint:

[0171] Single-view geometric regularization can maintain the consistency between depth and normal geometry. However, the geometric structures of multiple views are not completely consistent. Therefore, it is necessary to introduce multi-view geometric regularization to ensure the global consistency of the geometric structure. The embodiment of the present application uses the photometric multi-view consistency constraint based on plane patches to supervise the geometric structure. Specifically, the embodiment of the present application can render the normal and distance of each pixel to the plane. Then, the optimization of these geometric parameters can be achieved through the inter-view geometric consistency based on patches. For each pixel point p r , we can transform it to an adjacent viewpoint:

[0172]

[0173] where is the homogeneous coordinate of the pixel point p n , and the homogeneous H rn is calculated as:

[0174]

[0175] where R rn and T rn are the relative transformations from the reference view to the neighboring view. n r is the normal vector in the reference view, d r is the distance in the reference view, and K r is the internal parameter of the image acquisition device in the reference view.

[0176] To focus on geometric details, the embodiment of the present application converts the color image I into a grayscale image to supervise the geometric parameters. Then, the normalized cross-correlation (NCC) of the patches in the reference frame and the neighboring frame is used as an index to evaluate the photometric consistency:

[0177]

[0178] where is the valid region checked by the geometric consistency constraint, and pr is the pixel grayscale value from the reference perspective, H rn p n represents the pixel grayscale value in the reference perspective obtained by the new perspective through the above conversion matrix. Due to occlusion, the conversion operation may bring inconsistencies. Therefore, the embodiments of the present application will re-convert adjacent frames to the reference frame and use a threshold to filter out regions with obvious errors. The regions where these errors exceed the threshold are regarded as occlusion regions:

[0179]

[0180] Finally, the multi-view consistency constraint consists of two parts, the multi-view photometric constraint and the multi-view geometric consistency constraint:

[0181]

[0182] S213. Determine the constraint expression according to the minimum axis of the Gaussian kernel, the appearance loss, the geometric consistency of the depth map and the surface normal map, and the multi-view consistency constraint.

[0183] Accumulate according to the minimum axis of the Gaussian kernel, the target loss function, the geometric consistency of the depth map and the surface normal map, and the multi-view consistency constraint as the overall constraint expression

[0184]

[0185] S214. Use the training set of each image partition and the corresponding constraint expression as training data.

[0186] S215. Train a Gaussian model according to the training data to generate a scene reconstruction Gaussian model.

[0187] When training the Gaussian model in the embodiments of the present application, by adding a regularization term, a mesh can be generated according to the optimized Gaussian model; visual renderings and depth maps can be rendered from different advantageous positions, and then these rendered images and depth maps can be fused into a Truncated Signed Distance Function (TSDF) volume, and finally a high-quality three-dimensional surface mesh and point cloud can be created.

[0188] The method provided by the embodiments of this application can verify the performance through experiments. In the experiment, the side length of the 4K aerial image was reduced to one-fourth of the original size and aligned using the comparison method. Subsequently, Pixel-SFM (Lindenberger et al, 2021) was used to obtain the initial point cloud from the aerial images and perform Manhattan World alignment to make the y-axis perpendicular to the world coordinate axis of the ground plane. In the rubble, building, residence, and sci-art scenes, the entire scene was divided into 4×2 blocks, while for the largest scene - the campus, it was divided into 4×4 blocks. To ensure sufficient convergence, each block was trained for 120,000 iterations.

[0189] To compare the surface reconstruction results, Neuralangelo (Li et al, 2023), NeuS (Wang et al, 2021), and SuGaR (Antoine et al, 2023) were selected as comparison methods. Neuralangelo and Neus are methods based on the NeRF framework, while SuGaR is a method that relies on 3DGS. In addition, as a supplement to the above methods, VastGaussian (Lin et al, 2024) and MegaNeRF (Haithem et al, 2022) were used as additional comparison methods to analyze the rendering results.

[0190] Exemplarily, Figure 3 A visualization example diagram of different levels of effects is provided. The rendering images and normal maps at different levels of the same scene are visualized as rendering entities.

[0191] Exemplarily, Figure 4a An example diagram of the influence of the loss term on the final optimization process is provided, where the loss term is baseline; Figure 4b Another example diagram of the influence of the loss term on the final optimization process is provided, where the loss term is W / app, that is, the loss term is the appearance loss; Figure 4c Another example diagram of the influence of the loss term on the final optimization process is provided, where the loss term is W / (app+flatten), that is, the loss term is the appearance loss and the minimum axis loss of the Gaussian kernel; Figure 4d Another example diagram of the influence of the loss term on the final optimization process is provided, where the loss term is W / (app+flatten+local), that is, the loss term is the appearance loss, the minimum axis loss of the Gaussian kernel, and the geometric consistency of the depth map and the surface normal map; Figure 4eAnother example diagram showing the impact of the loss term on the final optimization process is provided. Here, the loss term is W / (app+flatten+local+mv), that is, the loss term is the appearance loss, the minimum axis loss of the Gaussian kernel, the geometric consistency of the depth map and the surface normal map, and the multi-view consistency constraint; Figure 4a It is an example diagram of the impact of the loss term in the prior art on the final optimization process. Figure 4b - 4e It is an example diagram of the impact of the loss term in the embodiment of the present application on the final optimization process, and the loss term gradually increases. Initially, using only the image loss as supervision cannot obtain a satisfactory surface reconstruction effect. Adding an appearance model can reduce some artifacts. However, without additional geometric supervision, adding flat regularization to the geometric structure will lead to a decrease in the expressiveness of the model. Nevertheless, adding the local loss can improve the surface quality. Finally, the introduction of multi-view regularization further improves the surface reconstruction performance, highlighting the superiority of the method in the embodiment of the present application in surface reconstruction.

[0192] Figure 5 An example diagram for comparing visualization results is provided. Among them, the first picture in the first row is the original picture, which is the collected ground truth data, and the remaining 5 pictures are the rendered views obtained by different rendering methods. The images in the second row are the corresponding depth maps from the same perspective, and the images in the third row are the corresponding normal maps from the same perspective; among them, GigaGS is the method provided by the embodiment of the present application, and Neuralangelo, SuGaR, NeuS, and PGSR are the methods in the prior art. Figure 5 The rendering results of the new perspective reconstruction method and the corresponding depth maps and normal maps are compared. It can be seen that the method in the embodiment of the present application is superior to the existing surface reconstruction methods in terms of surface texture and scene geometry. The existing NeRF-based methods lack fine details and show blurred and incorrect structures in image rendering. Similarly, the existing 3DGS-based methods are also troubled by artifacts, resulting in unsatisfactory rendering results.

[0193] In quantitative experiments, to maintain consistency, the embodiments of this application adopt the same dataset partitioning method as MegaNeRF. Visual quality metrics, namely PSNR, SSIM, and LPIPS, are used to compare the rendering quality on the test set. Exemplarily, Table 1 is a table showing the quantitative results of a rendering quality. Table 1 shows the SSIM↑, PSNR↑, and LPIPS↓ of the test views; and the best, second-best, and third-best results are indicated by bold, italic, and underlined (wavy) text. - indicates that training cannot continue due to insufficient memory. In Table 1, the test set results of the above methods are quantitatively compared, and two other large-scale novel view synthesis (NVS) methods are also added, only for the purpose of comparing the rendering quality. The last row of Table 1 is the method provided by the embodiments of this application, and the rest are methods in the prior art. It can be seen that the method provided by the embodiments of this application significantly improves the rendering quantitative results of the existing surface reconstruction methods and at the same time achieves performance comparable to that of the NVS methods.

[0194]

[0195] Table 1 Table showing the quantitative results of rendering quality

[0196] Exemplarily, Figure 6 A visualization example diagram of surface reconstruction is provided. The method provided by the embodiments of this application can extract high-quality meshes while ensuring high-quality rendering. This ability facilitates broader applications such as navigation, simulation, and virtual reality (VR).

[0197] Comprehensive experiments conducted on various datasets have demonstrated the effectiveness of the method proposed by the embodiments of this application in large-scale scene surface reconstruction.

[0198] The embodiment of the present application provides a method for generating a scene reconstruction Gaussian model. By using 3DGS to perform high-quality surface reconstruction of large-scale scenes, an efficient and scalable partitioning strategy is realized to meet the computational requirements for processing large-scale scenes. Different from traditional methods that rely on spatial distance metrics, the present application introduces a new grouping mechanism based on the mutual visibility of spatial regions captured by scene cameras, so that the scene can be divided into overlapping blocks that can be processed in parallel. Each block is independently optimized, thus realizing the distributed processing of scene data. Subsequently, the optimized blocks are seamlessly merged to reconstruct the complete scene, ensuring computational efficiency without affecting the reconstruction accuracy. Secondly, the present application proposes a new method that utilizes multi-view photometric and geometric consistency constraints within the framework of Level of Details (LoD). This method aims to strengthen the protection of geometric details at different scales of the reconstructed scene. By integrating the LoD representation into the constraint formulation, it can be ensured that the reconstruction process maintains fidelity and consistency in scenes of different complexities. Utilizing photometric and geometric information from multiple views, the method provided by the embodiment of the present application helps to robustly reconstruct intricate scene details while reducing artifacts and inconsistencies.

[0199] Embodiment III

[0200] Figure 7 The flowchart of a scene reconstruction method provided by Embodiment III of the present invention. This embodiment is applicable to the situation of performing high-precision scene reconstruction. This method can be executed by a scene reconstruction device, which can be implemented in the form of hardware and / or software. The scene reconstruction Gaussian model generation device can be configured in an electronic device. As Figure 7 shown, this method includes:

[0201] S301. Obtain the image sequence to be constructed.

[0202] In this embodiment, the image sequence to be constructed can be understood as the image sequence used for scene construction. The image sequence to be constructed includes at least one image, and all the images are sorted according to a certain rule to form the image sequence to be constructed.

[0203] When constructing a scene, images can be collected in advance and formed into an image sequence, and the collected image sequence is used as the image sequence to be constructed. For example, by installing an image acquisition device on a drone, the drone moves according to control or according to preset parameters such as trajectory and speed, and the image acquisition device performs image acquisition at a certain acquisition period to form an image sequence.

[0204] S302. Input the image sequence to be constructed into the pre-generated scene reconstruction Gaussian model to generate reconstructed scene data. Among them, the scene reconstruction Gaussian model is generated by using the scene reconstruction Gaussian model generation method described in any embodiment of the present application.

[0205] Pre-generate a scene reconstruction Gaussian model according to the scene reconstruction Gaussian model generation method provided in any embodiment of the present application. Take the image sequence to be constructed as the input and input it into the scene reconstruction Gaussian model. The scene reconstruction Gaussian model processes the image sequence to be constructed according to the parameters obtained during the training process to generate reconstructed scene data, completing the scene reconstruction. The reconstructed scene data can be used to represent the reconstructed scene.

[0206] The scene reconstruction Gaussian model can render visual renderings and depth maps from different advantageous positions, and then fuse these rendered images and depth maps into a Truncated Signed Distance Function (TSDF) volume, and finally create a high-quality three-dimensional surface mesh and point cloud to obtain the reconstructed scene data and complete the scene reconstruction.

[0207] The embodiment of the present invention provides a scene reconstruction method, which solves the problem of poor scene reconstruction quality. Pre-generate a high-precision scene reconstruction Gaussian model, process the image sequence to be constructed according to the scene reconstruction Gaussian model, and complete the scene reconstruction, which can ensure the calculation efficiency and reconstruction accuracy, improve the model reconstruction quality, and can be applied to the reconstruction of large-scale scenes.

[0208] Embodiment 4

[0209] Figure 8 It is a schematic structural diagram of a scene reconstruction Gaussian model generation device provided in Embodiment 4 of the present invention. As Figure 8 shown, the device includes: a training sequence acquisition module 41 and a model generation module 42.

[0210] The training sequence acquisition module 41 is used to acquire at least one training image sequence;

[0211] The model generation module 42 is used to, for each training image sequence, determine training data according to the training image sequence, and train a Gaussian model according to the training data to generate a scene reconstruction Gaussian model;

[0212] Among them, the model generation module 42 includes:

[0213] The anchor Gaussian set construction unit is used to determine point cloud data according to the training image sequence and construct an anchor Gaussian set according to the point cloud data;

[0214] A training set generation unit, configured to partition the to-be-trained image sequence to obtain each image partition, project the anchor points in the anchor Gaussian set, determine all observable anchor points corresponding to each image partition, and generate a training set for each image partition;

[0215] A constraint expression determination unit, configured to determine the constraint expression of each to-be-trained image in the to-be-trained image sequence according to the to-be-trained image sequence and three-dimensional Gaussian parameters;

[0216] A training data determination unit, configured to use the training set of each image partition and the corresponding constraint expression as training data.

[0217] An embodiment of the present invention provides a scene reconstruction Gaussian model generation device, which solves the problem of poor scene reconstruction quality. By partitioning the to-be-trained image sequence, after obtaining the image partition, the image acquisition device corresponding to each image partition is determined, and then the anchor points in the anchor Gaussian set are projected respectively to determine all the anchor points that can be observed by each image acquisition device. Finally, according to the projection result, all the observable anchor points corresponding to the image partition are determined to form a training set. By projecting in this way, it can be ensured that each partition has sufficient supervision, that is, it can be ensured that each partition can obtain the maximum degree of supervision, and each image acquisition device can present a complete image; at the same time, the constraint expression of each to-be-trained image is determined, and the training set and the constraint expression are used as training data to train the Gaussian model to generate a scene reconstruction Gaussian model. The scene reconstruction Gaussian model generated by the method provided in the embodiment of the present application can ensure the calculation efficiency and reconstruction accuracy when performing scene reconstruction, improve the model reconstruction quality, and can be applied to the reconstruction of large-scale scenes.

[0218] Optionally, the training set generation unit is specifically configured to: for each image acquisition device in each image partition, determine all the anchor points in the image partition, and use all the anchor points in the image partition as the observable anchor points of the image acquisition device; determine the other anchor points outside the image partition as alternative anchor points; for each alternative anchor point, determine the visibility of the anchor Gaussian set and the level corresponding to the alternative anchor point. If the visibility of the level is less than the level corresponding to the anchor Gaussian set, determine whether the alternative anchor point can be projected into the visual cone of the image acquisition device; if so, determine the alternative anchor point as the observable anchor point of the image acquisition device; determine all the observable anchor points corresponding to the image partition according to all the observable anchor points of all the image acquisition devices corresponding to each image partition.

[0219] Optionally, the constraint expression determination unit includes:

[0220] A minimum axis determination subunit, configured to, for each to-be-trained image in the to-be-trained image sequence, determine the minimum axis of the Gaussian kernel according to the three-dimensional Gaussian parameters;

[0221] An appearance loss determination subunit, configured to perform image rendering according to the to-be-trained image and three-dimensional Gaussian parameters, determine a rendered image, determine a simulated illumination image according to the rendered image and its corresponding embedding, and determine an appearance loss according to the rendered image, the simulated illumination image, and the to-be-trained image;

[0222] A geometric consistency determination subunit, configured to determine the geometric consistency of a depth map and a surface normal map according to the depth map and the surface normal map corresponding to the to-be-trained image;

[0223] A multi-view consistency determination subunit, configured to perform grayscale conversion on the rendered image corresponding to the to-be-trained image to obtain a grayscale image, determine a multi-view photometric constraint according to the grayscale image, determine a multi-view geometric consistency constraint according to the to-be-trained image, and determine a multi-view consistency constraint based on the multi-view photometric constraint and the multi-view geometric consistency constraint;

[0224] A constraint expression determination subunit, configured to determine a constraint expression according to the minimum axis of a Gaussian kernel, the appearance loss, the geometric consistency of the depth map and the surface normal map, and the multi-view consistency constraint.

[0225] Optionally, a minimum axis determination subunit is specifically configured to: determine the shortest axis of all three-dimensional Gaussian spheres according to the three-dimensional Gaussian parameters; calculate the average value of the shortest axes of all three-dimensional Gaussian spheres as the minimum axis of the Gaussian kernel.

[0226] Optionally, the appearance loss determination subunit is specifically configured to: perform pixel color adjustment on the rendered image and its corresponding embedding according to a pre-determined appearance model to obtain a pixel color adjustment value; multiply the pixel color adjustment value by the rendered image to obtain a simulated illumination image.

[0227] Optionally, the appearance loss determination subunit is specifically configured to: calculate a first loss according to the to-be-trained image and the simulated illumination image; calculate a second loss according to the rendered image and the to-be-trained image; perform weighting on the first loss and the second loss to determine the appearance loss.

[0228] Optionally, the geometric consistency determination subunit is specifically configured to: for each pixel point in the to-be-trained image, obtain the previous depth of the previous pixel point corresponding to the pixel point, the next depth of the next pixel point corresponding to the pixel point, the left depth of the left pixel point corresponding to the pixel point, and the right depth of the right pixel point corresponding to the pixel point from the corresponding depth map; calculate an uncertainty coefficient based on the previous depth, the next depth, the left depth, and the right depth; obtain the normal value corresponding to the pixel point from the corresponding surface normal map; determine the geometric consistency based on the previous depth, the next depth, the left depth, the right depth, in combination with the normal value and the uncertainty coefficient; and determine the geometric consistency of the depth map and the surface normal map based on the geometric consistency of each pixel point.

[0229] The scene reconstruction Gaussian model generation device provided by the embodiments of the present invention can execute the scene reconstruction Gaussian model generation method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0230] Embodiment 5

[0231] Figure 9 It is a schematic structural diagram of a scene reconstruction device provided in Embodiment 5 of the present invention. As Figure 9 shown, the device includes: a to-be-constructed sequence acquisition module 51 and a scene reconstruction module 52.

[0232] The to-be-constructed sequence acquisition module 51 is configured to acquire a to-be-constructed image sequence;

[0233] The scene reconstruction module 52 is configured to input the to-be-constructed image sequence into a pre-generated scene reconstruction Gaussian model to generate reconstructed scene data;

[0234] Among them, the scene reconstruction Gaussian model is generated by using the scene reconstruction Gaussian model generation method described in any embodiment of the present application.

[0235] The embodiments of the present invention provide a scene reconstruction device, which solves the problem of poor scene reconstruction quality. A high-precision scene reconstruction Gaussian model is pre-generated, and the to-be-constructed image sequence is processed according to the scene reconstruction Gaussian model to complete scene reconstruction, which can ensure the calculation efficiency and the reconstruction accuracy, improve the model reconstruction quality, and can be applied to the reconstruction of large-scale scenes.

[0236] The scene construction device provided by the embodiments of the present invention can execute the scene construction method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0237] Embodiment 6

[0238] Figure 10A schematic structural diagram of an electronic device is provided. The electronic device 60 can be used to implement the method provided by any embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0239] As Figure 10 shown, the electronic device 60 includes at least one processor 61, and a memory communicatively connected to the at least one processor 61, such as a read-only memory (ROM) 62, a random access memory (RAM) 63, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 61 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 62 or the computer program loaded from the storage unit 68 into the random access memory (RAM) 63. In the RAM 63, various programs and data required for the operation of the electronic device 60 can also be stored. The processor 61, the ROM 62, and the RAM 63 are connected to each other through a bus 64. The input / output (I / O) interface 65 is also connected to the bus 64.

[0240] Multiple components in the electronic device 60 are connected to the I / O interface 65, including: an input unit 66, such as a keyboard, a mouse, etc.; an output unit 67, such as various types of displays, speakers, etc.; a storage unit 68, such as a disk, an optical disc, etc.; and a communication unit 69, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 69 allows the electronic device 60 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0241] The processor 61 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 61 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 61 executes the various methods and processes described above, such as the scene reconstruction Gaussian model generation method or the scene construction method.

[0242] In some embodiments, the scene reconstruction Gaussian model generation method or the scene construction method may be implemented as a computer program, which is tangibly embodied in a computer-readable storage medium, such as storage unit 68. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 60 via the ROM 62 and / or the communication unit 69. When the computer program is loaded into the RAM 63 and executed by the processor 61, one or more steps of the scene reconstruction Gaussian model generation method or the scene construction method described above may be performed. Alternatively, in other embodiments, the processor 61 may be configured to execute the scene reconstruction Gaussian model generation method or the scene construction method by any other suitable means (e.g., by means of firmware).

[0243] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0244] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0245] An embodiment of the present invention provides a computer program product, which includes a computer program that, when executed by a processor, implements the scene reconstruction Gaussian model generation method or the scene construction method described in any embodiment of the present invention.

[0246] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0247] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0248] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0249] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0250] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0251] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating a Gaussian model for scene reconstruction, characterized in that: include: Obtain at least one image sequence to be trained; For each image sequence to be trained, training data is determined according to the image sequence to be trained, a Gaussian model is trained according to the training data, and a scene reconstruction Gaussian model is generated; Wherein, determining the training data according to the image sequence to be trained includes: Determining point cloud data according to the image sequence to be trained, and constructing anchor Gaussian sets of different levels according to the point cloud data; Partitioning the image sequence to be trained to obtain image partitions, determining the image acquisition device corresponding to each image partition, projecting the anchor points in the anchor Gaussian sets at different levels to determine all observable anchor points of each image acquisition device, determining all observable anchor points corresponding to each image partition according to the projection results, and generating a training set for each image partition; Determine a constraint expression of each image to be trained in the training image sequence according to the image sequence to be trained and the three-dimensional Gaussian parameters; The training set of each image partition and the corresponding constraint expression are used as training data; The projecting of the anchor points in the anchor Gaussian sets at different levels to determine all observable anchor points of each of the image acquisition devices, and determining all observable anchor points corresponding to each image partition according to the projection results, includes: For each image acquisition device in each image partition, determine all anchor points in the image partition, and use all anchor points in the image partition as observable anchor points of the image acquisition device; Determine other anchor points outside the image partition as candidate anchor points; For each candidate anchor point, determine the visibility of the anchor Gaussian set and the layer corresponding to the candidate anchor point, and if the visibility of the layer is less than the layer corresponding to the anchor Gaussian set, determine whether the candidate anchor point can be projected to the viewing cone of the image acquisition device; If so, determining the candidate anchor point as an observable anchor point of the image acquisition device; All observable anchor points corresponding to the image partition are determined according to all observable anchor points of all image acquisition devices corresponding to each of the image partitions.

2. The method according to claim 1, characterized in that The step of determining the constraint expression of each image to be trained in the training image sequence according to the image sequence to be trained and the three-dimensional Gaussian parameters includes: For each to-be-trained image in the training image sequence, determining the minimum axis of the Gaussian kernel according to the three-dimensional Gaussian parameters; Perform image rendering according to the image to be trained and the three-dimensional Gaussian parameters to determine a rendered image, determine a simulated illumination image according to the rendered image and its corresponding embedding, and determine an appearance loss according to the rendered image, the simulated illumination image, and the image to be trained; Determining the geometric consistency of the depth map and the surface normal map according to the depth map and the surface normal map corresponding to the image to be trained; Performing grayscale conversion on the rendered image corresponding to the image to be trained to obtain a grayscale image, determining a multi-view photometric constraint according to the grayscale image, determining a multi-view geometric consistency constraint according to the image to be trained, and determining a multi-view consistency constraint based on the multi-view photometric constraint and the multi-view geometric consistency constraint; The constraint expression is determined based on the minimum axis of the Gaussian kernel, the appearance loss, the geometric consistency of the depth map and the surface normal map, and the multi-view consistency constraint.

3. The method according to claim 2, characterized in that The step of determining the minimum axis of the Gaussian kernel according to the three-dimensional Gaussian parameters includes: Determine the shortest axes of all three-dimensional Gaussian spheres according to the three-dimensional Gaussian parameters; Calculate the average of the shortest axes of all three-dimensional Gaussian spheres as the minimum axis of the Gaussian kernel.

4. The method according to claim 2, characterized in that: The step of determining a simulated lighting image based on the rendered image and its corresponding embedding includes: Performing pixel color adjustment on the rendered image and its corresponding embedding according to a predetermined appearance model to obtain a pixel color adjustment value; The pixel color adjustment value is multiplied by the rendered image to obtain a simulated lighting image.

5. The method according to claim 2, characterized in that: The determining the appearance loss according to the rendered image, the simulated illumination image and the image to be trained comprises: Calculate a first loss according to the image to be trained and the simulated illumination image; Calculate a second loss based on the rendered image and the image to be trained; The first loss and the second loss are weighted to determine an appearance loss.

6. The method according to claim 2, characterized in that The step of determining the geometric consistency of the depth map and the surface normal map according to the depth map and the surface normal map corresponding to the image to be trained includes: For each pixel in the image to be trained, obtain from the corresponding depth map the previous depth of the previous pixel corresponding to the pixel, the next depth of the next pixel corresponding to the pixel, the left depth of the left pixel corresponding to the pixel, and the right depth of the right pixel corresponding to the pixel; Calculate an uncertainty coefficient based on the previous depth, the next depth, the left depth, and the right depth; Obtaining the normal value corresponding to the pixel point from the corresponding surface normal map; Determine geometric consistency based on the previous depth, the next depth, the left depth, and the right depth in combination with the normal value and the uncertainty coefficient; The geometric consistency of the depth map and the surface normal map is determined based on the geometric consistency of each of the pixel points.

7. A scene construction method, characterized in that: include: Obtain the image sequence to be constructed; Inputting the image sequence to be constructed into a pre-generated scene reconstruction Gaussian model to generate reconstructed scene data; The scene reconstruction Gaussian model is generated by using the scene reconstruction Gaussian model generation method according to any one of claims 1 to 6.

8. A scene reconstruction Gaussian model generation device, characterized in that: include: A training sequence acquisition module, used to acquire at least one training image sequence; A model generation module, used for determining training data for each image sequence to be trained according to the image sequence to be trained, training a Gaussian model according to the training data, and generating a scene reconstruction Gaussian model; Wherein, the model generation module includes: An anchor Gaussian set construction unit, used to determine point cloud data according to the image sequence to be trained, and to construct anchor Gaussian sets of different levels according to the point cloud data; A training set generating unit, used for partitioning the image sequence to be trained to obtain image partitions, determining the image acquisition device corresponding to each image partition, projecting the anchor points in the anchor Gaussian sets at different levels, determining all observable anchor points of each image acquisition device, determining all observable anchor points corresponding to each image partition according to the projection results, and generating a training set for each image partition; A constraint expression determination unit, used to determine the constraint expression of each image to be trained in the training image sequence according to the image sequence to be trained and the three-dimensional Gaussian parameters; A training data determination unit, used to use the training set of each image partition and the corresponding constraint expression as training data; The training set generation unit is specifically used to: for each image acquisition device in each image partition, determine all anchor points in the image partition, and use all anchor points in the image partition as observable anchor points of the image acquisition device; determine other anchor points outside the image partition as candidate anchor points; for each candidate anchor point, determine the anchor Gaussian set and the visibility of the level corresponding to the candidate anchor point, if the visibility of the level is less than the level corresponding to the anchor Gaussian set, determine whether the candidate anchor point can be projected to the visual cone of the image acquisition device; if so, determine the candidate anchor point as the observable anchor point of the image acquisition device; determine all observable anchor points corresponding to the image partition according to all observable anchor points of all image acquisition devices corresponding to each of the image partitions.

9. A scene construction device, characterized in that: include: A sequence acquisition module to be constructed is used to obtain an image sequence to be constructed; A scene reconstruction module, used for inputting the image sequence to be constructed into a pre-generated scene reconstruction Gaussian model to generate reconstructed scene data; The scene reconstruction Gaussian model is generated by using the scene reconstruction Gaussian model generation method according to any one of claims 1 to 6.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor, and a memory communicatively coupled to the at least one processor; Wherein, the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the scene reconstruction Gaussian model generation method described in any one of claims 1 to 6 or the scene construction method described in claim 7.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the scene reconstruction Gaussian model generation method described in any one of claims 1 to 6 or the scene construction method described in claim 7 when executed.

12. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the scene reconstruction Gaussian model generation method according to any one of claims 1 to 6 or the scene construction method according to claim 7.

Citation Information

Patent Citations

  • Three-dimensional scene reconstruction method and electronic equipment

    CN118365805A

  • Method, computer device and storage medium for real-time urban scene reconstruction

    US20220351463A1