Three-dimensional scene reconstruction method and device and electronic equipment
By building and converting sparse point cloud models, efficient reconstruction and rendering of dynamic three-dimensional scenes is achieved, and the problems of large computing overhead and large storage occupancy in the existing technology are solved, high-quality rendering effect is maintained and storage needs are reduced.
Patent Information
- Application Number
- CN202510038727.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
AI Technical Summary
When dealing with dynamic scenarios, the existing three-dimensional scene reconstruction method has large computing overhead, large storage space and long training time, making it difficult to effectively reduce storage demand and maintain high-quality rendering effects.
By building a sparse point cloud model of the target scene, voxelization is performed to obtain a 4D anchor point model, converted to a 4D Gaussian point cloud model, and converted to a 3D Gaussian point cloud model under a specific time slice for training, reducing storage requirements and maintaining high-quality rendering.
This method only needs to store a relatively small number of 4D anchor points and a small fully connected network, reducing the storage overhead of dynamic scene reconstruction, while ensuring the rendering quality and speed of Gaussian technology, and improving the availability of three-dimensional dynamic scene reconstruction technology.
Smart Images

Figure CN119963732A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the technical field of three-dimensional scene reconstruction, and more specifically, to a three-dimensional scene reconstruction method, device and electronic device. Background Art
[0002] 3D scene reconstruction is a technology that converts objects and scenes in the real world into digital 3D models through acquisition, calculation and rendering. Traditional 3D scene reconstruction methods include data acquisition, feature extraction and matching, sparse reconstruction, dense reconstruction, mesh reconstruction, texture mapping, rendering and display. This technology is widely used and plays an important role in virtual reality, augmented reality, architecture, game design, medical imaging and other fields.
[0003] Traditional 3D scene reconstruction is often oriented towards static scenes and obtains static results. In recent years, with the breakthrough of new technologies such as Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS), a new paradigm of 3D scene reconstruction has been proposed, and it has become possible to reconstruct a dynamic 3D scene. These technologies are no longer based on reconstructing and rendering meshes, but directly modeling the interaction between light and the scene to generate a voxel or Gaussian-based scene representation. This method allows the capture of complex light and shadow changes, transparent objects, and dynamic objects, and is more efficient and accurate when processing dynamic scenes. New perspective synthesis can be performed at any new perspective, realizing browsing of 3D scenes by looking.
[0004] 3DGS technology uses a set of anisotropic three-dimensional Gaussian ellipsoids to represent the scene. Each three-dimensional Gaussian is composed of many parameters, including the center position μ∈R 3 and a covariance matrix ∑∈R 3×3 In order to ensure that a valid positive semidefinite covariance matrix is obtained during the optimization process, the covariance matrix ∑ is decomposed into a scaling matrix S and a rotation matrix R, which together describe the geometric characteristics of a three-dimensional Gaussian ellipsoid: ∑ = RSS T R T The scaling matrix is a diagonal matrix, while the rotation matrix is represented by quaternions in the calculation. In order to render an image of a certain perspective, the three-dimensional Gaussian volume is projected onto a two-dimensional plane based on the perspective transformation matrix W and the affine transformation matrix J: ∑′=JWΣW T J T , and then use volume alpha blending to integrate the weighted colors in depth order to get the final rendered color of a pixel at that perspective: By calculating the difference between the rendered image and the real image and performing gradient descent optimization, a 3D scene can be represented by millions of 3D Gaussian volumes, which can realize arbitrary roaming browsing of the scene from a new perspective. In the real outdoor scene, 3DGS technology can achieve a peak signal-to-noise ratio (PSNR) value of 27.21 at the test perspective, achieving a realistic 3D scene browsing effect.
[0005] 3DGS technology is a technology launched for the three-dimensional reconstruction and new perspective synthesis tasks of static scenes. In order to expand this technology to the reconstruction of dynamic three-dimensional scenes, the existing method is to directly increase the time transformation dimension of the three-dimensional Gaussian and expand it to four dimensions. This method has a huge computational overhead during the optimization process, occupies more storage space, and takes longer to train. Summary of the invention
[0006] An objective of the embodiments of the present disclosure is to provide a three-dimensional scene reconstruction method, device and electronic device.
[0007] According to a first aspect of an embodiment of the present disclosure, a three-dimensional scene reconstruction method is provided, comprising:
[0008] Constructing a sparse point cloud model of the target scene according to multi-view images associated with the target scene; the multi-view images are image sequences obtained by shooting the target scene from multiple perspectives within a target time interval;
[0009] voxelize the sparse point cloud model to obtain a 4D anchor point model;
[0010] Convert the 4D anchor point model into a 4D Gaussian point cloud model;
[0011] Determine a target time slice corresponding to a target shooting moment in the target time interval;
[0012] The 4D Gaussian point cloud model under the target time slice is converted into a 3D Gaussian point cloud model at the target shooting moment, and a plurality of 3D Gaussian points in the 3D Gaussian point cloud model are projected onto the camera imaging plane to obtain a rendered image, and the 3D Gaussian point cloud model is trained.
[0013] Optionally, the method further includes:
[0014] Acquire, from the multi-view images, true values for expressing attribute information of points on the objects in the target scene at corresponding viewing angles;
[0015] During the training of the 3D Gaussian point cloud model, the calculated values of the attribute information of the points on the object represented by the multiple pixels in the rendered image are obtained, and the corresponding true values are used as the optimization targets of the calculated values of the obtained attribute information. The training process of the 3D Gaussian point cloud model is supervised to complete the dynamic modeling process of the target scene.
[0016] Optionally, the attribute information includes at least one of pixel value, brightness, contrast, and structure.
[0017] Optionally, converting the 4D anchor point model into a 4D Gaussian point cloud model includes:
[0018] Initializing the attribute feature vector of each anchor point in the 4D anchor point model;
[0019] Processing the attribute feature vector of each anchor point in the 4D anchor point model based on a fully connected neural network to obtain 4D Gaussian parameters of multiple Gaussian bodies corresponding to the anchor points;
[0020] The 4D Gaussian point cloud model is obtained according to the 4D Gaussian parameters of each Gaussian body.
[0021] Optionally, converting the 4D Gaussian point cloud model of the target time slice into a 3D Gaussian point cloud model of the target shooting moment includes:
[0022] Determine the 3D Gaussian distribution of the target at the time of shooting according to the 4D Gaussian parameters of the 4D Gaussian point cloud model at the target time slice;
[0023] A 3D Gaussian point cloud model of the target at the shooting moment is obtained according to the 3D Gaussian distribution of the target at the shooting moment.
[0024] Optionally, the Gaussian parameters include at least one of opacity, color, size, rotation, and center position.
[0025] Optionally, the training of the 3D Gaussian point cloud model includes:
[0026] Optimizing the attribute feature vector of each anchor point in the fully connected neural network and the 4D anchor point model.
[0027] Optionally, the determining a target time slice corresponding to a target shooting moment in the target time interval includes:
[0028] Traversing the shooting moments in the target time interval, and taking the currently traversed shooting moment as the target shooting moment;
[0029] Determine a target time slice corresponding to the target shooting moment.
[0030] Optionally, the method further includes:
[0031] According to the optimized attribute feature vector of each anchor point in the 4D anchor point model and the optimized fully connected neural network, an image sequence of the target scene under the test viewing angle within the target time interval is obtained.
[0032] According to a second aspect of the present disclosure, a three-dimensional scene reconstruction device is provided, comprising:
[0033] A sparse point cloud construction module, used to construct a sparse point cloud model of the target scene according to multi-view images associated with the target scene; the multi-view images are image sequences obtained by shooting the target scene from multiple perspectives within a target time interval;
[0034] A point cloud voxelization module, used for voxelizing the sparse point cloud model to obtain a 4D anchor point model;
[0035] An anchor point Gaussian conversion module, used to convert the 4D anchor point model into a 4D Gaussian point cloud model;
[0036] A time slice determination module, used to determine a target time slice corresponding to a target shooting moment in the target time interval;
[0037] The Gaussian model training module is used to convert the 4D Gaussian point cloud model under the target time slice into a 3D Gaussian point cloud model at the target shooting moment, and to train the 3D Gaussian point cloud model by projecting multiple 3D Gaussian points in the 3D Gaussian point cloud model onto the camera imaging plane to obtain a rendered image.
[0038] According to a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the method described in the first aspect of the present disclosure under the control of the computer program.
[0039] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect of the present disclosure is implemented.
[0040] Through the embodiments of the present invention, a sparse point cloud model of the target scene is constructed according to the multi-view images associated with the target scene; the sparse point cloud model is voxelized to obtain a 4D anchor point model; the 4D anchor point model is converted into a 4D Gaussian point cloud model; the target time slice corresponding to the target shooting moment in the target time interval is determined; the 4D Gaussian point cloud model under the target time slice is converted into a 3D Gaussian point cloud model of the target shooting moment, and a rendered image is obtained by projecting multiple 3D Gaussian points in the 3D Gaussian point cloud model onto the camera imaging plane, and the 3D Gaussian point cloud model is trained, so that the three-dimensional scene reconstruction method only needs to store a relatively small number of 4D anchor points and a small fully connected network, thereby reducing the storage overhead of dynamic scene reconstruction, and can effectively reduce the storage requirement expenditure for reconstructing three-dimensional dynamic scenes using Gaussian technology, while also ensuring the rendering quality and rendering speed that Gaussian technology should have, and improving the availability of technology based on Gaussian reconstruction of three-dimensional dynamic scenes.
[0041] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0043] Figure 1 is a block diagram showing a hardware configuration of an electronic device that can implement an embodiment of the present disclosure;
[0044] Figure 2 is a flow chart of a three-dimensional scene reconstruction method according to an embodiment of the present disclosure;
[0045] Figure 3 is a flowchart of an example of a three-dimensional scene reconstruction method according to an embodiment of the present disclosure;
[0046] Figure 4 is a block diagram of a three-dimensional scene reconstruction device according to an embodiment of the present disclosure;
[0047] Figure 5 is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0048] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention unless otherwise specifically stated.
[0049] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0050] Technologies, methods and equipment known to persons of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and equipment should be considered part of the specification.
[0051] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0052] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0053] <Hardware Configuration>
[0054] Figure 1 1 is a block diagram showing a hardware configuration of an electronic device 1000 that can implement an embodiment of the present disclosure.
[0055] The electronic device 1000 may be a portable computer, a desktop computer, a mobile phone, a tablet computer, etc. Figure 1 As shown, the electronic device 1000 may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, and the like. Among them, the processor 1100 may be a processor CPU, a microprocessor MCU, and the like. The memory 1200 includes, for example, a ROM (read-only memory), a RAM (random access memory), a non-volatile memory such as a hard disk, and the like. The interface device 1300 includes, for example, a USB interface, a headphone interface, and the like. The communication device 1400 is, for example, capable of wired or wireless communication, and may specifically include Wifi communication, Bluetooth communication, 2G / 3G / 4G / 5G communication, and the like. The display device 1500 is, for example, a liquid crystal display screen, a touch display screen, and the like. The input device 1600 may include, for example, a touch screen, a keyboard, a somatosensory input, and the like. The user may input / output voice information through the speaker 1700 and the microphone 1800.
[0056] Figure 1 The electronic device shown is merely illustrative and does not in any way imply any limitation on the present disclosure, its application or use. In the embodiments of the present disclosure, the memory 1200 of the electronic device 1000 is used to store instructions, and the instructions are used to control the processor 1100 to operate to perform any one of the methods provided in the embodiments of the present disclosure. It should be understood by those skilled in the art that although Figure 1In the electronic device 1000, multiple devices are shown, but the present disclosure may only involve some of the devices, for example, the electronic device 1000 only involves the processor 1100 and the memory 1200. A technician can design instructions according to the scheme disclosed in the present disclosure. How instructions control the processor to operate is well known in the art, so it will not be described in detail here.
[0057] <Method Example>
[0058] The present disclosure provides a 3D scene reconstruction method, which can be implemented by an electronic device. For example, the electronic device can be Figure 1 An electronic device 1000 is shown.
[0059] Figure 2 The figure is a flowchart of a three-dimensional scene reconstruction method according to an embodiment of the present disclosure.
[0060] like Figure 2 As shown, the three-dimensional scene reconstruction method includes steps S2100 to S2500 as shown below:
[0061] Step S2100, constructing a sparse point cloud model of the target scene based on the multi-view images associated with the target scene, where the multi-view images are image sequences obtained by shooting the target scene from multiple perspectives within a target time interval.
[0062] In this embodiment, multiple cameras may take photos of the target scene to be modeled from multiple perspectives within a target time interval to obtain an image sequence corresponding to each camera, for example, a photo. Specifically, multiple cameras take photos at the same shooting frequency, and in the image sequences obtained by the multiple cameras, the same frame of image corresponds to the same shooting time.
[0063] In order to ensure the accuracy of modeling, the overlap between each image can be required to be relatively high, that is, the difference in viewing angle between different images cannot be too large, so as to reflect the smoothness of the change in viewing angle. In addition, the camera parameters can be kept unchanged during shooting at different viewing angles.
[0064] In one embodiment, pose estimation (estimating the pose of the camera under various viewing angles) can be performed based on the captured image sequence to obtain a sparse point cloud of the scene (the point cloud is obtained through pose estimation).
[0065] In another embodiment, a sparse point cloud model of the target scene may be constructed using structured light (Structure from Motion, SfM) technology.
[0066] Specifically, a feature extraction algorithm (such as SIFT, ORB or SURF) is used to detect and describe key feature points of images from each perspective. A feature point matching algorithm (such as a robust matching method based on RANSAC) is used to match feature points from different perspectives to obtain corresponding feature point pairs. A sparse point cloud model is obtained by triangulating and reconstructing the geometric relationship of paired matching points and performing global optimization.
[0067] In the sparse point cloud model of this embodiment, each point is a key point in the target scene (for example, if there is a refrigerator in the scene, the sparse point cloud may include key points such as the four corners of the refrigerator or the textured area in the middle). In this sparse point cloud model, each point includes coordinate attributes and color attributes. The coordinate attributes are the coordinates X, Y, and Z, and the color attributes are the color values R (Red), G (Green), and B (Blue).
[0068] Step S2200, voxelizing the sparse point cloud model to obtain a 4D anchor point model.
[0069] The sparse point cloud model is voxelized, specifically, the appropriate voxel size is selected according to factors such as the density of the point cloud in the sparse point cloud model, application requirements, and computing resources; the three-dimensional space is divided into a uniform voxel grid with the selected voxel size as a unit; each point in the sparse point cloud model is assigned to a corresponding voxel; the attributes of the voxel are calculated according to information such as the number and coordinates of the points falling into each voxel; and a voxel cloud representation is generated according to the attribute information of the voxel.
[0070] In this embodiment, for each coordinate dimension i, the closest distance dp of each point to the corresponding dimension of other points can be calculated from the sparse point cloud model, and the dp values of 60% of the positions are selected according to the dp arranged from large to small as the voxel resolution V of the voxel grid in the i dimension. i , to ensure that the anchor point distribution is well representative of the scene space and reduce memory overhead.
[0071] For example, if the scene bounding box has an extent of [X min ,X max ],[Y min ,Y max ],[Z min ,Z max ], the selected voxel resolution is (V x ,V y ,V z ), then the size of each voxel is d x =(X max -X min ) / V , Dy =(Y max -Y min ) / V , D z =(Z max -Z min ) / V z .
[0072] For each voxel unit, an anchor point is placed at its center, and the shooting time corresponding to the point located in the voxel unit in the sparse point cloud model is used as the time attribute of the corresponding anchor point, so as to obtain a 4D anchor point model.
[0073] In the 4D anchor model, each anchor point can be represented by coordinates X, Y, Z and time attributes.
[0074] This embodiment can normalize and spatially constrain the initial anchor point positions through voxelization, thereby achieving preliminary equalization of the anchor point distribution. In addition, since the amount of sparse point cloud data is usually large, voxelization can compress the data to a certain extent, reduce the burden of data storage and processing, and retain the basic geometric features of the point cloud. Moreover, the voxelized point cloud has a regular spatial structure, which makes it easier to perform various subsequent processing operations, thereby improving the efficiency and stability of the algorithm.
[0075] Step S2300, converting the 4D anchor point model into a 4D Gaussian point cloud model.
[0076] In one embodiment of the present disclosure, converting the 4D anchor point model into a 4D Gaussian point cloud model includes steps S2310 to S2330 as shown below:
[0077] Step S2310, initializing the attribute feature vector of each anchor point in the 4D anchor point model.
[0078] In this embodiment, the dimension and initialization setting strategy of the attribute feature vector can be set according to specific implementation requirements. For example, the attribute feature vector can be set to 32 dimensions, and the value of each dimension can be randomly initialized from a standard Gaussian distribution (mean 0, variance 1), or randomly sampled from a uniform distribution of [-0.1, 0.1].
[0079] For example, if the attribute feature vector f of each anchor point is defined as a one-dimensional vector with a length of 32, an initial value can be generated for each dimension through a pseudo-random number generator. For example, the attribute feature vector f of the i-th anchor point i =[f i1 ,f i2 ,…,f i32 ], where each f ijAll of them can be randomly selected from the uniform distribution U(-0.1,0.1). In the subsequent training process, these features will be continuously iteratively adjusted through optimization algorithms such as gradient descent to adapt to the subsequent 4D Gaussian representation and rendering tasks.
[0080] Step S2320, processing the attribute feature vector of each anchor point in the 4D anchor point model based on a fully connected neural network to obtain 4D Gaussian parameters of multiple Gaussian bodies corresponding to the anchor points.
[0081] By inputting the attribute feature vector f of each 4D anchor point into the fully connected neural network (Multi-Layer Perceptron, MLP), the 4D Gaussian parameters of the corresponding anchor point can be obtained.
[0082] The fully connected network is composed of multiple layers of linear transformations (i.e., weight matrices and bias terms) and nonlinear activation functions. Assume that each anchor point has a feature vector where d f is the dimension of the attribute feature vector. The attribute feature vector f is used as the input of the MLP, and after being processed by several layers of perceptrons (such as 3 to 5 layers, each layer has a certain number of neurons, such as 256 or 512), the 4D Gaussian parameters of the corresponding anchor point are output.
[0083] In this embodiment, the MLP can decode the attribute feature vector f of each 4D anchor point into all Gaussian attributes of five 4D Gaussians (4DGaussians).
[0084] In one embodiment of the present disclosure, the Gaussian parameters include at least one of opacity (Opacity, o), color (Color, c), size (Scale, s), rotation (Rotation, q), and center position (Center, μ).
[0085] Opacity indicates the degree of light penetration of the 4D Gaussian body.
[0086] The color parameter represents the color distribution characteristics of the 4D Gaussian body at this space-time position, that is, the RGB value.
[0087] The size parameter represents the 4D length scale vector of the 4D Gaussian, which is used to control the extent of the expansion of the Gaussian distribution in the 4D space.
[0088] The rotation parameters represent the rotation parameters of the 4D Gaussian body in 4D space, which can be represented by a 4D rotor or other 4D rotation representation methods (such as and meet the normalization requirements).
[0089] The center position represents the center coordinates of the 4D Gaussian body. Contains 3D space coordinates and time coordinates.
[0090] In this process, each 4D anchor point outputs several sets of parameters corresponding to it through MLP, so that one anchor point corresponds to multiple 4D Gaussian volumes. In this way, one anchor point f is expanded to generate multiple sets of 4D Gaussian parameters, providing rich primitives for subsequent 4D scene representation.
[0091] In this embodiment, the fully connected neural network can obtain the 4D Gaussian parameters of the corresponding anchor point according to the attribute feature vector of each anchor point in the 4D anchor point model, and the 3D spatial coordinates and time coordinates of each anchor point in the 4D anchor point model.
[0092] Based on the fully connected neural network, the attribute feature vector of each anchor point in the 4D anchor point model is processed to obtain the 4D Gaussian parameters of the five Gaussian bodies corresponding to the anchor points.
[0093] In an example, if a 4-layer MLP is selected, the number of neurons in each layer is 256, and a nonlinear activation function is added after each layer, the input is (Assuming the attribute feature dimension is 32), after MLP, the output vector of 4D Gaussian attribute parameters is obtained.
[0094] Step S2330, obtaining a 4D Gaussian point cloud model according to the 4D Gaussian parameters of each Gaussian body.
[0095] In this embodiment, the 4D Gaussian parameters of all Gaussian bodies constitute a 4D Gaussian point cloud model.
[0096] Step S2400: determining a target time slice corresponding to a target shooting moment in a target time interval.
[0097] In this embodiment, the target time slice can be obtained by tracing back the first time length and tracing back the second time length based on the target shooting time, wherein the first time length and the second time length can be the same or different, which is not limited here. For example, the first time length and the second time length can both be 1 second.
[0098] In one embodiment of the present disclosure, the target shooting moment may be a shooting moment in a target time interval specified by the user according to his / her own needs, and the corresponding target time slice is determined according to the specified target shooting moment.
[0099] In another embodiment of the present disclosure, determining a target time slice corresponding to a target shooting moment in a target time interval includes: traversing the shooting moments in the target time interval and taking the currently traversed shooting moment as the target shooting moment; and determining the target time slice corresponding to the target shooting moment.
[0100] Step S2500, converting the 4D Gaussian point cloud model under the target time slice into a 3D Gaussian point cloud model at the target shooting moment, and rendering an image by projecting multiple 3D Gaussian points in the 3D Gaussian point cloud model onto the camera imaging plane to train the 3D Gaussian point cloud model.
[0101] In one embodiment of the present disclosure, a 4D Gaussian point cloud model under a target time slice is converted into a 3D Gaussian point cloud model at the target shooting moment, including: determining a 3D Gaussian distribution at the target shooting moment according to a 4D Gaussian parameter of the 4D Gaussian point cloud model under the target time slice; and obtaining a 3D Gaussian point cloud model at the target shooting moment according to the 3D Gaussian distribution at the target shooting moment.
[0102] In this embodiment, the 3D Gaussian distribution of the target shooting moment may be obtained according to the 4D Gaussian parameters of the time coordinate under the target time slice.
[0103] The smaller the difference between the time coordinate and the target shooting time, the greater the contribution of the Gaussian body with higher opacity to the 3D representation of the target shooting time, while the Gaussian body with a time coordinate far away from the target shooting time has less influence on the presentation of the target shooting time.
[0104] Specifically, for each 4D Gaussian (μ, Σ, o, c) (where μ is the center position parameter, Σ is the 4D covariance matrix, which can be synthesized by the scale parameter s and the rotation parameter q, o is the opacity, and c is the color parameter), the 3D Gaussian is obtained by the following steps: According to the 4D Gaussian point cloud model of the time coordinate under the target time slice, the center position and covariance of the 3D Gaussian at the target shooting time are calculated (that is, the conditional distribution of the 4D Gaussian is solved to obtain the 3D distribution). The obtained 3D Gaussian will have the corresponding 3D position, scale and rotation parameters, as well as the apparent characteristics determined by the color and opacity.
[0105] If MLP processes the attribute feature vector of a 4D anchor point into five 4D Gaussian volumes, then, after a time slicing operation at the target shooting moment, five corresponding 3D Gaussian volumes can be obtained as the basic elements for representing the target scene at the target shooting moment.
[0106] In one embodiment of the present disclosure, the method further includes: obtaining true values for expressing the attribute information of points on objects in the target scene at corresponding viewpoints from multi-view images; in the process of training the 3D Gaussian point cloud model, obtaining calculated values of the attribute information of points on the object represented by multiple pixels in the rendered image, using the corresponding true values as optimization targets for the calculated values of the acquired attribute information, supervising the training process of the 3D Gaussian point cloud model, and completing the dynamic modeling process of the target scene.
[0107] In this embodiment, a given camera perspective in a multi-perspective image is used to perform 3D Gaussian rendering, and a rendering technology based on alpha-blending is used to obtain a rendered image.
[0108] Among them, alpha-blending rendering is a semi-transparent overlay technology based on depth sorting. The 3D Gaussians distributed along the ray direction are accumulated and superimposed according to the front-back relationship, and opacity and color are fused in a weighted manner to generate the final pixel color. Specifically, given a projection matrix of a camera view image, the 3D Gaussian is projected onto a 2D plane and forward accumulated according to its opacity and color attributes (e.g., ∑c i α i ∏ j<i (1-α j )), and finally obtain a rendered image corresponding to the real camera perspective.
[0109] In one embodiment of the present disclosure, the attribute information includes at least one of pixel value, brightness, contrast, and structure.
[0110] Specifically, the corresponding true value is used as the optimization target of the calculated value of the acquired attribute information, and the training process of the 3D Gaussian point cloud model is supervised. The loss function between the rendered image and the real image of the data set can be calculated, and the 3D Gaussian point cloud model is trained with the goal of minimizing the value of the loss function.
[0111] The loss function of this embodiment includes L1 loss and structural similarity index (SSIM) loss.
[0112] Specifically, the color attributes of the corresponding pixels of the rendered image R and the real image I are compared point by point to calculate the L1 loss Where N is the total number of pixels in the image, R p is the color attribute of the pth pixel in the rendered image, I p is the color attribute of the p-th pixel in the real image.
[0113] The similarity between the rendered image and the real image is evaluated by comparing brightness, contrast, and structure. SSIM loss can be defined The closer the SSIM value is to 1, the more similar the rendered image R is to the real image I.
[0114] The final total loss can be a weighted combination of the two losses, such as where λ 1 is the weight parameter.
[0115] Based on the calculated loss, the gradient descent method is used to back-propagate the gradient and update the parameters of the 3D Gaussian under the target time slice.
[0116] The loss is derived with respect to the 3D Gaussian parameters to obtain gradient information. The gradient is used to make minor adjustments to parameters such as opacity, color, scale, rotation, and position, so that the next rendering result is closer to the real image, thereby gradually optimizing the 3D Gaussian point cloud model.
[0117] The optimization process is extended from a single time slice to multiple time slices and gradually traced back to the 4D representation level, thereby completing the optimized reconstruction of the entire dynamic scene.
[0118] Specifically, step S2500 may be performed for different time slices to continuously accumulate optimizations of the 3D Gaussian point cloud model in each time slice.
[0119] In one embodiment of the present disclosure, training a 3D Gaussian point cloud model includes optimizing a fully connected neural network and an attribute feature vector of each anchor point in a 4D anchor point model.
[0120] When the 3D Gaussian point cloud model is optimized at each shooting moment, these updated parameters will be used to update the 4D Gaussian point cloud model (i.e., reverse mapping and adjustment of the 4D Gaussian point cloud model). Subsequently, the gradient is further back-propagated to the 4D anchor point and MLP network:
[0121] The attribute feature vector f of the 4D anchor point is adjusted so that the anchor point feature can more effectively generate a 4D Gaussian point cloud model that matches the real scene.
[0122] The parameters of the MLP are updated so that the MLP can output more accurate 4D Gaussian parameters according to the updated attribute feature vector f in the next iteration. After multiple rounds of iteration and update, the optimized 4D anchor points and MLP network parameters can be finally obtained, so that the 4D Gaussian point cloud model generated therefrom can better match the real image when projected as a 3D Gaussian point cloud model at any time.
[0123] Through this embodiment, the optimization of the 3D Gaussian point cloud model from the local time slice is realized, and it is gradually traced back to the representation level of the 4D Gaussian point cloud model, and the attribute feature vector and the MLP network of the anchor point are optimized as a whole, thereby completing the high-quality reconstruction and representation of the dynamic three-dimensional scene.
[0124] In one embodiment of the present disclosure, the method further includes: obtaining an image sequence of the target scene under a test perspective within a target time interval based on an optimized attribute feature vector of each anchor point in the 4D anchor point model and an optimized fully connected neural network.
[0125] Through this embodiment, based on the obtained reconstructable dynamic three-dimensional scene, a high-fidelity image similar to the real scene can be rendered when any time slice and camera viewing angle are given.
[0126] Through the embodiments of the present invention, a sparse point cloud model of the target scene is constructed according to the multi-view images associated with the target scene; the sparse point cloud model is voxelized to obtain a 4D anchor point model; the 4D anchor point model is converted into a 4D Gaussian point cloud model; the target time slice corresponding to the target shooting moment in the target time interval is determined; the 4D Gaussian point cloud model under the target time slice is converted into a 3D Gaussian point cloud model of the target shooting moment, and a rendered image is obtained by projecting multiple 3D Gaussian points in the 3D Gaussian point cloud model onto the camera imaging plane, and the 3D Gaussian point cloud model is trained, so that the three-dimensional scene reconstruction method only needs to store a relatively small number of 4D anchor points and a small fully connected network, thereby reducing the storage overhead of dynamic scene reconstruction, and can effectively reduce the storage requirement expenditure for reconstructing three-dimensional dynamic scenes using Gaussian technology, while also ensuring the rendering quality and rendering speed that Gaussian technology should have, and improving the availability of technology based on Gaussian reconstruction of three-dimensional dynamic scenes.
[0127] <Example>
[0128] Figure 3 FIG. 1 is a flow chart of an example of a method for reconstructing a three-dimensional scene according to an embodiment of the present disclosure. Figure 3 As shown, the three-dimensional scene reconstruction method may include the following steps:
[0129] Step S3001, obtaining a multi-view image associated with a target scene, where the multi-view image is an image sequence obtained by shooting the target scene from multiple perspectives within a target time interval.
[0130] Step S3002: construct a sparse point cloud model of the target scene based on the multi-view images associated with the target scene.
[0131] Step S3003, voxelize the sparse point cloud model to obtain a 4D anchor point model.
[0132] Step S3004, initializing the attribute feature vector of each anchor point in the 4D anchor point model.
[0133] Step S3005 , processing the attribute feature vector of each anchor point in the 4D anchor point model based on a fully connected neural network to obtain 4D Gaussian parameters of multiple Gaussian volumes corresponding to the anchor points.
[0134] Step S3006, obtaining a 4D Gaussian point cloud model according to the 4D Gaussian parameters of each Gaussian body.
[0135] Step S3007, traverse the shooting moments in the target time interval, and use the currently traversed shooting moment as the target shooting moment.
[0136] Step S3008, determining the target time slice corresponding to the target shooting moment.
[0137] Step S3009: convert the 4D Gaussian point cloud model under the target time slice into a 3D Gaussian point cloud model at the target shooting moment.
[0138] Step S3010: Rendering an image by projecting multiple 3D Gaussian points in the 3D Gaussian point cloud model onto a camera imaging plane.
[0139] Step S3011, obtaining calculated values of attribute information of points on the object represented by a plurality of pixel points in the rendered image.
[0140] Step S3012, obtaining from the multi-view images the true value of the attribute information of the points on the objects in the target scene at the corresponding viewpoints.
[0141] Step S3013, training the 3D Gaussian point cloud model by taking the corresponding true value as the optimization target of the calculated value of the acquired attribute information.
[0142] Step S3014, determining whether the shooting moments in the target time interval have been traversed, if so, executing step S3015; if not, continuing to execute step S3007.
[0143] Step S3015, completing the dynamic modeling process of the target scene, and optimizing the attribute feature vector of each anchor point in the fully connected neural network and the 4D anchor point model.
[0144] <Device Example>
[0145] This embodiment provides a three-dimensional scene reconstruction device, such as Figure 4 As shown, the three-dimensional scene reconstruction device 4000 includes a sparse point cloud construction module 4100, a point cloud voxelization module 4200, an anchor point Gaussian transformation module 4300, a time slice determination module 4400 and a Gaussian model training module 4500.
[0146] The sparse point cloud construction module 4100 is used to construct a sparse point cloud model of the target scene based on the multi-view images associated with the target scene; the multi-view images are image sequences obtained by shooting the target scene from multiple perspectives within a target time interval.
[0147] The point cloud voxelization module 4200 is used to perform voxelization processing on the sparse point cloud model to obtain a 4D anchor point model.
[0148] The anchor point Gaussian conversion module 4300 is used to convert the 4D anchor point model into a 4D Gaussian point cloud model.
[0149] The time slice determination module 4400 is used to determine a target time slice corresponding to a target shooting moment in the target time interval.
[0150] The Gaussian model training module 4500 is used to convert the 4D Gaussian point cloud model under the target time slice into a 3D Gaussian point cloud model at the target shooting moment, and train the 3D Gaussian point cloud model by projecting multiple 3D Gaussian points in the 3D Gaussian point cloud model onto the camera imaging plane to obtain a rendered image.
[0151] In one embodiment of the present disclosure, the 3D scene reconstruction device 4000 further includes:
[0152] A true value acquisition module, used to acquire, from the multi-view images, true values for expressing attribute information of points on objects in the target scene at corresponding viewing angles;
[0153] The Gaussian model training module 4500 is used to obtain the calculated values of the attribute information of the points on the object represented by multiple pixels in the rendered image during the training of the 3D Gaussian point cloud model, and to use the corresponding true value as the optimization target of the calculated value of the obtained attribute information, so as to supervise the training process of the 3D Gaussian point cloud model and complete the dynamic modeling process of the target scene.
[0154] In one embodiment of the present disclosure, the attribute information includes at least one of pixel value, brightness, contrast, and structure.
[0155] In one embodiment of the present disclosure, the anchor point Gaussian transformation module 4300 is used to:
[0156] Initializing the attribute feature vector of each anchor point in the 4D anchor point model;
[0157] Processing the attribute feature vector of each anchor point in the 4D anchor point model based on a fully connected neural network to obtain 4D Gaussian parameters of multiple Gaussian bodies corresponding to the anchor points;
[0158] The 4D Gaussian point cloud model is obtained according to the 4D Gaussian parameters of each Gaussian body.
[0159] In one embodiment of the present disclosure, the Gaussian model training module 4500 is used to:
[0160] Determine the 3D Gaussian distribution of the target at the time of shooting according to the 4D Gaussian parameters of the 4D Gaussian point cloud model at the target time slice;
[0161] A 3D Gaussian point cloud model of the target at the shooting moment is obtained according to the 3D Gaussian distribution of the target at the shooting moment.
[0162] In one embodiment of the present disclosure, the Gaussian parameters include at least one of opacity, color, size, rotation, and center position.
[0163] In one embodiment of the present disclosure, the training of the 3D Gaussian point cloud model includes:
[0164] Optimizing the attribute feature vector of each anchor point in the fully connected neural network and the 4D anchor point model.
[0165] In one embodiment of the present disclosure, determining the target time slice corresponding to the target shooting moment in the target time interval includes:
[0166] Traversing the shooting moments in the target time interval, and taking the currently traversed shooting moment as the target shooting moment;
[0167] Determine a target time slice corresponding to the target shooting moment.
[0168] In one embodiment of the present disclosure, the 3D scene reconstruction device 4000 further includes:
[0169] The image sequence generation module is used to obtain an image sequence of the target scene under a test viewing angle within the target time interval according to the optimized attribute feature vector of each anchor point in the 4D anchor point model and the optimized fully connected neural network.
[0170] <Electronic Equipment Embodiment>
[0171] This embodiment provides an electronic device. In one aspect, the electronic device may include the aforementioned in another aspect, such as Figure 5 As shown, the electronic device 5000 may include a processor 5100 and a memory 5200, the memory 5200 is used to store a computer program, and the processor 5100 is used to control the electronic device to execute the method of any embodiment of the present disclosure under the control of the computer program.
[0172] <Readable Storage Medium Embodiment>
[0173] This embodiment provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method described in any method embodiment of the present disclosure is executed.
[0174] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0175] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0176] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0177] The computer program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present invention.
[0178] Various aspects of the present invention are described herein with reference to the flow charts and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each box of the flow chart and / or block diagram and the combination of each box in the flow chart and / or block diagram can be implemented by computer-readable program instructions.
[0179] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0180] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0181] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a part of a module, a program segment or an instruction, and a part of the module, a program segment or an instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that it is equivalent to implement it by hardware, implement it by software, and implement it by combining software and hardware.
[0182] Embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the marketplace, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. A three-dimensional scene reconstruction method, characterized in that: include: Constructing a sparse point cloud model of the target scene according to the multi-view images associated with the target scene; The multi-view image is an image sequence obtained by shooting the target scene from multiple perspectives within a target time interval; voxelize the sparse point cloud model to obtain a 4D anchor point model; Convert the 4D anchor point model into a 4D Gaussian point cloud model; Determine a target time slice corresponding to a target shooting moment in the target time interval; The 4D Gaussian point cloud model under the target time slice is converted into a 3D Gaussian point cloud model at the target shooting moment, and a plurality of 3D Gaussian points in the 3D Gaussian point cloud model are projected onto the camera imaging plane to obtain a rendered image, and the 3D Gaussian point cloud model is trained.
2. The method according to claim 1, characterized in that The method further comprises: Acquire, from the multi-view images, true values for expressing attribute information of points on the objects in the target scene at corresponding viewing angles; During the training of the 3D Gaussian point cloud model, the calculated values of the attribute information of the points on the object represented by the multiple pixels in the rendered image are obtained, and the corresponding true values are used as the optimization targets of the calculated values of the obtained attribute information. The training process of the 3D Gaussian point cloud model is supervised to complete the dynamic modeling process of the target scene.
3. The method according to claim 2, characterized in that The attribute information includes at least one of pixel value, brightness, contrast, and structure.
4. The method according to claim 1, characterized in that: The converting the 4D anchor point model into a 4D Gaussian point cloud model comprises: Initializing the attribute feature vector of each anchor point in the 4D anchor point model; Processing the attribute feature vector of each anchor point in the 4D anchor point model based on a fully connected neural network to obtain 4D Gaussian parameters of multiple Gaussian bodies corresponding to the anchor points; The 4D Gaussian point cloud model is obtained according to the 4D Gaussian parameters of each Gaussian body.
5. The method according to claim 4, characterized in that The step of converting the 4D Gaussian point cloud model of the target time slice into a 3D Gaussian point cloud model of the target shooting moment includes: Determine the 3D Gaussian distribution of the target at the time of shooting according to the 4D Gaussian parameters of the 4D Gaussian point cloud model at the target time slice; A 3D Gaussian point cloud model of the target at the shooting moment is obtained according to the 3D Gaussian distribution of the target at the shooting moment.
6. The method according to claim 4, characterized in that The Gaussian parameters include at least one of opacity, color, size, rotation, and center position.
7. The method according to claim 4, characterized in that The training of the 3D Gaussian point cloud model comprises: Optimizing the attribute feature vector of each anchor point in the fully connected neural network and the 4D anchor point model.
8. The method according to claim 7, characterized in that The method further comprises: According to the optimized attribute feature vector of each anchor point in the 4D anchor point model and the optimized fully connected neural network, an image sequence of the target scene under the test viewing angle within the target time interval is obtained.
9. A three-dimensional scene reconstruction device, characterized in that: include: A sparse point cloud construction module, used to construct a sparse point cloud model of the target scene according to the multi-view images associated with the target scene; The multi-view image is an image sequence obtained by shooting the target scene from multiple perspectives within a target time interval; A point cloud voxelization module, used for voxelizing the sparse point cloud model to obtain a 4D anchor point model; An anchor point Gaussian conversion module, used to convert the 4D anchor point model into a 4D Gaussian point cloud model; A time slice determination module, used to determine a target time slice corresponding to a target shooting moment in the target time interval; The Gaussian model training module is used to convert the 4D Gaussian point cloud model under the target time slice into a 3D Gaussian point cloud model at the target shooting moment, and to train the 3D Gaussian point cloud model by projecting multiple 3D Gaussian points in the 3D Gaussian point cloud model onto the camera imaging plane to obtain a rendered image.
10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the method according to any one of claims 1 to 8 under the control of the computer program.
Citation Information
Cited By
Three-dimensional scene reconstruction method based on Gaussian splashing
CN120495543A
Monocular video scene dynamic three-dimensional reconstruction method based on optical flow
CN120747366A
Semantic segmentation method, electronic equipment and storage medium
CN120876845A